Method #2, Golden datasets, 2 min read
Golden datasets
Provenance-tracked test cases from the distribution you actually serve, including the rare and the high-stakes.
A small set with honest gold beats a large set nobody checked.
Contents

What it is
A golden dataset is not a handful of favourite prompts. It is a versioned, provenance-tracked sample of the distribution you will actually serve, including the cases you would rather not look at.
This is for teams whose current eval set is a spreadsheet of happy paths, or who have some production data and need it turned into a set a score can rest on.
From production, with a chain of custody
I sample from logs or historical cases under the data rules you already have. Your data stays in the agreed environment. Anonymisation is part of the artifact, not an afterthought. Every item carries provenance: where it came from, who labeled it, which version of the rubric applied.
The set is sliced. A single average hides the slice that fails. Challenge cases (rare, high-risk, adversarial) are constructed when production does not yet contain them, then validated by a human. Synthetic data is a supplement, never the whole set.
A pipeline audit without a set is a reading of whatever happened to be in the room. The dataset is how that reading becomes repeatable when the model, the prompt, or the index moves.
A download of MMLU or a public RAG bench is not a golden set for your product. Those numbers are useful as a smoke test. They do not know your documents, your users, or the one policy sentence that must never be invented.
Labels you can defend
Ground truth is a claim. I treat it as one: guidelines, more than one annotator where the stake requires it, agreement measured, disagreements adjudicated. The dataset card records that process so a later auditor is not asked to take the score on faith.
Where the label is preference rather than truth (two fluent answers, one better), the set still carries the guideline version and the pair. Rankings without a protocol are taste.
Domain experts label the critical slice. Crowd labels, if you already have them, are a starting point. They do not replace the cases that would fail a release. The card says which is which.
Versioned like software
When the product, the policy, or the model changes, the set changes with it, or it is frozen and named. Access control belongs with the rest of your evaluation assets. A golden set that cannot be reproduced is a demo.
Retention follows your existing records policy. In regulated settings the evaluation set is itself an artifact that may need to be produced later. I write that down at the start, not after a request from audit.
You keep the set. I leave a card, a schema, and a sampling rule your team can extend. A vendor-hosted eval dump you cannot export is not a ledger.
What the set must carry
Minimum fields. Further columns follow the rubric.
| Field | Why it exists |
|---|---|
| Provenance | Source system, time, sampling rule |
| Slice | Happy path, edge, adversarial, demographic |
| Label and guideline version | What “correct” meant on that date |
| Agreement | Where humans disagreed, and how it was closed |
| Access and retention | Who may see it, and when it is destroyed |
How I usually run it
I build the golden set as part of the evaluation itself, including the cases production under-samples. When almost no usable data exists yet, I start by building a first set of cases with sourced gold answers.