Method #1, Evaluation strategy, 3 min read
Evaluation strategy
Failure modes, task taxonomies, rubrics, and the written line between ship, fix these, and defer.
Start from the failure you fear, then write the line.
Contents

What it is
An evaluation strategy is the written system that turns opaque GenAI behaviour into a decision: ship, fix these, or defer. It starts from the failure modes and the decision you must own. The metrics a tool happens to expose come later.
This is for AI-forward startups, small businesses with a first GenAI pipeline, and mid-size teams that need a written line for “good enough to ship”, without a multi-month governance programme.
Six evaluation principles
A named list, so a reader can cite a method rather than a paragraph.
- Start from the failure and the decision. Not from the metrics a tool happens to expose. Name what must never ship, then write the line.
- Match the evaluator to the failure class. Code for deterministic checks, a judge for qualitative known failures, an agent-judge for trajectories, a human for high-stakes nuance.
- Calibrate and version everything. Judges, rubrics, datasets, and thresholds are software artefacts. A judge prompt change is a model change.
- Keep the offline set and the live sample distinct. Both are required. Offline is the reference. The live sample is how you learn the reference is no longer the job.
- Make every score a decision. Ship, fix, defer, or roll back. A metric that cannot issue one of those four is décor.
- Design for auditability. Who authored the rubric, what data was used, how agreement was measured, what the release threshold was. If it cannot be shown, it is not evidence.
Start from the failure, not the score
Public leaderboards do not predict your distribution. A model that looks calm on a general benchmark can still invent a citation, skip a policy, or take the wrong tool on the one case that matters. Strategy work begins with the failures that would actually hurt: product, engineering, risk, and domain experts in the same room, naming what must never ship.
From that list I write a task taxonomy and slices (happy path, edge, adversarial, demographic), so the later dataset and harness have somewhere to live. A score without a slice is a number you cannot act on.
Ownership is named in the same pass. If no one owns the rubric when the model card changes, the strategy was a workshop, not a system. The memo says who may move a threshold, and who must be in the room when they do.
Rubrics and the release gate
The rubric is the judgment, written down. Pointwise or pairwise, pass/fail or graded, single-turn or trajectory: the form follows the failure. Each dimension has a definition, a method, and a threshold. Faithfulness is not groundedness. Calibration is not confidence. Operational risk is not a style score.
The release policy is the line: what is good enough to ship, what must be fixed first, what waits. Decision thresholds live here, in a go-live memo the organisation keeps after I leave.
When the line is not written yet, writing it is the first piece of work: the line, the dimensions agreed, the first cases named. When the line exists, I score a slice against it, and then add the taxonomy and the mapping that let a harness enforce it.
Mapped to regulation and the business
Where the system is high-stakes, metrics map to internal policy, model-risk language, and, where it applies, obligations such as the EU AI Act. The mapping is explicit. A dashboard that cannot be read by legal is not a governance artifact.
I do not issue certificates. I produce the evaluation design and the first evidence pack so your risk function has something other than a vendor narrative. What they then file is theirs.
The pack names the system, the data it saw, who authored the rubric, and who may move a gate. That is evidence. A slide that says “AI Act ready” is not.
What the strategy contains
Typical artifacts. Scope follows the use case, not a template.
| Artifact | What it decides |
|---|---|
| Failure-mode map | Which errors are in scope, and who owns them |
| Task taxonomy and slices | What is tested, including the rare path |
| Rubric specification | Dimensions, methods, and agreement rules |
| Release policy | Ship, fix these, or defer, and when to roll back |
| Metric-to-risk map | How a score becomes a sentence a committee can read |
How I usually run it
When “good enough” is not written down yet, I start by writing it: the line, the dimensions agreed, the first cases named. When the feature is already clear, I write the line at the start of the evaluation.