Method #3, Evaluation harnesses, 2 min read
Evaluation harnesses
Repeatable scoring in CI. Judges calibrated. Gates that fire before the answer leaves the building.
A harness is worth it when someone reruns it.
Contents

What it is
A harness makes evaluation a regression signal: the same dimensions, run before merge, nightly, and before release. Judges are calibrated against humans. The gate is a decision, not a dashboard.
This is for teams who already generate outputs and need a harness they can rerun, in CI if they have it or as a scripted run if they do not, without marrying a single vendor.
Tool-neutral, on purpose
DeepEval, Braintrust, Arize Phoenix, LangSmith, RAGAS, Vertex Evaluate, custom Python: I select and configure what fits the failure class and your stack. Independence means no preferential relationship. The architecture should still stand if the tool changes.
Deterministic checks belong in code. Qualitative known failures belong to an LLM-as-judge, itself monitored for drift. Trajectories belong to an agent-judge or a human. Matching the evaluator to the failure is the work; the library is an implementation detail.
I will not pretend that a single framework covers faithfulness, policy, and tool use equally well. The harness is a composition. If a vendor later wraps the same composition, you still own the gates.
If you do not have CI yet, the first harness is a scripted run you can execute again. A YAML file and a command that prints a slice report is a harness. A screenshot of a vendor dashboard is not.
Calibrate the judge
An uncalibrated judge is another fluent system. I design the judge prompt, measure it against human scores, and keep measuring it. When the judge drifts, the harness has failed even if the candidate model has not.
Calibration is versioned with the rest of the stack. A judge prompt change is a model change. It goes through the same gate it is meant to enforce, against a frozen human sample.
Judge cost sits next to human cost. A cheap judge that disagrees with experts on the critical slice is not cheaper. I pick the mix so the gate you actually own is one you can afford to rerun.
In the loop
Pre-merge for cheap checks, nightly for the full golden set, pre-release for the gate that leadership actually owns. Dashboards and alerts exist so a miss is an incident, not a surprise in production. Cost and latency sit next to quality; a cheaper model that fails the rubric is not cheaper.
The first working gate is usually a pilot: one use case, one golden set, one release policy. After that I make the gate ordinary: several use cases, CI that developers actually run, a scorecard a risk reader can open.
The scorecard names the slice that failed, not only a global number. Product, engineering, and whoever owns the release read the same page. A metric that only the evaluation person understands will not hold a ship decision.
Where the harness runs
Cadence is a starting point. Risk sets the clock.
| Gate | What it is for |
|---|---|
| Pre-merge | Deterministic checks and cheap judges on touched paths |
| Nightly | Full golden set, slice report, judge health |
| Pre-release | The written threshold: ship or hold |
| Judge monitor | Agreement with humans; drift of the evaluator itself |
How I usually run it
I leave the harness as a protocol you can rerun on the next release. When it is handed over to your team, it comes as an eval/ folder they run in CI.