martincousseau.com

Method #3, Evaluation harnesses, 2 min read

Evaluation harnesses

Repeatable scoring in CI. Judges calibrated. Gates that fire before the answer leaves the building.

IN SHORT

A harness is worth it when someone reruns it.

Dithered 1-bit test harness of linked stages illustration for the evaluation harnesses guide
PLATE 1. Generated plate · 1-bit · harness

01 · In brief

#

What it is

A harness makes evaluation a regression signal: the same dimensions, run before merge, nightly, and before release. Judges are calibrated against humans. The gate is a decision, not a dashboard.

This is for teams who already generate outputs and need a harness they can rerun, in CI if they have it or as a scripted run if they do not, without marrying a single vendor.

02 · Stack

#

Tool-neutral, on purpose

DeepEval, Braintrust, Arize Phoenix, LangSmith, RAGAS, Vertex Evaluate, custom Python: I select and configure what fits the failure class and your stack. Independence means no preferential relationship. The architecture should still stand if the tool changes.

Deterministic checks belong in code. Qualitative known failures belong to an LLM-as-judge, itself monitored for drift. Trajectories belong to an agent-judge or a human. Matching the evaluator to the failure is the work; the library is an implementation detail.

I will not pretend that a single framework covers faithfulness, policy, and tool use equally well. The harness is a composition. If a vendor later wraps the same composition, you still own the gates.

If you do not have CI yet, the first harness is a scripted run you can execute again. A YAML file and a command that prints a slice report is a harness. A screenshot of a vendor dashboard is not.

03 · Judges

#

Calibrate the judge

An uncalibrated judge is another fluent system. I design the judge prompt, measure it against human scores, and keep measuring it. When the judge drifts, the harness has failed even if the candidate model has not.

Calibration is versioned with the rest of the stack. A judge prompt change is a model change. It goes through the same gate it is meant to enforce, against a frozen human sample.

Judge cost sits next to human cost. A cheap judge that disagrees with experts on the critical slice is not cheaper. I pick the mix so the gate you actually own is one you can afford to rerun.

04 · Cadence

#

In the loop

Pre-merge for cheap checks, nightly for the full golden set, pre-release for the gate that leadership actually owns. Dashboards and alerts exist so a miss is an incident, not a surprise in production. Cost and latency sit next to quality; a cheaper model that fails the rubric is not cheaper.

The first working gate is usually a pilot: one use case, one golden set, one release policy. After that I make the gate ordinary: several use cases, CI that developers actually run, a scorecard a risk reader can open.

The scorecard names the slice that failed, not only a global number. Product, engineering, and whoever owns the release read the same page. A metric that only the evaluation person understands will not hold a ship decision.

05 · Artifacts

#

Where the harness runs

Cadence is a starting point. Risk sets the clock.

GateWhat it is for
Pre-mergeDeterministic checks and cheap judges on touched paths
NightlyFull golden set, slice report, judge health
Pre-releaseThe written threshold: ship or hold
Judge monitorAgreement with humans; drift of the evaluator itself

06 · Engagement

#

How I usually run it

I leave the harness as a protocol you can rerun on the next release. When it is handed over to your team, it comes as an eval/ folder they run in CI.