martincousseau.com

Method #6, Continuous evaluation, 2 min read

Continuous evaluation

A living scorecard. The same dimensions, rerun as models, prompts, and traffic drift.

IN SHORT

Re-check when the model, prompt or data moves.

Dithered 1-bit pulse line illustration for the continuous evaluation guide
PLATE 1. Generated plate · 1-bit · pulse

01 · In brief

#

What it is

Offline evaluation is necessary and not sufficient. Models, prompts, retrieval, and traffic move. Continuous evaluation is the same report, run again.

This is for teams that already have a first golden set and a gate. It is rarely the first run. A rerun of the existing evaluation on the same system is how it usually starts.

02 · Sampling

#

Sample the live system

I design sampling that is cheap enough to run and representative enough to trust, including the slices that matter. Online scoring reuses the same dimensions as the offline set. A second, prettier dashboard is not continuity.

Production is a different distribution. The golden set stays the reference; the live sample is how you learn that the reference is no longer the job. Both stay in the scorecard.

Sampling is a data decision. Rate, retention, and who may see live outputs follow your existing rules, never a vendor’s default. A live sample that cannot be destroyed on request is not a sample you should have taken.

03 · Drift

#

Drift is a decision

Quality drift, behavioural drift, and judge drift are distinguished. An alert without an owner is décor. Signals connect to incident response: investigate, hold, roll back the prompt, the model, or the retrieval.

A quiet move in one slice (one language, one product line, one tool) is the usual incident. Continuity that only watches the global average will miss it.

Observability tells you latency, cost, and whether a trace existed. Continuous evaluation scores the output against the same rubric you used to ship. You need both. One dashboard that mixes them will hide the miss.

04 · Trade-offs

#

Cost, latency, and judgment together

A model that is cheaper and slower to fail is not an improvement. Joint monitoring keeps the trade-off visible, so a cost optimisation cannot silently spend the threshold.

Reruns are how this stays alive: new use cases, new models, the same dimensions. The first written result is a beginning. Drift is the rest of the work.

This is not a monitoring product. I stand the first evaluation up first. Then I scope the cadence as reruns: the same dimensions, run again, with an owner for the alert.

A weekly slice report with a named owner beats a real-time wall nobody reads. Cadence follows risk. The point is that last month’s threshold still means something this month.

05 · Artifacts

#

What stays alive

The living scorecard. Same dimensions as the offline rubric.

SignalWhat it is for
Production sampleThe distribution you actually serve
Slice reportWhere quality moved, not only that it moved
Drift alertA change large enough to own
Rollback pathPrompt, model, index: who may revert it

06 · Engagement

#

How I usually run it

I rerun the evaluation on the same system after the first result: a new slice, or a check after fixes. When the work is handed over, an eval/ folder lets your team run it between my reruns.