Method #6, Continuous evaluation, 2 min read
Continuous evaluation
A living scorecard. The same dimensions, rerun as models, prompts, and traffic drift.
Re-check when the model, prompt or data moves.
Contents

What it is
Offline evaluation is necessary and not sufficient. Models, prompts, retrieval, and traffic move. Continuous evaluation is the same report, run again.
This is for teams that already have a first golden set and a gate. It is rarely the first run. A rerun of the existing evaluation on the same system is how it usually starts.
Sample the live system
I design sampling that is cheap enough to run and representative enough to trust, including the slices that matter. Online scoring reuses the same dimensions as the offline set. A second, prettier dashboard is not continuity.
Production is a different distribution. The golden set stays the reference; the live sample is how you learn that the reference is no longer the job. Both stay in the scorecard.
Sampling is a data decision. Rate, retention, and who may see live outputs follow your existing rules, never a vendor’s default. A live sample that cannot be destroyed on request is not a sample you should have taken.
Drift is a decision
Quality drift, behavioural drift, and judge drift are distinguished. An alert without an owner is décor. Signals connect to incident response: investigate, hold, roll back the prompt, the model, or the retrieval.
A quiet move in one slice (one language, one product line, one tool) is the usual incident. Continuity that only watches the global average will miss it.
Observability tells you latency, cost, and whether a trace existed. Continuous evaluation scores the output against the same rubric you used to ship. You need both. One dashboard that mixes them will hide the miss.
Cost, latency, and judgment together
A model that is cheaper and slower to fail is not an improvement. Joint monitoring keeps the trade-off visible, so a cost optimisation cannot silently spend the threshold.
Reruns are how this stays alive: new use cases, new models, the same dimensions. The first written result is a beginning. Drift is the rest of the work.
This is not a monitoring product. I stand the first evaluation up first. Then I scope the cadence as reruns: the same dimensions, run again, with an owner for the alert.
A weekly slice report with a named owner beats a real-time wall nobody reads. Cadence follows risk. The point is that last month’s threshold still means something this month.
What stays alive
The living scorecard. Same dimensions as the offline rubric.
| Signal | What it is for |
|---|---|
| Production sample | The distribution you actually serve |
| Slice report | Where quality moved, not only that it moved |
| Drift alert | A change large enough to own |
| Rollback path | Prompt, model, index: who may revert it |
How I usually run it
I rerun the evaluation on the same system after the first result: a new slice, or a check after fixes. When the work is handed over, an eval/ folder lets your team run it between my reruns.