martincousseau.com

Method #4, Human evaluation, 2 min read

Human evaluation

Calibrated expert review where automation is not enough, with agreement measured.

IN SHORT

People read the rows a metric can't.

Dithered 1-bit panel of reviewers illustration for the human evaluation guide
PLATE 1. Generated plate · 1-bit · panel

01 · In brief

#

What it is

Where automation cannot see the failure, a human must: trained, calibrated, and measured. Unstructured eyeballing is not a programme. Agreement is the score of the scoring.

This is for dimensions where automation is not enough: preference, nuance, or the slice an LLM-judge keeps missing. I design the programme as part of the evaluation; your subject-matter experts usually rate.

02 · Raters

#

Guidelines before raters

A rater without a guideline is improvising. I write the guide from the rubric, train raters on it, and keep a calibration set that is scored again as people drift. New raters do not join a live queue until they match the standard.

Guidelines are short enough to use under time pressure and specific enough to survive a disagreement. If two trained people cannot apply a sentence the same way, the sentence is wrong.

The guide is an artifact your experts keep. It names examples of pass, hold, and fail for each dimension, with no paragraph of theory. If a new rater cannot use it on day one, it is not finished.

03 · Agreement

#

Agreement is part of the ledger

Inter-annotator agreement is tracked. Disagreements are adjudicated, not averaged away. Preference ranking and pairwise comparison are used when the question is “which is better”, not “is this true”. The method follows the decision.

Adjudication is logged. A later reader should see why the gold label is the gold label. Hidden consensus is how human evaluation becomes theatre.

Agreement is not a vanity statistic. If two trained people split on a policy item, the release gate cannot pretend the score is settled. That item is a hold until the guide or the model changes.

04 · Design

#

Hybrid by default

Humans are expensive. The design is usually hybrid: automatic scores on the bulk, human review on a sample, on disagreements, and on the critical slice. Cost is an evaluation parameter, not an embarrassment.

The sample is not random courtesy. It is stratified by slice and by the auto-score’s uncertainty. That is how a small panel still sees the cases that would fail a gate.

Your people usually rate. I design the programme, the calibration loop, and the report. I am not a labelling farm, and I do not take a cut from a crowd vendor. The capability is meant to stay with you.

The output is a programme you can run again: the guide, the calibration set, the sampling rule, and an agreement report a later auditor can read. A week of unstructured comments in a spreadsheet is not that.

05 · Artifacts

#

What the programme holds

These artifacts are what make a human score repeatable.

ArtifactPurpose
Rater guidelinesThe same definition of the dimension, in writing
Calibration setWhether people still agree with last month
Agreement reportAgreement, adjudication log, rater notes
Sampling ruleWhat the humans see, and what the auto-score covers

06 · Engagement

#

How I usually run it

I design the programme as part of the evaluation, then you keep the cadence. Before I report any result, I read the rows the average hides.