martincousseau.com

Method #5, Agent evaluation, 2 min read

Agent evaluation

Process quality: tool calls, recovery, hand-offs, and then the final answer.

IN SHORT

Score the path, not only the final answer.

Dithered 1-bit graph of linked nodes illustration for the agent evaluation guide
PLATE 1. Generated plate · 1-bit · graph

01 · In brief

#

What it is

Agents fail in the process: the wrong tool, the wrong argument, no recovery, a bad hand-off. A final-answer score will bless a trajectory that should never have been allowed to finish.

This is for teams whose agent experiments have matured enough to score tool calls and recovery. I start with the core feature, unless the agent is already the job.

02 · Trajectories

#

Score the path

I evaluate tool selection, argument correctness, use of retrieved context, and what happens after an error. Task completion is necessary and not sufficient. Intermediate state, whether the agent knew it was lost, is part of the rubric.

Traces are the unit of work. A last-message corpus will not show a forbidden call that was later overwritten by a polite summary. If you cannot replay the path, you cannot score it.

A RAG faithfulness score on the final sentence can still bless a trajectory that queried the wrong system, looped on an empty retrieval, or skipped the human hand-off. Agents fail in the process. That is why this guide exists.

03 · Judges

#

Judges that can see a trace

Agent-as-judge designs are used when a single-turn LLM-judge cannot see the process. They are calibrated like any other judge. Humans remain on the trajectories that would be indefensible to wave through.

A process judge can still be fluent and wrong. I hold it to a human sample of traces, including recoveries and refusals, not only the successes a demo would pick.

Replay is the offline set: held-out tasks, frozen tools, a trace you can run again. Live sampling of production trajectories is continuous evaluation. Do not mix the two numbers on one chart without saying so.

04 · Hand-offs

#

Coordination is a failure surface

Multi-agent systems add hand-offs. I score whether the next agent received what it needed, whether ownership was clear, and whether the system stopped when it should have asked a person.

Escalation is a dimension. An agent that “finishes” by guessing instead of handing to a human has failed the threshold even if the answer happens to be right.

When the agent is the job, the work still produces a rubric, a replay set, and a gate. A demo of tool calling is not a gate.

05 · Artifacts

#

Dimensions on a trajectory

Not every agent needs every row. The failure you name chooses.

DimensionWhat is scored
Tool selectionThe right tool, at the right step
ArgumentsParameters match the schema and the intent
RecoveryBehaviour after a tool error or empty retrieval
CompletionThe task actually finished, on the evidence
Hand-offState and ownership across agents or humans

06 · Engagement

#

How I usually run it

I focus on one agent family, once the core feature has already been evaluated. I test through a guest seat or a scoped staging key, never a private key.