martincousseau.com

7 GUIDES

How I work

// one conclusion each

I build LLM systems and the evaluations that show whether they are ready to ship. I name the failure before the metric, run code checks before asking a model for an opinion, and read the rows the average hides. Each guide below starts from the one-line conclusion it argues for.

  1. #1 · STRATEGYGUIDE

    Evaluation strategy

    Start from the failure you fear, then write the line.

    Fig. 1.1 · illustration, not data

  2. #2 · DATASETSGUIDE

    Golden datasets

    A small set with honest gold beats a large set nobody checked.

    Fig. 1.2 · illustration, not data

  3. #3 · HARNESSESGUIDE

    Evaluation harnesses

    A harness is worth it when someone reruns it.

    Fig. 1.3 · illustration, not data

  4. #4 · HUMANGUIDE

    Human evaluation

    People read the rows a metric can't.

    Fig. 1.4 · illustration, not data

  5. #5 · AGENTSGUIDE

    Agent evaluation

    Score the path, not only the final answer.

    Fig. 1.5 · illustration, not data

  6. #6 · CONTINUOUSGUIDE

    Continuous evaluation

    Re-check when the model, prompt or data moves.

    Fig. 1.6 · illustration, not data

  7. #7 · RISKGUIDE

    Risk & compliance

    Evidence your risk team can read, not a badge.

    Fig. 1.7 · illustration, not data

Contact

// one line, one action

Tell me what you’re about to ship.

contact@martincousseau.com

$ mail contact@martincousseau.com

Warsaw · English / French · open to B2B engagements