# Martin Cousseau > Martin Cousseau is an AI Engineer in Warsaw, Poland: "I build LLM systems and the evaluations that tell you whether they are ready to ship." Contact: contact@martincousseau.com. Personal site of Martin Cousseau. French, based in Warsaw, Poland; works in English and French. Pages are written by Martin Cousseau unless a byline says otherwise (the paper lists all of its authors in published order). Other people share the name Martin Cousseau; the profiles listed below identify this one. Site: https://martincousseau.com Contact: contact@martincousseau.com LinkedIn: https://www.linkedin.com/in/cousseaumartin/ GitHub: https://github.com/martin-cousseau Hugging Face: https://huggingface.co/martincousseau Google Scholar: https://scholar.google.com/citations?user=KzMJqUcAAAAJ ORCID: https://orcid.org/0009-0000-5228-3565 arXiv: https://arxiv.org/abs/2505.07289 YouTube: https://www.youtube.com/@martin-cousseau Medium: https://medium.com/@martin-cousseau X: https://x.com/m_cousseau TikTok: https://www.tiktok.com/@martin.cousseau --- # Extraction Arena Source: https://martincousseau.com/research/extraction-arena Type: Note Published: 2026-08-24 (updated 2026-10-05) Original: Can we trust a model with a rescue procedure? — https://x.com/m_cousseau/status/2097740397170634883 Code: https://github.com/martin-cousseau/extraction-arena Topics: extraction, thresholds Conclusion: A strong average can hide a procedure the model gets entirely wrong. Gate the field you fear, not the mean. ## Summary Three vision models read Tesla’s Cybertruck first-responder documentation and filled a structured record. The headline scores look ready to ship (93, 90, 90). The water-rescue procedure does not. I recorded a [walkthrough of the harness](/research/extraction-arena-walkthrough) behind this note, and the code is open source: [martin-cousseau/extraction-arena](https://github.com/martin-cousseau/extraction-arena). I keep the model answers frozen and change only the scoring rule, so I can see whether I have a **model problem** or a **definition problem**. Later I add a fourth pipeline, LlamaParse through [LlamaExtract](https://www.llamaindex.ai/llamaextract), and read that run the same way. The business question does not move. Would you accept a high extraction score if the one procedure a crew would follow still scores **zero**? ## A go / no-go decision, not just a leaderboard For product owners deciding go / no-go, subject-matter experts deciding whether order and wording are allowed to move on a rescue step, or engineers who have to turn that decision into weights. If you ship document extraction into operations, this is the meeting you skip at your cost. The average will look fine. The field a crew would actually follow might not. ## This is not a clean text document Tesla publishes official rescue information for the Cybertruck. A rescue sheet ([ISO 17840-1](https://www.iso.org/standard/78461.html)) is the quick page at the scene. An Emergency Response Guide ([ISO 17840-3](https://www.iso.org/standard/67353.html)) is the longer procedure set used to train and operate. The four pages the models saw mix **diagrams**, **pictograms**, **numbered procedures**, and **prohibitions** in capital letters. A wrong cut zone or a skipped disable step is an operational failure, not just a formatting issue. I selected and scored the following fields: cut zones, high-voltage disable, fire, submersion, and tow. The first experiment ran in four steps. 1. Each page is rendered as an image (the models see what a person sees on paper). 2. Every model gets the same empty form; the correct answers are never in the prompt. 3. **GLM-5V-Turbo**, **GPT-5.4 Mini**, and **Grok 4.5** each return their own JSON; I map those answers onto one shared form. 4. Then I score field by field, because the **business risk is not the same** on manufacturer, on a prohibition, and on a five-step procedure. ## I already have the answers. I am changing the question. I am not asking the models again. I am changing how a correct answer is defined for one field. I zoomed in on the five-step procedure for a submerged Cybertruck. Product owners pick the scoring rule that matches the risk. Subject-matter experts say whether order and wording are allowed to move. How close must the wording be? **Exact or partial**. If too strict, a useful extraction looks like a failure. If too loose, a dangerous paraphrase looks like a pass. Example: **exact match** fails “Wear PPE” against the gold “Wear appropriate PPE for water rescue.” **Partial match** still scores the overlapping words. Does order count? **Sequence or set**. A rescue procedure is a sequence of actions. **Sequence** compares step 1 to gold step 1. **Set** treats the list as a bag. Example: one model opened at gold step 4. Under sequence that is a miss in position 1. Under set, the step is merely present, and a crew could skip PPE. The question, **which mistake to surface first**, is a **product decision**, not a scoring one. **Precision**: of everything the model wrote, how much was on the sheet — the lens for a prohibition. **Recall**: of everything the sheet required, how much the model captured — the lens for a procedure. And **F1** holds both costs. ## Three stories. None of them improved the models. None of the three vision models returned the five-step order. In the illustration below a filled cell is a step the model actually wrote, a fused cell is two gold steps crushed into one line, an empty cell is a skipped step. - One model opened at gold step 4. - The other two started on PPE, fused remove-from-water with the high-voltage disable, then stopped. All three saw the submersion paragraph. None produced a five-step ordered list. I then read those frozen answers under three scenarios. Only the **scoring rule** changes. - Exact wording, order required: precision 0, recall 0, F1 0. - Flexible wording, order still required: 44, 25, 33. - Flexible wording, order ignored: 81, 47, 58. All of these results are true. **They just answer different questions.** I did not improve the models, but I changed the question. For this field a kinder pairing rule increases the score and rewards an extracted procedure no crew should follow. ## The mix still looks shippable. A well-designed metric should let you trust the headline. That is the goal. A composed score only earns that trust if its weights match the risk of the work. But in this demo they do not: manufacturer, vehicle type and other easy fields pull the extraction score up to 93, while the **water-rescue** procedure scores zero under the strict rule. The demo mix is weight 0.25 on the exact-field rate, weight 0.35 on average partial credit, and weight 0.40 on average F1, taken across every field on the sheet. With this mix, **GLM-5V-Turbo** lands at 93. **GPT-5.4 Mini** at 90. **Grok 4.5** at 90. A composed score is the right tool to record a run and follow a trend. It is the wrong tool for a ship / no-ship decision when those weights have not been set to the **operational risk**. Use it to follow a trend across runs or tune the weights until important fields can move the headline. ## The veto test, in five steps. People ask how to “choose weights.” Weights are a claim about which mistake you are willing to average away. Here is the method I use: 1. **Write the feared failure in one sentence**, with a subject. “The extractor skips PPE on a submerged-vehicle procedure.” If you cannot write this, you cannot set weights and you are just collecting scores. 2. **Tag every field Gate, Credit, or Noise.** **Gate**: a wrong answer means do not ship. Often a veto — exact match. **Credit**: partial overlap is allowed and useful (a warning list, a paraphrasable instruction). **Noise**: easy fields that inflate the average (manufacturer, model year, door count). For example, on the Cybertruck sheet: **submersion ordered-steps** is Gate. **Fire prohibition** is Gate. **Manufacturer** is Noise. **Occupant-access** methods are Credit. 3. **Run the veto test.** Set the gate field to zero. Leave every other field at the score you actually observed. Recompute the headline. If that headline still looks like a ship (80, 90, 93) the weights are wrong and they have not been set to the risk. This is the test this demo failed: **procedure at 0, composed extraction score at 93.** 4. **Retune** until a zero on the gate field pulls the headline below the written line. You can raise the gate field’s weight in the average or stop averaging it. For example, across the 34 fields of this demo the procedure has a 5% say. If the procedure gets 40% of the headline, the same run lands near 58 and it no longer looks shippable. **The models did not get worse.** 5. **Write the line as a conjunction.** “We ship if composed ≥ T and the gate field ≥ G.” Product owners pick the conjunction. Subject-matter experts pick G. Engineers encode both. If your current mix cannot fail the veto test you should change the mix (rather than argue about the model). ## The scorer did not blink The first experiment was three vision pipelines on page images. After that I added LlamaExtract, through LlamaParse. LlamaIndex ships four extract modes: **turbo**, **cost-effective**, **agentic**, and **agentic plus**. Same rescue-sheet schema. Same scorer. With the same mix that produced 93, 90, 90 on vision I get the results below. **87%** is the best of the four, but it is lower than every previous vision headline. At least the ranking follows my instinct and picks agentic plus. A ranking against vision still picks GLM-5V-Turbo at 93. The strict procedure still fails on every mode. ## Do not ship on 93 Seven runs. Vision: 93, 90, 90. LlamaParse: 79, 83, 84, 87. Agentic plus emerged as the best of the four LlamaParse pipelines, yet it remains below every vision headline. The exact sequence mode for the **water-rescue** procedure returned **0** on **every single run**. While the Extraction Score showed movement, the deployment gate did not. In conclusion, frame the **release criteria** as a strict conjunction. **Ship only if the headline clears [threshold](/glossary#threshold) T and the gate field clears threshold G.** Display the worst-performing field first on the dashboard. When an 87 baseline defines the trend and the rescue procedure holds the deciding veto, the directive is clear: **do not ship**. ## The walkthrough Three vision models, one gold rescue sheet, and a score that moves when the question changes. The same argument, recorded. There is also a [page for the walkthrough](/research/extraction-arena-walkthrough). ## Extraction Arena: Evaluate Vision LLMs for Document Extraction ## The harness Extraction Arena is a public TypeScript repository, [martin-cousseau/extraction-arena](https://github.com/martin-cousseau/extraction-arena), on the `main` branch: a document extraction evaluation framework. Two Node projects: a backend that turns pages into images, and a frontend that scores fields against gold answers. The note is the argument; the repository is the instrument. ```bash git clone https://github.com/martin-cousseau/extraction-arena.git ``` | Path | What it holds | | -------------------- | ------------------------------------------------------ | | `backend/` | Express. PDF to PNG at 300 DPI, then the vision calls. | | `frontend/` | Datasets, judges, field scores, the comparison. | | `docker-compose.yml` | The whole app in containers. | | `README.md` | Quick start and the JSON contract. | | `AGENTS.md` | Conventions for the repo. | The same piece also ran on [X](https://x.com/m_cousseau/status/2097740397170634883), [LinkedIn](https://www.linkedin.com/pulse/can-we-trust-model-rescue-procedure-martin-cousseau-ugwrf/) and [TikTok](https://www.tiktok.com/@martin.cousseau/video/7682673820016184608). --- # Extraction Arena: Evaluate Vision LLMs for Document Extraction Source: https://martincousseau.com/research/extraction-arena-walkthrough Type: Video Published: 2026-08-21 Video: https://www.youtube.com/watch?v=QXWN8WyvPmI Duration: PT9M43S Topics: extraction, thresholds Conclusion: A walkthrough of the harness behind the note. ## Summary A recorded walkthrough of Extraction Arena, the open-source harness I use to evaluate vision LLMs on document extraction, field by field. It is the same argument as [the Extraction Arena note](/research/extraction-arena), recorded: three vision models, one gold rescue sheet, and a score that moves when the question changes. The code is on [GitHub](https://github.com/martin-cousseau/extraction-arena). ## Extraction Arena: Evaluate Vision LLMs for Document Extraction Also on [YouTube](https://www.youtube.com/watch?v=QXWN8WyvPmI), on my channel [@martin-cousseau](https://www.youtube.com/@martin-cousseau). --- # Semantic Retention and Extreme Compression in LLMs: Can We Have Both? Source: https://martincousseau.com/research/srcr Type: Paper Published: 2025-05-12 (updated 2026-10-05) Authors (published order): Stanislas Laborde, Martin Cousseau, Antoun Yaacoub, Lionel Prevost Venue: IJCNN 2025 arXiv: https://arxiv.org/abs/2505.07289 DOI: https://doi.org/10.1109/IJCNN64981.2025.11227279 PDF: https://arxiv.org/pdf/2505.07289 Topics: compression Conclusion: Introduces SrCr, a way to measure how much meaning survives compression. ## Summary My co-authors and I introduced SrCr, the Semantic Retention Compression Rate: a way to measure how much meaning survives when a model is compressed. SrCr is a metric that quantifies the trade-off between model compression and semantic preservation. In the paper we examine joint compression, that is, how strategically combining pruning and quantization can yield better performance-to-compression ratios than either method alone, and we use SrCr to optimize pruning-quantization configurations. The abstract reports that the authors' recommended combination achieves, on average, a 20% performance increase compared to an equivalent quantization-only model at the same theoretical compression rate. I co-authored the paper with Stanislas Laborde, Antoun Yaacoub and Lionel Prevost; I am the second of the four authors. Stanislas Laborde and I contributed equally, and I started the project. It was accepted at IJCNN 2025, the International Joint Conference on Neural Networks (Rome, Italy, 30 June to 5 July 2025), and posted to arXiv on 12 May 2025. The measure also shaped how I think about evaluation more broadly. The companion note, [What Survives Compression](/research/what-survives-compression), covers that habit: ask what meaning survived before you ask how fluent the remainder sounds. ## Abstract Abstract of the paper as posted on arXiv (2505.07289) by Stanislas Laborde, Martin Cousseau, Antoun Yaacoub and Lionel Prevost: > The exponential growth in Large Language Model (LLM) deployment has intensified the need for efficient model compression techniques to reduce computational and memory costs. While pruning and quantization have shown promise, their combined potential remains largely unexplored. In this paper, we examine joint compression and how strategically combining pruning and quantization could yield superior performance-to-compression ratios compared to single-method approaches. Recognizing the challenges in accurately assessing LLM performance, we address key limitations of previous evaluation frameworks and introduce the Semantic Retention Compression Rate (SrCr), a novel metric that quantifies the trade-off between model compression and semantic preservation, facilitating the optimization of pruning-quantization configurations. Experiments demonstrate that our recommended combination achieves, on average, a 20% performance increase compared to an equivalent quantization-only model at the same theoretical compression rate. ## Read it - [arXiv:2505.07289](https://arxiv.org/abs/2505.07289): the preprint. The arXiv version includes an appendix with 6 result tables and has 10 pages, 15 figures and 7 tables. - [IEEE proceedings, via DOI 10.1109/IJCNN64981.2025.11227279](https://doi.org/10.1109/IJCNN64981.2025.11227279): the published version in the Proceedings of the 2025 International Joint Conference on Neural Networks (IJCNN), pages 1 to 9. - [PDF](https://arxiv.org/pdf/2505.07289): the arXiv PDF. ## Cite ```bibtex @inproceedings{laborde2025semantic, author = {Laborde, Stanislas and Cousseau, Martin and Yaacoub, Antoun and Prevost, Lionel}, title = {Semantic Retention and Extreme Compression in {LLMs}: Can We Have Both?}, booktitle = {2025 International Joint Conference on Neural Networks ({IJCNN})}, year = {2025}, pages = {1--9}, address = {Rome, Italy}, publisher = {IEEE}, doi = {10.1109/IJCNN64981.2025.11227279}, eprint = {2505.07289}, archivePrefix = {arXiv}, primaryClass = {cs.CL} } ``` --- # What Survives Compression Source: https://martincousseau.com/research/what-survives-compression Type: Note Published: 2025-05-12 Original: Semantic Retention and Extreme Compression in LLMs: Can We Have Both? — https://arxiv.org/abs/2505.07289 Topics: compression, judges Conclusion: Ask what meaning survived, not how fluent the remainder sounds. ## Summary Compression is a useful metaphor for what generative systems do to source. A long context becomes a short answer. A document becomes a schema. A policy becomes a refusal. The question that matters is not whether the remainder is elegant. It is what survived. At IJCNN 2025, my co-authors and I introduced SrCr — Semantic Retention Compression Rate — a measure of what remains after compression. The paper is about models. The habit it recommends is broader: score the retained meaning before you score the style of the remainder. That habit is also how I like to work: measurement before fluency, and the name on a report is the name on the work. ## What SrCr is for It is easy to compress a text and keep a sentence that sounds like the original. It is harder to keep the claim, the constraint, the exception. SrCr asks how much of the meaning survived the squeeze, not how pretty the squeeze looks. The same question applies when a RAG system answers from a pile of pages, or an extractor fills a schema from a scan. The remainder can be fluent and still be a different document. ## Measurement before fluency Fluency is cheap now. It is the wrong north star for a system that touches money, identity, or medical fact. The evaluation has to ask whether the claims are still the claims, whether the absences are still absences, whether the refusal is still the policy. A judge that rewards well-written error will pass a system that should have been held. Calibrate the judge against that failure, or do not use a judge. ## From the paper to a run In practice the analogue is simple. Name the meaning that must survive — the fields, the citations, the policy clauses. Measure retention on those, on the real distribution. Write the line. The rest of the report is how the number was made. [The paper](/research/srcr) remains the longer argument. This note is only the habit it left in my work. --- # Evaluation strategy Source: https://martincousseau.com/method/evaluation-strategy Type: Guide Topics: thresholds, compliance Conclusion: Start from the failure you fear, then write the line. ## What it is An evaluation strategy is the written system that turns opaque GenAI behaviour into a decision: ship, fix these, or defer. It starts from the failure modes and the decision you must own. The metrics a tool happens to expose come later. This is for AI-forward startups, small businesses with a first GenAI pipeline, and mid-size teams that need a written line for “good enough to ship”, without a multi-month governance programme. ## Six evaluation principles A named list, so a reader can cite a method rather than a paragraph. 1. **Start from the failure and the decision.** Not from the metrics a tool happens to expose. Name what must never ship, then write the line. 2. **Match the evaluator to the failure class.** Code for deterministic checks, a judge for qualitative known failures, an agent-judge for trajectories, a human for high-stakes nuance. 3. **Calibrate and version everything.** Judges, rubrics, datasets, and thresholds are software artefacts. A judge prompt change is a model change. 4. **Keep the offline set and the live sample distinct.** Both are required. Offline is the reference. The live sample is how you learn the reference is no longer the job. 5. **Make every score a decision.** Ship, fix, defer, or roll back. A metric that cannot issue one of those four is décor. 6. **Design for auditability.** Who authored the rubric, what data was used, how agreement was measured, what the release threshold was. If it cannot be shown, it is not evidence. ## Start from the failure, not the score Public leaderboards do not predict your distribution. A model that looks calm on a general benchmark can still invent a citation, skip a policy, or take the wrong tool on the one case that matters. Strategy work begins with the failures that would actually hurt: product, engineering, risk, and domain experts in the same room, naming what must never ship. From that list I write a task taxonomy and slices (happy path, edge, adversarial, demographic), so the later dataset and harness have somewhere to live. A score without a slice is a number you cannot act on. Ownership is named in the same pass. If no one owns the rubric when the model card changes, the strategy was a workshop, not a system. The memo says who may move a threshold, and who must be in the room when they do. ## Rubrics and the release gate The rubric is the judgment, written down. Pointwise or pairwise, pass/fail or graded, single-turn or trajectory: the form follows the failure. Each dimension has a definition, a method, and a threshold. [Faithfulness](/glossary#faithfulness) is not [groundedness](/glossary#groundedness). [Calibration](/glossary#calibration) is not confidence. Operational risk is not a style score. The release policy is the line: what is good enough to ship, what must be fixed first, what waits. Decision thresholds live here, in a go-live memo the organisation keeps after I leave. When the line is not written yet, writing it is the first piece of work: the line, the dimensions agreed, the first cases named. When the line exists, I score a slice against it, and then add the taxonomy and the mapping that let a harness enforce it. ## Mapped to regulation and the business Where the system is high-stakes, metrics map to internal policy, model-risk language, and, where it applies, obligations such as the EU AI Act. The mapping is explicit. A dashboard that cannot be read by legal is not a governance artifact. I do not issue certificates. I produce the evaluation design and the first evidence pack so your risk function has something other than a vendor narrative. What they then file is theirs. The pack names the system, the data it saw, who authored the rubric, and who may move a gate. That is evidence. A slide that says “AI Act ready” is not. ## What the strategy contains Typical artifacts. Scope follows the use case, not a template. | Artifact | What it decides | | ------------------------ | --------------------------------------------------- | | Failure-mode map | Which errors are in scope, and who owns them | | Task taxonomy and slices | What is tested, including the rare path | | Rubric specification | Dimensions, methods, and agreement rules | | Release policy | Ship, fix these, or defer, and when to roll back | | Metric-to-risk map | How a score becomes a sentence a committee can read | ## How I usually run it When “good enough” is not written down yet, I start by writing it: the line, the dimensions agreed, the first cases named. When the feature is already clear, I write the line at the start of the evaluation. --- # Golden datasets Source: https://martincousseau.com/method/golden-datasets Type: Guide Topics: datasets Conclusion: A small set with honest gold beats a large set nobody checked. ## What it is A golden dataset is not a handful of favourite prompts. It is a versioned, provenance-tracked sample of the distribution you will actually serve, including the cases you would rather not look at. This is for teams whose current eval set is a spreadsheet of happy paths, or who have some production data and need it turned into a set a score can rest on. ## From production, with a chain of custody I sample from logs or historical cases under the data rules you already have. Your data stays in the agreed environment. Anonymisation is part of the artifact, not an afterthought. Every item carries provenance: where it came from, who labeled it, which version of the rubric applied. The set is sliced. A single average hides the slice that fails. Challenge cases (rare, high-risk, adversarial) are constructed when production does not yet contain them, then validated by a human. Synthetic data is a supplement, never the whole set. A pipeline audit without a set is a reading of whatever happened to be in the room. The dataset is how that reading becomes repeatable when the model, the prompt, or the index moves. A download of MMLU or a public RAG bench is not a golden set for your product. Those numbers are useful as a smoke test. They do not know your documents, your users, or the one policy sentence that must never be invented. ## Labels you can defend Ground truth is a claim. I treat it as one: guidelines, more than one annotator where the stake requires it, agreement measured, disagreements adjudicated. The dataset card records that process so a later auditor is not asked to take the score on faith. Where the label is preference rather than truth (two fluent answers, one better), the set still carries the guideline version and the pair. Rankings without a protocol are taste. Domain experts label the critical slice. Crowd labels, if you already have them, are a starting point. They do not replace the cases that would fail a release. The card says which is which. ## Versioned like software When the product, the policy, or the model changes, the set changes with it, or it is frozen and named. Access control belongs with the rest of your evaluation assets. A golden set that cannot be reproduced is a demo. Retention follows your existing records policy. In regulated settings the evaluation set is itself an artifact that may need to be produced later. I write that down at the start, not after a request from audit. You keep the set. I leave a card, a schema, and a sampling rule your team can extend. A vendor-hosted eval dump you cannot export is not a ledger. ## What the set must carry Minimum fields. Further columns follow the rubric. | Field | Why it exists | | --------------------------- | --------------------------------------------- | | Provenance | Source system, time, sampling rule | | Slice | Happy path, edge, adversarial, demographic | | Label and guideline version | What “correct” meant on that date | | Agreement | Where humans disagreed, and how it was closed | | Access and retention | Who may see it, and when it is destroyed | ## How I usually run it I build the golden set as part of the evaluation itself, including the cases production under-samples. When almost no usable data exists yet, I start by building a first set of cases with sourced gold answers. --- # Evaluation harnesses Source: https://martincousseau.com/method/evaluation-harnesses Type: Guide Topics: judges, thresholds Conclusion: A harness is worth it when someone reruns it. ## What it is A harness makes evaluation a regression signal: the same dimensions, run before merge, nightly, and before release. Judges are calibrated against humans. The gate is a decision, not a dashboard. This is for teams who already generate outputs and need a harness they can rerun, in CI if they have it or as a scripted run if they do not, without marrying a single vendor. ## Tool-neutral, on purpose DeepEval, Braintrust, Arize Phoenix, LangSmith, RAGAS, Vertex Evaluate, custom Python: I select and configure what fits the failure class and your stack. Independence means no preferential relationship. The architecture should still stand if the tool changes. Deterministic checks belong in code. Qualitative known failures belong to an LLM-as-judge, itself monitored for drift. Trajectories belong to an agent-judge or a human. Matching the evaluator to the failure is the work; the library is an implementation detail. I will not pretend that a single framework covers faithfulness, policy, and tool use equally well. The harness is a composition. If a vendor later wraps the same composition, you still own the gates. If you do not have CI yet, the first harness is a scripted run you can execute again. A YAML file and a command that prints a slice report is a harness. A screenshot of a vendor dashboard is not. ## Calibrate the judge An uncalibrated judge is another fluent system. I design the judge prompt, measure it against human scores, and keep measuring it. When the judge drifts, the harness has failed even if the candidate model has not. [Calibration](/glossary#calibration) is versioned with the rest of the stack. A judge prompt change is a model change. It goes through the same gate it is meant to enforce, against a frozen human sample. Judge cost sits next to human cost. A cheap judge that disagrees with experts on the critical slice is not cheaper. I pick the mix so the gate you actually own is one you can afford to rerun. ## In the loop Pre-merge for cheap checks, nightly for the full golden set, pre-release for the gate that leadership actually owns. Dashboards and alerts exist so a miss is an incident, not a surprise in production. Cost and latency sit next to quality; a cheaper model that fails the rubric is not cheaper. The first working gate is usually a pilot: one use case, one golden set, one release policy. After that I make the gate ordinary: several use cases, CI that developers actually run, a scorecard a risk reader can open. The scorecard names the slice that failed, not only a global number. Product, engineering, and whoever owns the release read the same page. A metric that only the evaluation person understands will not hold a ship decision. ## Where the harness runs Cadence is a starting point. Risk sets the clock. | Gate | What it is for | | ------------- | ---------------------------------------------------------- | | Pre-merge | Deterministic checks and cheap judges on touched paths | | Nightly | Full golden set, slice report, judge health | | Pre-release | The written [threshold](/glossary#threshold): ship or hold | | Judge monitor | Agreement with humans; drift of the evaluator itself | ## How I usually run it I leave the harness as a protocol you can rerun on the next release. When it is handed over to your team, it comes as an `eval/` folder they run in CI. --- # Human evaluation Source: https://martincousseau.com/method/human-evaluation Type: Guide Topics: human review, judges Conclusion: People read the rows a metric can't. ## What it is Where automation cannot see the failure, a human must: trained, calibrated, and measured. Unstructured eyeballing is not a programme. Agreement is the score of the scoring. This is for dimensions where automation is not enough: preference, nuance, or the slice an LLM-judge keeps missing. I design the programme as part of the evaluation; your subject-matter experts usually rate. ## Guidelines before raters A rater without a guideline is improvising. I write the guide from the rubric, train raters on it, and keep a calibration set that is scored again as people drift. New raters do not join a live queue until they match the standard. Guidelines are short enough to use under time pressure and specific enough to survive a disagreement. If two trained people cannot apply a sentence the same way, the sentence is wrong. The guide is an artifact your experts keep. It names examples of pass, hold, and fail for each dimension, with no paragraph of theory. If a new rater cannot use it on day one, it is not finished. ## Agreement is part of the ledger Inter-annotator agreement is tracked. Disagreements are adjudicated, not averaged away. Preference ranking and pairwise comparison are used when the question is “which is better”, not “is this true”. The method follows the decision. Adjudication is logged. A later reader should see why the gold label is the gold label. Hidden consensus is how human evaluation becomes theatre. Agreement is not a vanity statistic. If two trained people split on a policy item, the release gate cannot pretend the score is settled. That item is a hold until the guide or the model changes. ## Hybrid by default Humans are expensive. The design is usually hybrid: automatic scores on the bulk, human review on a sample, on disagreements, and on the critical slice. Cost is an evaluation parameter, not an embarrassment. The sample is not random courtesy. It is stratified by slice and by the auto-score’s uncertainty. That is how a small panel still sees the cases that would fail a gate. Your people usually rate. I design the programme, the calibration loop, and the report. I am not a labelling farm, and I do not take a cut from a crowd vendor. The capability is meant to stay with you. The output is a programme you can run again: the guide, the calibration set, the sampling rule, and an agreement report a later auditor can read. A week of unstructured comments in a spreadsheet is not that. ## What the programme holds These artifacts are what make a human score repeatable. | Artifact | Purpose | | ---------------- | --------------------------------------------------- | | Rater guidelines | The same definition of the dimension, in writing | | Calibration set | Whether people still agree with last month | | Agreement report | Agreement, adjudication log, rater notes | | Sampling rule | What the humans see, and what the auto-score covers | ## How I usually run it I design the programme as part of the evaluation, then you keep the cadence. Before I report any result, I read the rows the average hides. --- # Agent evaluation Source: https://martincousseau.com/method/agent-evaluation Type: Guide Topics: agents, judges Conclusion: Score the path, not only the final answer. ## What it is Agents fail in the process: the wrong tool, the wrong argument, no recovery, a bad hand-off. A final-answer score will bless a trajectory that should never have been allowed to finish. This is for teams whose agent experiments have matured enough to score tool calls and recovery. I start with the core feature, unless the agent is already the job. ## Score the path I evaluate tool selection, argument correctness, use of retrieved context, and what happens after an error. Task completion is necessary and not sufficient. Intermediate state, whether the agent knew it was lost, is part of the rubric. Traces are the unit of work. A last-message corpus will not show a forbidden call that was later overwritten by a polite summary. If you cannot replay the path, you cannot score it. A RAG [faithfulness](/glossary#faithfulness) score on the final sentence can still bless a trajectory that queried the wrong system, looped on an empty retrieval, or skipped the human hand-off. Agents fail in the process. That is why this guide exists. ## Judges that can see a trace Agent-as-judge designs are used when a single-turn LLM-judge cannot see the process. They are calibrated like any other judge. Humans remain on the trajectories that would be indefensible to wave through. A process judge can still be fluent and wrong. I hold it to a human sample of traces, including recoveries and refusals, not only the successes a demo would pick. Replay is the offline set: held-out tasks, frozen tools, a trace you can run again. Live sampling of production trajectories is continuous evaluation. Do not mix the two numbers on one chart without saying so. ## Coordination is a failure surface Multi-agent systems add hand-offs. I score whether the next agent received what it needed, whether ownership was clear, and whether the system stopped when it should have asked a person. Escalation is a dimension. An agent that “finishes” by guessing instead of handing to a human has failed the [threshold](/glossary#threshold) even if the answer happens to be right. When the agent is the job, the work still produces a rubric, a replay set, and a gate. A demo of tool calling is not a gate. ## Dimensions on a trajectory Not every agent needs every row. The failure you name chooses. | Dimension | What is scored | | -------------- | ----------------------------------------------- | | Tool selection | The right tool, at the right step | | Arguments | Parameters match the schema and the intent | | Recovery | Behaviour after a tool error or empty retrieval | | Completion | The task actually finished, on the evidence | | Hand-off | State and ownership across agents or humans | ## How I usually run it I focus on one agent family, once the core feature has already been evaluated. I test through a guest seat or a scoped staging key, never a private key. --- # Continuous evaluation Source: https://martincousseau.com/method/continuous-evaluation Type: Guide Topics: drift, thresholds Conclusion: Re-check when the model, prompt or data moves. ## What it is Offline evaluation is necessary and not sufficient. Models, prompts, retrieval, and traffic move. Continuous evaluation is the same report, run again. This is for teams that already have a first golden set and a gate. It is rarely the first run. A rerun of the existing evaluation on the same system is how it usually starts. ## Sample the live system I design sampling that is cheap enough to run and representative enough to trust, including the slices that matter. Online scoring reuses the same dimensions as the offline set. A second, prettier dashboard is not continuity. Production is a different distribution. The golden set stays the reference; the live sample is how you learn that the reference is no longer the job. Both stay in the scorecard. Sampling is a data decision. Rate, retention, and who may see live outputs follow your existing rules, never a vendor’s default. A live sample that cannot be destroyed on request is not a sample you should have taken. ## Drift is a decision Quality drift, behavioural drift, and judge drift are distinguished. An alert without an owner is décor. Signals connect to incident response: investigate, hold, roll back the prompt, the model, or the retrieval. A quiet move in one slice (one language, one product line, one tool) is the usual incident. Continuity that only watches the global average will miss it. Observability tells you latency, cost, and whether a trace existed. Continuous evaluation scores the output against the same rubric you used to ship. You need both. One dashboard that mixes them will hide the miss. ## Cost, latency, and judgment together A model that is cheaper and slower to fail is not an improvement. Joint monitoring keeps the trade-off visible, so a cost optimisation cannot silently spend the [threshold](/glossary#threshold). Reruns are how this stays alive: new use cases, new models, the same dimensions. The first written result is a beginning. Drift is the rest of the work. This is not a monitoring product. I stand the first evaluation up first. Then I scope the cadence as reruns: the same dimensions, run again, with an owner for the alert. A weekly slice report with a named owner beats a real-time wall nobody reads. Cadence follows risk. The point is that last month’s threshold still means something this month. ## What stays alive The living scorecard. Same dimensions as the offline rubric. | Signal | What it is for | | ----------------- | ------------------------------------------- | | Production sample | The distribution you actually serve | | Slice report | Where quality moved, not only that it moved | | Drift alert | A change large enough to own | | Rollback path | Prompt, model, index: who may revert it | ## How I usually run it I rerun the evaluation on the same system after the first result: a new slice, or a check after fixes. When the work is handed over, an `eval/` folder lets your team run it between my reruns. --- # Risk & compliance Source: https://martincousseau.com/method/risk-compliance Type: Guide Topics: compliance Conclusion: Evidence your risk team can read, not a badge. ## What it is Some failures are not quality in the ordinary sense. They are policy, leakage, or a sentence you cannot defend. This is for teams whose first evaluation already exists and who now need a scored red team or a policy file. ## Red-teaming with a catalog Adversarial tests are designed from the failure-mode map: jailbreaks, leakage, disallowed advice, prompt injection, retrieval of the wrong document. Findings are a catalog with patches, not a slide of scary examples. A focused red team is this work in concentrated form. I re-test after the patch. A red team that does not return is a performance. The catalog is versioned with the system it was run against. I score policy, leakage, and the cases that would fail a release, against the same rubric the rest of the evaluation uses. A PDF of jailbreaks without a [threshold](/glossary#threshold) is a scare file. ## Policy, fairness, safety Instruction adherence is scored against system, policy, and format constraints. Bias and demographic performance are sliced, not averaged. Content safety is a dimension with a threshold, including “zero critical” where that is the only acceptable line. Operational risk (policy, leakage, escalation) is this family of failures. A high [faithfulness](/glossary#faithfulness) score does not excuse a single critical leak. Fairness work is sliced evidence, not a slogan. I report the slice that fails and the sample it rests on. I do not publish a single “bias score” that cannot be acted on. ## Evidence that can leave the room Audit-ready packages (eval factsheets, or the format your model-risk function already uses) record who authored the rubric, what data was used, how agreement was measured, and what the release threshold was. If it cannot be shown, it is not evidence. I write in the register the committee already reads. I do not invent a parallel vocabulary that only the vendor understands. This is not a legal opinion and not an EU AI Act certificate. Counsel decides what the file means. The point of the work is that they have a scored catalog instead of a vendor narrative. The file is yours. I do not keep a parallel copy of findings, and I do not sell a compliance badge. If a committee asks “who wrote the rubric, and on which data?”, the factsheet answers in a sentence. ## What the file contains For internal risk, legal, and external review when required. | Artifact | Reader | | ------------------------- | ---------------------------- | | Failure catalog | Engineering and security | | Policy adherence scores | Product and compliance | | Fairness and slice report | Risk and the business owner | | Eval factsheet | Model risk, audit, regulator | ## How I usually run it I scope the work to policy, leakage, and escalation, once a first evaluation exists. My access is scoped and time-boxed. --- # Glossary Source: https://martincousseau.com/glossary Updated: 2026-09-26 Five terms I use in every evaluation report: faithfulness, groundedness, calibration, threshold, ledger. Each is defined once here so it can be cited. - Faithfulness: Faithfulness is whether a claim is supported by the context that was actually retrieved, and not by the model’s prior. (https://martincousseau.com/glossary#faithfulness) - Groundedness: Groundedness is the absence of unsourced assertion. Fluent and unsupported is still a miss. (https://martincousseau.com/glossary#groundedness) - Calibration: Calibration is whether stated confidence matches observed accuracy, of the model or of the judge. (https://martincousseau.com/glossary#calibration) - Threshold: The threshold is the written line: above it, ship; below it, fix or defer. A score without that line is a dashboard. (https://martincousseau.com/glossary#threshold) - Ledger: A ledger is the table in my evaluation reports: the dimensions, the scores, the line, and who may move each gate. (https://martincousseau.com/glossary#ledger) --- # About Source: https://martincousseau.com/about Updated: 2026-10-05 ## In brief I'm Martin Cousseau, a French AI Engineer based in Warsaw. I build LLM systems and the evaluations that tell you whether they are ready to ship, and I take on B2B projects that need both. - **What I do:** build LLM systems, and design the evaluations that show whether they are ready to ship. - **Based in:** Warsaw, Poland. - **Languages:** English and French. - **Research:** second author of a paper on semantic retention under LLM compression, published at IJCNN 2025. - **Public work:** two open evaluation harnesses, one for document extraction and one for a helpdesk agent, with their gold datasets. Details below. - **Contact:** ## Background An engineering degree in Paris, LLM evaluation at work since 2025, and public evaluation work on the side. ### Experience - **GenAI Evaluation, Hitachi Rail** (full-time, September 2025 to now). I developed an evaluation harness for LLM-based features, created dedicated metrics and deterministic scores, built LLM-as-a-judge pipelines and automated the reports for stakeholders. - **Junior Software Engineer, GlobalLogic** (full-time, remote from Warsaw, September 2025 to now). I build AI pipelines and evaluation frameworks for LLM-based apps. - **Full Stack & Generative AI intern, Hitachi Rail** (February to August 2025, Vélizy-Villacoublay, France, hybrid). I automated complex compliance workflows with RAG and LLMs. - **Full Stack & Generative AI intern, Wemersive** (April to December 2024, remote, Los Angeles). - **Research intern, Talan** (June to July 2023, Paris, on site). A technical internship at Talan's research and innovation centre, focused on LLMs and their applications. - **Technical support intern, Atos** (July 2021, Boulogne-Billancourt, France). ### Education - **ESIEA, Paris:** Master of Science in Computer Science (September 2020 to August 2025). In my third year I led a scientific and technical project at the school's research lab, the Learning, Data and Robotics (LDR) lab. In my fourth year I was project leader for a collaboration between the LDR lab and a partner. ESIEA funded my research there and gave access to resources, GPUs in particular. - **Centria University of Applied Sciences, Finland:** a semester of computer science (September to December 2022). - **Dorset College Dublin:** a semester of mathematics and computer science (September to December 2021). ### Public record What I can date from my own repositories, papers and videos: - **2023:** public notebooks on 4-bit quantized inference for Falcon 7B and on embedding models. - **2024:** a public study of pruning and quantization on Llama-2-7b (repository created in March). - **2025:** the IJCNN paper below, on arXiv in May and published at IJCNN in Rome (30 June to 5 July). - **2026:** Extraction Arena (June) and Refund Arena (September), evaluation harnesses with public gold datasets. ## Research I am the second author of “Semantic Retention and Extreme Compression in LLMs: Can We Have Both?”, published at IJCNN 2025. ### How the paper came about After a research internship at Talan's research and innovation centre in Paris (June to July 2023, on LLMs and their applications), I started an AI project with one of their AI researchers while I was doing a one-year master's. I then brought the project to my school's research lab, the Learning, Data and Robotics (LDR) ESIEA Lab in Paris, and invited Stanislas Laborde to join. We started the serious work at the end of 2024. Stanislas and I contributed equally. He is first author by agreement because he is doing a PhD. My school funded my trip to Rome for the conference. The authors, in published order, are Stanislas Laborde, Martin Cousseau, Antoun Yaacoub and Lionel Prevost, all at ESIEA's LDR lab in Paris. The paper appeared at the 2025 International Joint Conference on Neural Networks (IJCNN) in Rome and is published by IEEE. **SrCr in one line:** the Semantic Retention Compression Rate is a metric that quantifies the trade-off between how far a model is compressed and how much of its meaning it keeps. The paper reports that its recommended combination of compression techniques gives, on average, a 20% performance increase over an equivalent quantization-only model at the same theoretical compression rate. Experiments ran on NVIDIA A10 GPUs and were scored with lm-evaluation-harness on benchmarks that include MMLU-Pro and BBH. [Read it on arXiv](https://arxiv.org/abs/2505.07289) · [IEEE record (DOI 10.1109/IJCNN64981.2025.11227279)](https://doi.org/10.1109/IJCNN64981.2025.11227279) · [Paper page on this site](/research/srcr) ## How I work Four habits that shape how I evaluate a system. 1. **I name the failure before the metric.** I ask which failure would stop the system from shipping. It gets its own scorer and its own line in the results, next to the average. 2. **Deterministic checks first, judges named.** Code checks run before any model is asked for an opinion. When a model judges, I say which one. 3. **I read the rows the average hides.** The verdict comes from the person who read the rows, not from the average alone. 4. **I publish what I learn.** Conclusions up front, method inside. The [research](/research) and [method](/method) pages are written that way. ## Open source & writing Everything here is public, linked and dated. ### Code - [Extraction Arena](https://github.com/martin-cousseau/extraction-arena) (June 2026): a DocAI evaluation harness. It scores document-extraction pipelines, such as LlamaParse and LlamaExtract and vision models, against a per-document golden dataset. TypeScript, MIT. - [Refund Arena](https://github.com/martin-cousseau/refund-arena) (September 2026): an eval plate for a helpdesk agent, with one shop, one policy and six orders. The write is the score. Python, Apache-2.0. ### Datasets on Hugging Face - [Cybertruck-Rescue-Sheet](https://huggingface.co/datasets/martincousseau/Cybertruck-Rescue-Sheet) (June 2026): a one-document gold extraction of the public four-page Cybertruck first-responder rescue sheet, ISO 17840-style and not certified. MIT. - [refund-arena](https://huggingface.co/datasets/martincousseau/refund-arena) (September 2026): the gold for the helpdesk agent above. Apache-2.0. - [mermaid-persona-queries](https://huggingface.co/datasets/martincousseau/mermaid-persona-queries) (September 2026): 50 hand-written user queries, across 5 personas, for evaluating text-to-Mermaid diagram generation. CC-BY-4.0. ### Video and writing - 21 August 2026: [Extraction Arena: Evaluate Vision LLMs for Document Extraction](https://www.youtube.com/watch?v=QXWN8WyvPmI), a 9 min 43 s walkthrough on YouTube. - 7 September 2026: [Cybertruck: Extraction Failure with LlamaParse](https://www.youtube.com/watch?v=LT4fSzy886s), a YouTube Short. - 9 September 2026: [Can we trust a model with a rescue procedure?](https://medium.com/@martin-cousseau/can-we-trust-a-model-with-a-rescue-procedure-ddf5184ef608), an article on Medium.