# Martin Cousseau > Martin Cousseau is an AI Engineer in Warsaw, Poland: "I build LLM systems and the evaluations that tell you whether they are ready to ship." Contact: contact@martincousseau.com. Personal site of Martin Cousseau. French, based in Warsaw, Poland; works in English and French. Pages are written by Martin Cousseau unless a byline says otherwise (the paper lists all of its authors in published order). Other people share the name Martin Cousseau; the profiles listed below identify this one. ## Research - [Extraction Arena](https://martincousseau.com/research/extraction-arena): Three vision models and four LlamaParse modes read the same rescue sheet. I froze their answers and moved only the scoring rule to see what the headline hid. - [Extraction Arena: Evaluate Vision LLMs for Document Extraction](https://martincousseau.com/research/extraction-arena-walkthrough): A recorded walkthrough of Extraction Arena, the open-source harness I use to evaluate vision LLMs on document extraction, field by field. - [Semantic Retention and Extreme Compression in LLMs: Can We Have Both?](https://martincousseau.com/research/srcr): An IJCNN 2025 paper I co-authored on semantic retention under extreme LLM compression. It introduces SrCr, a measure of how much meaning survives. - [What Survives Compression](https://martincousseau.com/research/what-survives-compression): A note on SrCr, semantic retention under compression, and why an evaluation should ask what meaning survived before it asks how fluent the remainder sounds. ## Method - [Evaluation strategy](https://martincousseau.com/method/evaluation-strategy): Evaluation strategy for GenAI systems: failure modes, dedicated metrics, rubrics, and a written line for what is good enough to ship. - [Golden datasets](https://martincousseau.com/method/golden-datasets): Golden datasets and challenge sets for LLM, RAG, and agent evaluation: sampled from real use, versioned, and labeled so a score can be defended. - [Evaluation harnesses](https://martincousseau.com/method/evaluation-harnesses): Tool-neutral LLM evaluation harnesses: dedicated metrics, calibrated judges, and a pipeline you can rerun, so quality becomes a signal you can act on. - [Human evaluation](https://martincousseau.com/method/human-evaluation): Human and hybrid LLM evaluation: rater guidelines, calibration, and measured agreement, used where automation is not enough. - [Agent evaluation](https://martincousseau.com/method/agent-evaluation): Evaluation of multi-step agents: tool selection, arguments, recovery, and hand-offs. The failures a single-turn score misses. - [Continuous evaluation](https://martincousseau.com/method/continuous-evaluation): Continuous evaluation of production GenAI: sample the live system, watch quality drift, and rerun the same dimensions when the model, prompt, or data moves. - [Risk & compliance](https://martincousseau.com/method/risk-compliance): Risk and safety evaluation for GenAI: red-teaming, policy adherence, sliced fairness, and evidence a risk team can keep. ## Reference - [Glossary](https://martincousseau.com/glossary): My definitions of faithfulness, groundedness, calibration, threshold and ledger. Five terms I use in every evaluation report, written so they can be cited. - [About](https://martincousseau.com/about): French AI Engineer in Warsaw. I build LLM systems and the evaluations that show whether they are ready to ship, and I co-authored an IJCNN 2025 paper on semantic retention under LLM compression. ## Profiles - [LinkedIn](https://www.linkedin.com/in/cousseaumartin/): cousseaumartin - [GitHub](https://github.com/martin-cousseau): martin-cousseau - [Hugging Face](https://huggingface.co/martincousseau): martincousseau - [Google Scholar](https://scholar.google.com/citations?user=KzMJqUcAAAAJ): Martin Cousseau - [ORCID](https://orcid.org/0009-0000-5228-3565): 0009-0000-5228-3565 - [arXiv](https://arxiv.org/abs/2505.07289): 2505.07289 - [YouTube](https://www.youtube.com/@martin-cousseau): @martin-cousseau - [Medium](https://medium.com/@martin-cousseau): @martin-cousseau - [X](https://x.com/m_cousseau): @m_cousseau - [TikTok](https://www.tiktok.com/@martin.cousseau): @martin.cousseau ## Optional - [Full text](https://martincousseau.com/llms-full.txt): every page above as plain markdown in one file - [RSS feed](https://martincousseau.com/rss.xml): new research notes and method guides - [Sitemap](https://martincousseau.com/sitemap.xml): every URL with its last-modified date