martincousseau.com

About

// who writes here

I build LLM systems and the evaluations that tell you whether they are ready to ship.

Martin Cousseau at his desk, in front of three monitors showing code

01

#

In brief

I'm Martin Cousseau, a French AI Engineer based in Warsaw. I build LLM systems and the evaluations that tell you whether they are ready to ship, and I take on B2B projects that need both.

  • What I do: build LLM systems, and design the evaluations that show whether they are ready to ship.
  • Based in: Warsaw, Poland.
  • Languages: English and French.
  • Research: second author of a paper on semantic retention under LLM compression, published at IJCNN 2025.
  • Public work: two open evaluation harnesses, one for document extraction and one for a helpdesk agent, with their gold datasets. Details below.
  • Contact: contact@martincousseau.com

02

#

Background

An engineering degree in Paris, LLM evaluation at work since 2025, and public evaluation work on the side.

Experience

  • GenAI Evaluation, Hitachi Rail (full-time, September 2025 to now). I developed an evaluation harness for LLM-based features, created dedicated metrics and deterministic scores, built LLM-as-a-judge pipelines and automated the reports for stakeholders.
  • Junior Software Engineer, GlobalLogic (full-time, remote from Warsaw, September 2025 to now). I build AI pipelines and evaluation frameworks for LLM-based apps.
  • Full Stack & Generative AI intern, Hitachi Rail (February to August 2025, Vélizy-Villacoublay, France, hybrid). I automated complex compliance workflows with RAG and LLMs.
  • Full Stack & Generative AI intern, Wemersive (April to December 2024, remote, Los Angeles).
  • Research intern, Talan (June to July 2023, Paris, on site). A technical internship at Talan's research and innovation centre, focused on LLMs and their applications.
  • Technical support intern, Atos (July 2021, Boulogne-Billancourt, France).

Education

  • ESIEA, Paris: Master of Science in Computer Science (September 2020 to August 2025). In my third year I led a scientific and technical project at the school's research lab, the Learning, Data and Robotics (LDR) lab. In my fourth year I was project leader for a collaboration between the LDR lab and a partner. ESIEA funded my research there and gave access to resources, GPUs in particular.
  • Centria University of Applied Sciences, Finland: a semester of computer science (September to December 2022).
  • Dorset College Dublin: a semester of mathematics and computer science (September to December 2021).

Public record

What I can date from my own repositories, papers and videos:

  • 2023: public notebooks on 4-bit quantized inference for Falcon 7B and on embedding models.
  • 2024: a public study of pruning and quantization on Llama-2-7b (repository created in March).
  • 2025: the IJCNN paper below, on arXiv in May and published at IJCNN in Rome (30 June to 5 July).
  • 2026: Extraction Arena (June) and Refund Arena (September), evaluation harnesses with public gold datasets.

03

#

Research

I am the second author of “Semantic Retention and Extreme Compression in LLMs: Can We Have Both?”, published at IJCNN 2025.

How the paper came about

After a research internship at Talan's research and innovation centre in Paris (June to July 2023, on LLMs and their applications), I started an AI project with one of their AI researchers while I was doing a one-year master's. I then brought the project to my school's research lab, the Learning, Data and Robotics (LDR) ESIEA Lab in Paris, and invited Stanislas Laborde to join. We started the serious work at the end of 2024. Stanislas and I contributed equally. He is first author by agreement because he is doing a PhD. My school funded my trip to Rome for the conference.

The authors, in published order, are Stanislas Laborde, Martin Cousseau, Antoun Yaacoub and Lionel Prevost, all at ESIEA's LDR lab in Paris. The paper appeared at the 2025 International Joint Conference on Neural Networks (IJCNN) in Rome and is published by IEEE.

SrCr in one line: the Semantic Retention Compression Rate is a metric that quantifies the trade-off between how far a model is compressed and how much of its meaning it keeps.

The paper reports that its recommended combination of compression techniques gives, on average, a 20% performance increase over an equivalent quantization-only model at the same theoretical compression rate. Experiments ran on NVIDIA A10 GPUs and were scored with lm-evaluation-harness on benchmarks that include MMLU-Pro and BBH.

Read it on arXiv · IEEE record (DOI 10.1109/IJCNN64981.2025.11227279) · Paper page on this site

04

#

How I work

Four habits that shape how I evaluate a system.

  1. I name the failure before the metric. I ask which failure would stop the system from shipping. It gets its own scorer and its own line in the results, next to the average.
  2. Deterministic checks first, judges named. Code checks run before any model is asked for an opinion. When a model judges, I say which one.
  3. I read the rows the average hides. The verdict comes from the person who read the rows, not from the average alone.
  4. I publish what I learn. Conclusions up front, method inside. The research and method pages are written that way.

05

#

Open source & writing

Everything here is public, linked and dated.

Code

  • Extraction Arena (June 2026): a DocAI evaluation harness. It scores document-extraction pipelines, such as LlamaParse and LlamaExtract and vision models, against a per-document golden dataset. TypeScript, MIT.
  • Refund Arena (September 2026): an eval plate for a helpdesk agent, with one shop, one policy and six orders. The write is the score. Python, Apache-2.0.

Datasets on Hugging Face

  • Cybertruck-Rescue-Sheet (June 2026): a one-document gold extraction of the public four-page Cybertruck first-responder rescue sheet, ISO 17840-style and not certified. MIT.
  • refund-arena (September 2026): the gold for the helpdesk agent above. Apache-2.0.
  • mermaid-persona-queries (September 2026): 50 hand-written user queries, across 5 personas, for evaluating text-to-Mermaid diagram generation. CC-BY-4.0.

Video and writing

Contact

// one line, one action

Tell me what you’re about to ship.

contact@martincousseau.com

$ mail contact@martincousseau.com

Warsaw · English / French · open to B2B engagements