martincousseau.com

Note, , 6 min read, Updated

Extraction Arena

Can we trust a model with a rescue procedure?

IN SHORT

A strong average can hide a procedure the model gets entirely wrong. Gate the field you fear, not the mean.

One-bit dithered grid of form fields, the cover image for the Extraction Arena note.
PLATE 1. Generated plate · 1-bit · form-grid

01 · In brief

#

Summary

Three vision models read Tesla’s Cybertruck first-responder documentation and filled a structured record. The headline scores look ready to ship (93, 90, 90). The water-rescue procedure does not.

I recorded a walkthrough of the harness behind this note, and the code is open source: martin-cousseau/extraction-arena.

LedgerFIG. 01

Composed scores: GLM-5V-Turbo 93, GPT-5.4 Mini 90, Grok 4.5 90. Submersion exact sequence is 0. LlamaParse in mode order: turbo 83, cost-effective 79, agentic 84, agentic plus 87.

Composed score · three vision pipelines

  • 93GLM-5V-TurboZ.AI
  • 90GPT-5.4 MiniOpenAI
  • 90Grok 4.5xAI
Submersion · exact · sequenceThe procedure a crew would follow0, see the veto test

LlamaParse

FIG. 01. Headline scores look ready to ship. The water-rescue procedure does not, on vision or on LlamaParse.

I keep the model answers frozen and change only the scoring rule, so I can see whether I have a model problem or a definition problem. Later I add a fourth pipeline, LlamaParse through LlamaExtract, and read that run the same way.

The business question does not move. Would you accept a high extraction score if the one procedure a crew would follow still scores zero?

02

#

A go / no-go decision, not just a leaderboard

For product owners deciding go / no-go, subject-matter experts deciding whether order and wording are allowed to move on a rescue step, or engineers who have to turn that decision into weights. If you ship document extraction into operations, this is the meeting you skip at your cost. The average will look fine. The field a crew would actually follow might not.

03

#

This is not a clean text document

Tesla publishes official rescue information for the Cybertruck. A rescue sheet (ISO 17840-1) is the quick page at the scene. An Emergency Response Guide (ISO 17840-3) is the longer procedure set used to train and operate. The four pages the models saw mix diagrams, pictograms, numbered procedures, and prohibitions in capital letters. A wrong cut zone or a skipped disable step is an operational failure, not just a formatting issue. I selected and scored the following fields: cut zones, high-voltage disable, fire, submersion, and tow.

The first experiment ran in four steps.

  1. Each page is rendered as an image (the models see what a person sees on paper).
  2. Every model gets the same empty form; the correct answers are never in the prompt.
  3. GLM-5V-Turbo, GPT-5.4 Mini, and Grok 4.5 each return their own JSON; I map those answers onto one shared form.
  4. Then I score field by field, because the business risk is not the same on manufacturer, on a prohibition, and on a five-step procedure.
Rescue sheetFIG. 02

Four pages of the public Cybertruck rescue sheet, drawn as text. Layout scores cut zones. Disable scores high-voltage disable. Occupants marks occupant access as credit. Fire, water, and tow score fire and submersion as gates, and tow.

Tesla Cybertruck rescue sheet · four pages

  1. 01Layout · airbags · HV

    • Diagram
    • Legend
    Cut zones
    Scored

    The models saw this page as a picture of the vehicle, the airbags, and the high-voltage zones.

  2. 02Identify · immobilize · disable

    • Numbered procedure
    • Prohibition
    High-voltage disable
    Scored

    Identification, immobilization, and the disable sequence. A skipped disable step is an operational failure.

  3. 03Occupants · stored energy

    • Diagram
    • Numbered procedure
    Occupant access
    credit

    How a crew reaches the occupants, and where the stored energy sits. Occupant access is credit, not a veto.

  4. 04Fire · water · tow

    • Pictogram
    • Prohibition
    • Numbered procedure
    Fire
    gate
    Submersion
    gate
    Tow
    Scored

    Pictograms, prohibitions in capital letters, then the fire, submersion, and tow procedures.

FIG. 02. ISO 17840. Diagrams, pictograms, numbered procedures, prohibitions in capital letters. Tesla wording stays Tesla's. Not a client result.

04

#

I already have the answers. I am changing the question.

I am not asking the models again. I am changing how a correct answer is defined for one field. I zoomed in on the five-step procedure for a submerged Cybertruck. Product owners pick the scoring rule that matches the risk. Subject-matter experts say whether order and wording are allowed to move.

LexiconFIG. 03

Six lenses on one field: precision, recall, F1, sequence, set, and exact versus partial.

What the scores actually mean

Six lenses. One field.

01Precision
Of everything the model wrote, how much was actually on the sheet. Low precision means it invented steps. Right lens for a prohibition.
Precision is the lens for a prohibition. Invented steps fail here.
02Recall
Of everything the sheet required, how much the model captured. Low recall means it skipped steps. Right lens for a procedure.
Recall is the lens for a procedure. Skipped steps fail here.
03F1
The balance of the two. If the model invents half the procedure or skips half, this number falls. Not a popularity score.
F1 holds both costs. Inventing half or skipping half pulls it down.
04Sequence
Order is part of the answer. Step four sitting in position one is a miss. You do not start a water rescue by lifting the front.
One model opened at gold step 4. Under sequence that is a miss in position 1.
05Set
Order does not count. A list of warnings is correct if the items are present, even shuffled. Not for a procedure a crew must perform in order.
Under set, the step is merely present, and a crew could skip PPE.
06Exact / partial
Exact: the shipped value must be that wording. Partial: overlapping words still score. Gold step 1 is the PPE line.
Exact match fails "Wear PPE" against the gold "Wear appropriate PPE for water rescue." Partial match still scores the overlapping words.
FIG. 03. Precision is invented steps. Recall is skipped steps. Sequence is order. Exact is the wording a crew would act on.

How close must the wording be? Exact or partial. If too strict, a useful extraction looks like a failure. If too loose, a dangerous paraphrase looks like a pass. Example: exact match fails “Wear PPE” against the gold “Wear appropriate PPE for water rescue.” Partial match still scores the overlapping words.

Does order count? Sequence or set. A rescue procedure is a sequence of actions. Sequence compares step 1 to gold step 1. Set treats the list as a bag. Example: one model opened at gold step 4. Under sequence that is a miss in position 1. Under set, the step is merely present, and a crew could skip PPE.

The question, which mistake to surface first, is a product decision, not a scoring one.

Precision: of everything the model wrote, how much was on the sheet — the lens for a prohibition.

Recall: of everything the sheet required, how much the model captured — the lens for a procedure.

And F1 holds both costs.

05

#

Three stories. None of them improved the models.

None of the three vision models returned the five-step order.

In the illustration below a filled cell is a step the model actually wrote, a fused cell is two gold steps crushed into one line, an empty cell is a skipped step.

  • One model opened at gold step 4.
  • The other two started on PPE, fused remove-from-water with the high-voltage disable, then stopped.

All three saw the submersion paragraph. None produced a five-step ordered list.

AlignmentFIG. 04

Gold is five ordered steps: PPE, remove from water, continue HV disable, raise the front about 30 cm, store flat. One vision model opened at step 4. The other two fused remove with disable. Three LlamaParse modes left the list empty. Agentic plus wrote three lines, not five.

Ground truth. The five ordered steps on the gold record

  • Filled: a step the model wrote.
  • Fused: two gold steps crushed into one line.
  • Empty: a skipped step.

Gold record

  1. 01PPE for water rescue
  2. 02Remove from water
  3. 03Continue HV disable
  4. 04Raise front ~30 cm
  5. 05Store flat
  1. 01PPE for water rescue
  2. 02Remove from water
  3. 03Continue HV disable
  4. 04Raise front ~30 cm
  5. 05Store flat

One vision model

  1. PPEEmpty
  2. RemoveEmpty
  3. DisableEmpty
  4. RaiseFilled
  5. StoreEmpty

Opened at gold step 4, "raise the front." PPE and extraction never happen.

The other two

  1. PPEFilled
  2. Remove + disableFused
  3. RaiseEmpty
  4. StoreEmpty

Started on PPE, fused remove-from-water with the high-voltage disable, then stopped.

Turbo · cost-effective · agentic

  1. PPEEmpty
  2. RemoveEmpty
  3. DisableEmpty
  4. RaiseEmpty
  5. StoreEmpty

LlamaParse. Ordered steps empty on three modes. Precision 1, recall 0. Omission, not invention.

LlamaParse · agentic plus

  1. PPE · slot 2Filled
  2. Remove + disableFused
  3. Raise ‡Empty
  4. Store ‡Empty

Wrote three ordered steps. Opened on a platitude. Fused remove with disable. 30 cm still in another field.

FIG. 04. ‡ Captured as drainage_lift_requirement. A full partial on its own field, still not step 4 or 5.

I then read those frozen answers under three scenarios. Only the scoring rule changes.

  • Exact wording, order required: precision 0, recall 0, F1 0.
  • Flexible wording, order still required: 44, 25, 33.
  • Flexible wording, order ignored: 81, 47, 58.

All of these results are true. They just answer different questions.

I did not improve the models, but I changed the question. For this field a kinder pairing rule increases the score and rewards an extracted procedure no crew should follow.

Scoring rulesFIG. 05

Mean of three vision models on submersion ordered-steps. Exact sequence: precision 0, recall 0, F1 0. Partial sequence: 44, 25, 33. Partial set: 81, 47, 58.

Submersion ordered steps · mean of three models

  • Precision %
  • Recall %
  • F1 %
Submersion ordered-steps, mean of three vision models, under three scoring rules
Scoring rulePrecisionRecallF1
Exact · sequenceThe question you would hand to a crew.000
Partial · sequenceFlexible wording, order still required.442533
Partial · setFlexible wording, order ignored. A procedure no crew should follow.814758
FIG. 05. Exact · sequence is the question you would hand to a crew.

06

#

The mix still looks shippable.

A well-designed metric should let you trust the headline. That is the goal. A composed score only earns that trust if its weights match the risk of the work.

But in this demo they do not: manufacturer, vehicle type and other easy fields pull the extraction score up to 93, while the water-rescue procedure scores zero under the strict rule.

The demo mix is weight 0.25 on the exact-field rate, weight 0.35 on average partial credit, and weight 0.40 on average F1, taken across every field on the sheet.

FormulaFIG. 06

Demo mix: 25 percent exact gate, 35 percent partial credit, 40 percent mean F1. Extraction score equals 0.25 times gate plus 0.35 times partial plus 0.40 times F1.

Composed extraction score · demo weights

  • 25%Exact gate
  • 35%Partial credit
  • 40%Mean F1

Extraction score = 0.25·gate + 0.35·partial + 0.40·F1

Run the veto test
FIG. 06. A trend tool. Not a ship decision until the weights survive a veto test.

With this mix, GLM-5V-Turbo lands at 93. GPT-5.4 Mini at 90. Grok 4.5 at 90.

GaugesFIG. 07

Composed scores on the demo mix: GLM-5V-Turbo 93, GPT-5.4 Mini 90, Grok 4.5 90. Submersion ordered-steps is 0.

Composed score. Three vision pipelines

  • 93

    GLM-5V-Turbo

    Z.AI

  • 90

    GPT-5.4 Mini

    OpenAI

  • 90

    Grok 4.5

    xAI

Submersion ordered-steps0

A composed score is the right tool to record a run and follow a trend. It is the wrong tool for a ship / no-ship decision when those weights have not been set to the operational risk. Use it to follow a trend across runs or tune the weights until important fields can move the headline.

07

#

The veto test, in five steps.

People ask how to “choose weights.” Weights are a claim about which mistake you are willing to average away. Here is the method I use:

  1. Write the feared failure in one sentence, with a subject. “The extractor skips PPE on a submerged-vehicle procedure.” If you cannot write this, you cannot set weights and you are just collecting scores.
  2. Tag every field Gate, Credit, or Noise. Gate: a wrong answer means do not ship. Often a veto — exact match. Credit: partial overlap is allowed and useful (a warning list, a paraphrasable instruction). Noise: easy fields that inflate the average (manufacturer, model year, door count). For example, on the Cybertruck sheet: submersion ordered-steps is Gate. Fire prohibition is Gate. Manufacturer is Noise. Occupant-access methods are Credit.
  3. Run the veto test. Set the gate field to zero. Leave every other field at the score you actually observed. Recompute the headline. If that headline still looks like a ship (80, 90, 93) the weights are wrong and they have not been set to the risk. This is the test this demo failed: procedure at 0, composed extraction score at 93.
  4. Retune until a zero on the gate field pulls the headline below the written line. You can raise the gate field’s weight in the average or stop averaging it. For example, across the 34 fields of this demo the procedure has a 5% say. If the procedure gets 40% of the headline, the same run lands near 58 and it no longer looks shippable. The models did not get worse.
  5. Write the line as a conjunction. “We ship if composed ≥ T and the gate field ≥ G.” Product owners pick the conjunction. Subject-matter experts pick G. Engineers encode both. If your current mix cannot fail the veto test you should change the mix (rather than argue about the model).
Veto testFIG. 08

Veto test. The gate field stays at 0. At a 5% share the implied headline is 93.0. That still looks like a ship. At 40% it is 58.7. That is below 80. The conjunction holds because the gate field is 0.

Working session · veto test · default mix

Share of the headline given to the gate field

  • Looks like a ship
  • Below 80
  • 80
Implied headline for each share of the headline given to the gate field, gate field at 0
Share5%10%15%20%25%30%35%40%
Headline93.088.183.278.373.468.563.658.7
ReadingLooks like a shipLooks like a shipLooks like a shipBelow 80Below 80Below 80Below 80Below 80

Gate field (ordered-steps) is 0. Easy fields still carry the rest of the average.

ConjunctionHold

Ship only if the headline clears T and the gate field clears G. The gate field is 0, so the conjunction holds.

FIG. 08. Implied by the essay. The demo gives this field a 5% say and still prints 93. Every column keeps that rest mean and holds the gate at 0.

08

#

The scorer did not blink

The first experiment was three vision pipelines on page images. After that I added LlamaExtract, through LlamaParse. LlamaIndex ships four extract modes: turbo, cost-effective, agentic, and agentic plus. Same rescue-sheet schema. Same scorer.

With the same mix that produced 93, 90, 90 on vision I get the results below. 87% is the best of the four, but it is lower than every previous vision headline. At least the ranking follows my instinct and picks agentic plus. A ranking against vision still picks GLM-5V-Turbo at 93. The strict procedure still fails on every mode.

LlamaParseFIG. 09

LlamaParse composed scores: turbo 83, cost-effective 79, agentic 84, agentic plus 87. Exact sequence on submersion is 0 on every row. Cost-effective omitted fire.

LlamaIndex · LlamaParse · four extract modes

Four LlamaParse extract modes on the same gold
ModeScoreP / R / F1CreditsTimeSubmersionFire
Turbo23/34 exact · 688387 / 89 / 85140140 extract · no parse job17.34sEmpty · F1 0Held
Cost-effective22/34 exact · 657989 / 84 / 836040 parse + 20 extract111.49sEmpty · F1 0Omitted
Agentic22/34 exact · 658493 / 92 / 9010040 parse + 60 extract83.84sEmpty · F1 0Held
Agentic plus24/34 exact · 7187, best of the four92 / 93 / 9124040 parse + 200 extract93.97s3 steps · exact 0Held

Agentic plus. Submersion 3 steps · exact 0.

FIG. 09. Same gold, same 34-field scorer. Agentic plus is 87, best of the four, still under GLM-5V-Turbo. Exact sequence is 0 on every row.
Agentic plusFIG. 10

Agentic plus wrote three ordered steps, not five. It opened on a platitude, put PPE in the wrong slot, and fused remove-from-water with the high-voltage disable.

Agentic plus. submersion.ordered_steps. Three lines, not five

  1. 01

    Treat a submerged Cybertruck like any other submerged vehicle.

    Not a gold step

  2. 02

    Wear appropriate PPE for water rescue.

    Gold 01, wrong slot

  3. 03

    Remove the vehicle from the water and continue with normal high voltage disabling.

    Fused gold 02 + 03

FIG. 10. Gold 04 and 05 scored a full partial on drainage_lift_requirement. Still not steps 4 and 5. The other three modes left this list empty.

09

#

Do not ship on 93

Seven runs. Vision: 93, 90, 90. LlamaParse: 79, 83, 84, 87.

Agentic plus emerged as the best of the four LlamaParse pipelines, yet it remains below every vision headline. The exact sequence mode for the water-rescue procedure returned 0 on every single run. While the Extraction Score showed movement, the deployment gate did not.

In conclusion, frame the release criteria as a strict conjunction. Ship only if the headline clears threshold T and the gate field clears threshold G. Display the worst-performing field first on the dashboard. When an 87 baseline defines the trend and the rescue procedure holds the deciding veto, the directive is clear: do not ship.

10 · Video

#

The walkthrough

Three vision models, one gold rescue sheet, and a score that moves when the question changes. The same argument, recorded. There is also a page for the walkthrough.

VIDEO
Play video (opens YouTube in a new tab)

// the player loads from YouTube only when you press play

Extraction Arena: Evaluate Vision LLMs for Document ExtractionWatch on YouTube (opens in a new tab)

11 · Source

#

The harness

Extraction Arena is a public TypeScript repository, martin-cousseau/extraction-arena, on the main branch: a document extraction evaluation framework. Two Node projects: a backend that turns pages into images, and a frontend that scores fields against gold answers. The note is the argument; the repository is the instrument.

git clone https://github.com/martin-cousseau/extraction-arena.git
PathWhat it holds
backend/Express. PDF to PNG at 300 DPI, then the vision calls.
frontend/Datasets, judges, field scores, the comparison.
docker-compose.ymlThe whole app in containers.
README.mdQuick start and the JSON contract.
AGENTS.mdConventions for the repo.

The same piece also ran on X, LinkedIn and TikTok.