Note, , 6 min read, Updated
Extraction Arena
Can we trust a model with a rescue procedure?
A strong average can hide a procedure the model gets entirely wrong. Gate the field you fear, not the mean.
Contents

Summary
Three vision models read Tesla’s Cybertruck first-responder documentation and filled a structured record. The headline scores look ready to ship (93, 90, 90). The water-rescue procedure does not.
I recorded a walkthrough of the harness behind this note, and the code is open source: martin-cousseau/extraction-arena.
Composed scores: GLM-5V-Turbo 93, GPT-5.4 Mini 90, Grok 4.5 90. Submersion exact sequence is 0. LlamaParse in mode order: turbo 83, cost-effective 79, agentic 84, agentic plus 87.
Composed score · three vision pipelines
- 93GLM-5V-TurboZ.AI
- 90GPT-5.4 MiniOpenAI
- 90Grok 4.5xAI
LlamaParse
I keep the model answers frozen and change only the scoring rule, so I can see whether I have a model problem or a definition problem. Later I add a fourth pipeline, LlamaParse through LlamaExtract, and read that run the same way.
The business question does not move. Would you accept a high extraction score if the one procedure a crew would follow still scores zero?
A go / no-go decision, not just a leaderboard
For product owners deciding go / no-go, subject-matter experts deciding whether order and wording are allowed to move on a rescue step, or engineers who have to turn that decision into weights. If you ship document extraction into operations, this is the meeting you skip at your cost. The average will look fine. The field a crew would actually follow might not.
This is not a clean text document
Tesla publishes official rescue information for the Cybertruck. A rescue sheet (ISO 17840-1) is the quick page at the scene. An Emergency Response Guide (ISO 17840-3) is the longer procedure set used to train and operate. The four pages the models saw mix diagrams, pictograms, numbered procedures, and prohibitions in capital letters. A wrong cut zone or a skipped disable step is an operational failure, not just a formatting issue. I selected and scored the following fields: cut zones, high-voltage disable, fire, submersion, and tow.
The first experiment ran in four steps.
- Each page is rendered as an image (the models see what a person sees on paper).
- Every model gets the same empty form; the correct answers are never in the prompt.
- GLM-5V-Turbo, GPT-5.4 Mini, and Grok 4.5 each return their own JSON; I map those answers onto one shared form.
- Then I score field by field, because the business risk is not the same on manufacturer, on a prohibition, and on a five-step procedure.
Four pages of the public Cybertruck rescue sheet, drawn as text. Layout scores cut zones. Disable scores high-voltage disable. Occupants marks occupant access as credit. Fire, water, and tow score fire and submersion as gates, and tow.
Tesla Cybertruck rescue sheet · four pages
01Layout · airbags · HV
- Diagram
- Legend
- Cut zones
- Scored
The models saw this page as a picture of the vehicle, the airbags, and the high-voltage zones.
02Identify · immobilize · disable
- Numbered procedure
- Prohibition
- High-voltage disable
- Scored
Identification, immobilization, and the disable sequence. A skipped disable step is an operational failure.
03Occupants · stored energy
- Diagram
- Numbered procedure
- Occupant access
- credit
How a crew reaches the occupants, and where the stored energy sits. Occupant access is credit, not a veto.
04Fire · water · tow
- Pictogram
- Prohibition
- Numbered procedure
- Fire
- gate
- Submersion
- gate
- Tow
- Scored
Pictograms, prohibitions in capital letters, then the fire, submersion, and tow procedures.
I already have the answers. I am changing the question.
I am not asking the models again. I am changing how a correct answer is defined for one field. I zoomed in on the five-step procedure for a submerged Cybertruck. Product owners pick the scoring rule that matches the risk. Subject-matter experts say whether order and wording are allowed to move.
Six lenses on one field: precision, recall, F1, sequence, set, and exact versus partial.
What the scores actually mean
Six lenses. One field.
- 01Precision
- Of everything the model wrote, how much was actually on the sheet. Low precision means it invented steps. Right lens for a prohibition.
- Precision is the lens for a prohibition. Invented steps fail here.
- 02Recall
- Of everything the sheet required, how much the model captured. Low recall means it skipped steps. Right lens for a procedure.
- Recall is the lens for a procedure. Skipped steps fail here.
- 03F1
- The balance of the two. If the model invents half the procedure or skips half, this number falls. Not a popularity score.
- F1 holds both costs. Inventing half or skipping half pulls it down.
- 04Sequence
- Order is part of the answer. Step four sitting in position one is a miss. You do not start a water rescue by lifting the front.
- One model opened at gold step 4. Under sequence that is a miss in position 1.
- 05Set
- Order does not count. A list of warnings is correct if the items are present, even shuffled. Not for a procedure a crew must perform in order.
- Under set, the step is merely present, and a crew could skip PPE.
- 06Exact / partial
- Exact: the shipped value must be that wording. Partial: overlapping words still score. Gold step 1 is the PPE line.
- Exact match fails "Wear PPE" against the gold "Wear appropriate PPE for water rescue." Partial match still scores the overlapping words.
How close must the wording be? Exact or partial. If too strict, a useful extraction looks like a failure. If too loose, a dangerous paraphrase looks like a pass. Example: exact match fails “Wear PPE” against the gold “Wear appropriate PPE for water rescue.” Partial match still scores the overlapping words.
Does order count? Sequence or set. A rescue procedure is a sequence of actions. Sequence compares step 1 to gold step 1. Set treats the list as a bag. Example: one model opened at gold step 4. Under sequence that is a miss in position 1. Under set, the step is merely present, and a crew could skip PPE.
The question, which mistake to surface first, is a product decision, not a scoring one.
Precision: of everything the model wrote, how much was on the sheet — the lens for a prohibition.
Recall: of everything the sheet required, how much the model captured — the lens for a procedure.
And F1 holds both costs.
Three stories. None of them improved the models.
None of the three vision models returned the five-step order.
In the illustration below a filled cell is a step the model actually wrote, a fused cell is two gold steps crushed into one line, an empty cell is a skipped step.
- One model opened at gold step 4.
- The other two started on PPE, fused remove-from-water with the high-voltage disable, then stopped.
All three saw the submersion paragraph. None produced a five-step ordered list.
Gold is five ordered steps: PPE, remove from water, continue HV disable, raise the front about 30 cm, store flat. One vision model opened at step 4. The other two fused remove with disable. Three LlamaParse modes left the list empty. Agentic plus wrote three lines, not five.
Ground truth. The five ordered steps on the gold record
- Filled: a step the model wrote.
- Fused: two gold steps crushed into one line.
- Empty: a skipped step.
Gold record
- 01PPE for water rescue
- 02Remove from water
- 03Continue HV disable
- 04Raise front ~30 cm
- 05Store flat
- 01PPE for water rescue
- 02Remove from water
- 03Continue HV disable
- 04Raise front ~30 cm
- 05Store flat
One vision model
- PPEEmpty
- RemoveEmpty
- DisableEmpty
- RaiseFilled
- StoreEmpty
Opened at gold step 4, "raise the front." PPE and extraction never happen.
The other two
- PPEFilled
- Remove + disableFused
- RaiseEmpty
- StoreEmpty
Started on PPE, fused remove-from-water with the high-voltage disable, then stopped.
Turbo · cost-effective · agentic
- PPEEmpty
- RemoveEmpty
- DisableEmpty
- RaiseEmpty
- StoreEmpty
LlamaParse. Ordered steps empty on three modes. Precision 1, recall 0. Omission, not invention.
LlamaParse · agentic plus
- PPE · slot 2Filled
- Remove + disableFused
- Raise ‡Empty
- Store ‡Empty
Wrote three ordered steps. Opened on a platitude. Fused remove with disable. 30 cm still in another field.
I then read those frozen answers under three scenarios. Only the scoring rule changes.
- Exact wording, order required: precision 0, recall 0, F1 0.
- Flexible wording, order still required: 44, 25, 33.
- Flexible wording, order ignored: 81, 47, 58.
All of these results are true. They just answer different questions.
I did not improve the models, but I changed the question. For this field a kinder pairing rule increases the score and rewards an extracted procedure no crew should follow.
Mean of three vision models on submersion ordered-steps. Exact sequence: precision 0, recall 0, F1 0. Partial sequence: 44, 25, 33. Partial set: 81, 47, 58.
Submersion ordered steps · mean of three models
- Precision %
- Recall %
- F1 %
| Scoring rule | Precision | Recall | F1 |
|---|---|---|---|
| Exact · sequenceThe question you would hand to a crew. | 0 | 0 | 0 |
| Partial · sequenceFlexible wording, order still required. | 44 | 25 | 33 |
| Partial · setFlexible wording, order ignored. A procedure no crew should follow. | 81 | 47 | 58 |
The mix still looks shippable.
A well-designed metric should let you trust the headline. That is the goal. A composed score only earns that trust if its weights match the risk of the work.
But in this demo they do not: manufacturer, vehicle type and other easy fields pull the extraction score up to 93, while the water-rescue procedure scores zero under the strict rule.
The demo mix is weight 0.25 on the exact-field rate, weight 0.35 on average partial credit, and weight 0.40 on average F1, taken across every field on the sheet.
Demo mix: 25 percent exact gate, 35 percent partial credit, 40 percent mean F1. Extraction score equals 0.25 times gate plus 0.35 times partial plus 0.40 times F1.
Composed extraction score · demo weights
- 25%Exact gate
- 35%Partial credit
- 40%Mean F1
Extraction score = 0.25·gate + 0.35·partial + 0.40·F1
With this mix, GLM-5V-Turbo lands at 93. GPT-5.4 Mini at 90. Grok 4.5 at 90.
Composed scores on the demo mix: GLM-5V-Turbo 93, GPT-5.4 Mini 90, Grok 4.5 90. Submersion ordered-steps is 0.
Composed score. Three vision pipelines
- 93
GLM-5V-Turbo
Z.AI
- 90
GPT-5.4 Mini
OpenAI
- 90
Grok 4.5
xAI
Submersion ordered-steps0
A composed score is the right tool to record a run and follow a trend. It is the wrong tool for a ship / no-ship decision when those weights have not been set to the operational risk. Use it to follow a trend across runs or tune the weights until important fields can move the headline.
The veto test, in five steps.
People ask how to “choose weights.” Weights are a claim about which mistake you are willing to average away. Here is the method I use:
- Write the feared failure in one sentence, with a subject. “The extractor skips PPE on a submerged-vehicle procedure.” If you cannot write this, you cannot set weights and you are just collecting scores.
- Tag every field Gate, Credit, or Noise. Gate: a wrong answer means do not ship. Often a veto — exact match. Credit: partial overlap is allowed and useful (a warning list, a paraphrasable instruction). Noise: easy fields that inflate the average (manufacturer, model year, door count). For example, on the Cybertruck sheet: submersion ordered-steps is Gate. Fire prohibition is Gate. Manufacturer is Noise. Occupant-access methods are Credit.
- Run the veto test. Set the gate field to zero. Leave every other field at the score you actually observed. Recompute the headline. If that headline still looks like a ship (80, 90, 93) the weights are wrong and they have not been set to the risk. This is the test this demo failed: procedure at 0, composed extraction score at 93.
- Retune until a zero on the gate field pulls the headline below the written line. You can raise the gate field’s weight in the average or stop averaging it. For example, across the 34 fields of this demo the procedure has a 5% say. If the procedure gets 40% of the headline, the same run lands near 58 and it no longer looks shippable. The models did not get worse.
- Write the line as a conjunction. “We ship if composed ≥ T and the gate field ≥ G.” Product owners pick the conjunction. Subject-matter experts pick G. Engineers encode both. If your current mix cannot fail the veto test you should change the mix (rather than argue about the model).
Veto test. The gate field stays at 0. At a 5% share the implied headline is 93.0. That still looks like a ship. At 40% it is 58.7. That is below 80. The conjunction holds because the gate field is 0.
Working session · veto test · default mix
Share of the headline given to the gate field
- Looks like a ship
- Below 80
- 80
| Share | 5% | 10% | 15% | 20% | 25% | 30% | 35% | 40% |
|---|---|---|---|---|---|---|---|---|
| Headline | 93.0 | 88.1 | 83.2 | 78.3 | 73.4 | 68.5 | 63.6 | 58.7 |
| Reading | Looks like a ship | Looks like a ship | Looks like a ship | Below 80 | Below 80 | Below 80 | Below 80 | Below 80 |
Gate field (ordered-steps) is 0. Easy fields still carry the rest of the average.
ConjunctionHold
Ship only if the headline clears T and the gate field clears G. The gate field is 0, so the conjunction holds.
The scorer did not blink
The first experiment was three vision pipelines on page images. After that I added LlamaExtract, through LlamaParse. LlamaIndex ships four extract modes: turbo, cost-effective, agentic, and agentic plus. Same rescue-sheet schema. Same scorer.
With the same mix that produced 93, 90, 90 on vision I get the results below. 87% is the best of the four, but it is lower than every previous vision headline. At least the ranking follows my instinct and picks agentic plus. A ranking against vision still picks GLM-5V-Turbo at 93. The strict procedure still fails on every mode.
LlamaParse composed scores: turbo 83, cost-effective 79, agentic 84, agentic plus 87. Exact sequence on submersion is 0 on every row. Cost-effective omitted fire.
LlamaIndex · LlamaParse · four extract modes
| Mode | Score | P / R / F1 | Credits | Time | Submersion | Fire |
|---|---|---|---|---|---|---|
| Turbo23/34 exact · 68 | 83 | 87 / 89 / 85 | 140140 extract · no parse job | 17.34s | Empty · F1 0 | Held |
| Cost-effective22/34 exact · 65 | 79 | 89 / 84 / 83 | 6040 parse + 20 extract | 111.49s | Empty · F1 0 | Omitted |
| Agentic22/34 exact · 65 | 84 | 93 / 92 / 90 | 10040 parse + 60 extract | 83.84s | Empty · F1 0 | Held |
| Agentic plus24/34 exact · 71 | 87, best of the four | 92 / 93 / 91 | 24040 parse + 200 extract | 93.97s | 3 steps · exact 0 | Held |
Agentic plus. Submersion 3 steps · exact 0.
Agentic plus wrote three ordered steps, not five. It opened on a platitude, put PPE in the wrong slot, and fused remove-from-water with the high-voltage disable.
Agentic plus. submersion.ordered_steps. Three lines, not five
- 01
Treat a submerged Cybertruck like any other submerged vehicle.
Not a gold step
- 02
Wear appropriate PPE for water rescue.
Gold 01, wrong slot
- 03
Remove the vehicle from the water and continue with normal high voltage disabling.
Fused gold 02 + 03
Do not ship on 93
Seven runs. Vision: 93, 90, 90. LlamaParse: 79, 83, 84, 87.
Agentic plus emerged as the best of the four LlamaParse pipelines, yet it remains below every vision headline. The exact sequence mode for the water-rescue procedure returned 0 on every single run. While the Extraction Score showed movement, the deployment gate did not.
In conclusion, frame the release criteria as a strict conjunction. Ship only if the headline clears threshold T and the gate field clears threshold G. Display the worst-performing field first on the dashboard. When an 87 baseline defines the trend and the rescue procedure holds the deciding veto, the directive is clear: do not ship.
The walkthrough
Three vision models, one gold rescue sheet, and a score that moves when the question changes. The same argument, recorded. There is also a page for the walkthrough.

// the player loads from YouTube only when you press play
The harness
Extraction Arena is a public TypeScript repository, martin-cousseau/extraction-arena, on the main branch: a document extraction evaluation framework. Two Node projects: a backend that turns pages into images, and a frontend that scores fields against gold answers. The note is the argument; the repository is the instrument.
git clone https://github.com/martin-cousseau/extraction-arena.git
| Path | What it holds |
|---|---|
backend/ | Express. PDF to PNG at 300 DPI, then the vision calls. |
frontend/ | Datasets, judges, field scores, the comparison. |
docker-compose.yml | The whole app in containers. |
README.md | Quick start and the JSON contract. |
AGENTS.md | Conventions for the repo. |