How to Stop LLMs Inventing Fields in Data Extraction (2026)
The short answer
Most hallucinated fields in LLM extraction are not the model being creative — they are the model filling a required slot with the nearest plausible thing on the page. The fix is mostly structural: make fields nullable, tell the model in one sentence that null is the correct answer when the value is absent, verify each value with a cheap second model, and test with pages that contain decoys. A September 27, 2026 benchmark of 16 models found the single instruction “Use null for any field whose value is not on the page. Do not guess.” cut invented fields from 70.7% to 20.2% on average — every one of the 16 models improved.
Step 1: Make absence a legal answer
If your schema has price: number with no null option, the model will find a number. Use price: number | null (or Optional[float] in Pydantic, .nullable() in Zod) for every field that could be missing, and use structured-output or JSON-mode features so the schema is enforced rather than suggested. Then add the sentence, verbatim or close to it:
Use null for any field whose value is not on the page. Do not guess.
In the benchmark, on the “Was $493.00” page — an old price shown as a decoy — all 16 models returned 493 as the current price without that sentence; with it, one did.
Step 2: Test with twin pages and decoys
Build pairs of pages that differ by exactly one row: one shows the answer, the other does not. Both carry the same decoy. The benchmark’s decoys are a good starting set:
- “Was $493.00” — an old price, not the current price.
- “Fact-checked by Omar Tamm” — a reviewer, not the author.
- “Last updated September 7, 2020” — an update date, not the publication date.
- A press-inquiry email where a contact email is asked for (ambiguous — the benchmark excluded these from scoring, and you should decide your own policy).
- A venue listed as “TBA” — counts as made up if returned as a value.
Score only the pages where the field is missing. An honest extractor returns the value on page A and null on page B. Seven page types and 42 pairs were enough to separate the top of the field from the bottom; the middle overlaps at 95% confidence, so do not over-read rank order.
Step 3: Choose a model that stays honest
Made-up fields out of 36 missing, with the instruction, one run each, September 27, 2026:
| Model | With instruction | Without | Run cost (84 pages) |
|---|---|---|---|
| Gemini 3.8 Flash | 1 | 14 | $0.16 |
| GLM 5.3 | 1 | 18 | $0.17 |
| Hy3 | 3 | 22 | $0.036 |
| DeepSeek V4.1 Flash | 3 | 24 | $0.019 |
| GPT-6 Luna | 5 | 25 | $0.0049 |
| GLM 5.3 Flash | 5 | 22 | $0.018 |
| Claude Sonnet 5 | 5 | 24 | $0.11 |
| GPT-5.6 Sol | 6 | 30 | $0.057 |
| GPT-5.6 Luna | 7 | 28 | $0.010 |
| Qwen 3.8 27B | 7 | 30 | $0.10 |
| Claude Haiku 4.5 | 8 | 25 | $0.036 |
| MiniMax M3 | 8 | 27 | $0.027 |
| Gemma 4 31B | 13 | 26 | $0.0037 |
| MiMo 2.6 Flash | 13 | 26 | $0.0053 |
| Solar Pro 4 | 19 | 35 | $0.0028 |
Three takeaways. Price does not predict honesty: GPT-6 Luna at $0.10/$0.50 per million tokens is in the top tier at a thirtieth of Gemini 3.8 Flash’s run cost. Bigger frontier models are not automatically better: GPT-5.6 Sol invented more than GPT-6 Luna. And without the instruction, every model — including the best — invented most missing fields, so the instruction is not optional for any of them.
Step 4: Add a cheap verifier
After extraction, ask a small model one question per returned value: does the page support “the author is Omar Tamm”? Results from the same benchmark:
| Checker | Fabrications caught | Correct values wrongly rejected | Cost for 126 checks |
|---|---|---|---|
| GPT-6 Luna | 38 / 49 | 0 / 47 | $0.0049 |
| Jev 1.13 (TypeSafe) | 23 / 49 | 0 / 48 | $0.0024 |
Neither rejected a correct value. Luna caught 20 of Firecrawl’s 24 fabrications. Jev, a decision model, caught obvious decoys (wrong author, old price) but missed near-meaning substitutions — resting, cooking or total time returned as prep time, 0 of 6. If your fields have near-neighbours like that, use the general model as the checker. Either way the check costs a fraction of a cent per page, which is cheaper than the extraction it validates. Pricing reference for the checker models: GPT-6 Luna vs Gemini 3.8 Flash vs DeepSeek V4.1 Flash.
Step 5: Be careful with extraction APIs
Paid scraping-plus-extraction APIs bundle the fetch, the cleaning and the model, and in this test they fabricated more than raw models did: Firecrawl 24 of 36, ScrapingBee 16 of 36, ScrapeGraphAI 7 of 31 — versus a plain HTTP fetch, HTML stripped to text, and GPT-6 Luna at 5 of 36. Every Firecrawl fabrication copied the decoy. Caveats: free tiers only, and ScrapingBee has no prompt slot so the null rule had to go into field descriptions. If you use one, run your own twin-page test through it and keep the verifier on.
Step 6: Log and re-run
Log the source snippet the model used for each value, and re-run your twin-page suite whenever you change model, prompt or schema. The September 2026 numbers are one run per contestant on synthetic pages; your pages differ, and model behavior drifts between versions. A suite of 40-80 pages costs under a dollar to run on most of the models above.
Why this works
The models are not lying; they are completing a form. A required field plus an on-page number that looks like the answer is a strong prior. Nullable schema plus an explicit “null is correct” instruction changes the task from “find a value” to “decide whether a value exists,” which is a different and easier judgment. The verifier then catches the residual 20% at negligible cost. For the broader pattern of agents over-reasoning toward wrong answers, see what is the reasoning trap.
Last verified: September 28, 2026. Benchmark figures are from a single run on September 27, 2026; prices from vendor pages.