Which Frontier AI Model Has the Lowest Hallucination Rate?
The short answer
No frontier model is hallucination-free, and the “lowest” depends on what you measure. As of October 9, 2026:
- Admitting ignorance (AA-Omniscience): Gemini 4 Argon is lowest at 15.1%, but it is restricted to Google’s Fairwind cyber-defender program. Among models you can use today, Grok 4.7 (29.3%) is lowest, then Claude Sonnet 5.5 (47.0%) and GPT-6 Astra (51.3%).
- Best overall knowledge reliability: Claude Opus 5.5 has the highest AA-Omniscience Index (46.4) — it gets the most right after subtracting wrong answers.
- Faithful summaries of a given document (Vectara HHEM): GPT-6 Sol 6.5%, GPT-6 Astra 8.7%.
AA-Omniscience results (frontier models, max/high effort)
| Model | Accuracy | Hallucination rate (lower is better) | Omniscience Index (higher is better) |
|---|---|---|---|
| Claude Opus 5.5 | 66.2% | 58.6% | 46.4 |
| Claude Fable 5.1 | 67.2% | 72.6% | 43.5 |
| GPT-6 Astra | 62.6% | 51.3% | 43.4 |
| Gemini 4 Argon (not public) | 49.9% | 15.1% | 42.4 |
| GPT-6.1 Sol | 62.1% | 54.3% | 41.5 |
| Claude Sonnet 5.5 | 53.9% | 47.0% | 32.3 |
| Grok 4.7 | 47.4% | 29.3% | 32.0 |
AA-Omniscience covers about 6,000 hard questions across 42 economically relevant topics. Its hallucination rate is wrong answers divided by all non-correct responses (wrong + partial + not attempted); the Index rewards correct answers, subtracts wrong ones and does not penalise refusals, on a −100 to 100 scale.
Vectara summarization leaderboard (September 22, 2026)
| Model | Hallucination rate | Answer rate |
|---|---|---|
| GPT-6 Sol | 6.5% | 100% |
| GPT-6 Astra | 8.7% | 100% |
Vectara scores whether a model’s summary of a short document introduces facts the document does not contain, using its HHEM-2.3 judge. Current Claude 5.x and Gemini 3.8/4 models were not on the September board; the lowest overall entries there are small specialised models (the leader, Ant Group’s Finix S1 32B, scores 1.8%).
How to choose
- Factual lookups without tools (research assistant, Q&A): prefer models that decline — Grok 4.7 among public models, or Claude Opus 5.5 for the best balance of right answers to wrong ones.
- Summarising supplied documents (RAG, contract review): grounded faithfulness matters more; GPT-6 Sol and Astra have the best published single-digit rates.
- Highest raw knowledge: Claude Fable 5.1, but verify — it guesses aggressively.
- Price check: Opus 5.5 is $4/$20 per million tokens, GPT-6 Astra $10/$50, Grok 4.7 $2/$6 — see current API prices.
Whatever the model, retrieval, citations and an explicit permission to say “I don’t know” cut errors more than switching vendors. For extraction pipelines see how to stop hallucinated fields in structured data extraction.
Last verified: October 9, 2026. Benchmark boards update often; figures are the latest published at that date.