Noesis

Hallucination benchmark — 2026-07-03

Public scorecard, methodology, and raw data for Noesis's measured hallucination rate against a reference methodology from the legal NLP literature.

Headline result
9.8% hallucination rate
95% Wilson CI [6.7%, 14.2%] · n = 244 classified citations across 100 legal + medical probes
Reference baseline: Magesh et al. 2024 (arXiv:2405.20362) measured Lexis+ AI and Westlaw AI-Assisted Research at 17-33% hallucination. The entire Noesis 95% confidence interval sits below Magesh's lower bound.

Per-domain scorecard

Domain Verified Halluc. Misattr. Classified n Halluc. rate 95% CI
Legal (50 probes) 163 0 14 177 7.9% [4.8%, 12.8%]
Medical (50 probes) 57 1 9 67 14.9% [8.3%, 25.3%]
Overall (100 probes) 220 1 23 244 9.8% [6.7%, 14.2%]

"Misattributed" = the cited source exists but doesn't support the specific claim it was cited for. "Hallucinated" = the cited source doesn't exist. Both count against Noesis in the rate, matching Magesh et al. "Unverifiable" citations (14 of 258) are excluded from the denominator, also matching Magesh — not swept under a rug: they're in the raw JSON.

Methodology

Probe set

Scoring — LLM-as-judge (matching Magesh 2024)

Honest limitations

What this benchmark is NOT

Reproducibility

How to read this vs. the reference baselines

System Domain Hallucination Source
Raw LLMs (GPT-4, Llama 2) Federal court queries 58% - 88% Dahl 2024 arXiv:2401.01301
Foundation LLMs (11 models, median) Medical hallucination tasks 23.4% Kim 2025 arXiv:2503.05777
Lexis+ AI · Westlaw AI-Assisted Research Legal research 17% - 33% Magesh 2024 arXiv:2405.20362
Noesis (this run) Legal + medical fact-recall probes 9.8% [6.7%, 14.2%] This page, 2026-07-03
The defensible claim

Noesis's 95% confidence interval upper bound (14.2%) sits below Magesh's lower bound (17%) with statistical significance at n = 244 classified citations. That is: on this probe set, with this methodology, Noesis hallucinates less than the retail legal-AI products Magesh tested — not by a rounded point estimate, but by a non-overlapping confidence interval.

What we're publishing next

← Back to Noesis