Does a retrieval-augmented system get the right answer for the right reasons?
An AI system can give the right answer without any evidence for it. This project measures how often that happens when AI answers questions about company financial reports. It also tests how far that measurement can be trusted.
The full technical report, with every table, interval and design decision, is in TECHNICAL_REPORT.md.
- Retrieval-augmented generation (RAG) answers a question in two steps. First it searches a library of documents for relevant passages. Then a language model writes an answer from those passages.
- Correct means the answer matches the known right answer.
- Grounded means every statement in the answer is backed by the passages the model was shown.
- The two can come apart. A correct answer can come from the model's memory or a lucky guess. Checking accuracy alone cannot tell the difference.
-
Search is the weakest link.
- The best of 36 search setups found any of the evidence for only 26 of 126 questions.
- So the model mostly declines to answer: 74% of all questions in the final system.
- Given the right pages instead, its answer rate on those same questions rises by +69 pp (percentage points).
-
About a third of correct answers are not fully backed by the evidence.
- Even when handed the right pages, 31% of Claude Sonnet 5's correct answers contain a claim the pages do not support.
- But the automatic checker is often wrong about these. A human review of 30 of them found it right only 50% of the time.
- Accuracy alone would miss the problem. The checker alone would overstate it.
-
AI checkers agree with each other far more than with a human. (Agreement is measured with κ: 0 means chance, 1 means perfect.)
- On whether an answer is fully grounded, the checker agrees with itself on a re-run (κ = 0.92) and with a second AI model (κ = 0.72).
- It agrees with blind human labels much less: κ = 0.52. On individual claims, only 0.23.
- RAGAS, a widely used off-the-shelf tool, did no better (κ = 0.56).
- Lesson: AI checkers agreeing with each other is not proof that they are right.
-
Better search helps grounding, but only a little.
- The link between finding the evidence and giving a grounded answer is weak (correlation 0.30 for Claude Sonnet 5, 0.31 for Qwen2.5 3B).
- Finding the evidence mainly changes whether the model answers, not how well.
-
Asking the model to quote its sources changed nothing measurable.
- With the right pages, groundedness moved by 0 pp and citation quality by +1 pp.
- Over three repeated runs, the plain and "quote your source" prompts could not be told apart.
-
Telling the model it may decline stopped it correcting false premises.
- On questions built on made-up figures, Claude Sonnet 5 pointed out the false premise in 4 of 10 answers with the plain prompt.
- With a prompt that allows declining, it did so in 0 of 10. It declined instead.
Correct and grounded are separate properties. Every model and setting has answers in all four boxes. With the right pages in hand (bottom left), 110 of Claude Sonnet 5's 351 correct answers still fail the grounding check.
- 150 questions from FinanceBench, a public benchmark. Each has a known right answer and the exact passage that supports it.
- 72 company reports (annual 10-K and quarterly 10-Q filings) from the SEC's EDGAR database, 8,514 pages in total.
- 40 trick questions written by hand, in four kinds:
- answerable from one report;
- answerable only by combining two reports;
- not answerable from these reports at all;
- built on a false premise (a figure that does not exist).
Think of an open-book exam.
- Retrieved context (realistic). The system searches all the reports itself and passes its top 5 passages to the model. If the search misses, the model never sees the answer.
- Oracle context (evidence guaranteed). The model is handed 5 passages known to contain the answer. The search cannot fail.
- Why both? If the model fails even with the right pages, the model is at fault. If it fails only with its own search results, the search is at fault.
- 36 search setups were compared: 3 ways of cutting reports into passages × 3 text-embedding models × 4 search methods.
- The search methods: meaning-based (dense), keywords only (BM25), a mix of both, and the mix plus a re-ranking model.
- The best setup was fixed before any answers were generated.
- Two models: Claude Sonnet 5 through its API, and Qwen2.5 3B, a small open model run locally on a 4 GB graphics card.
- Four prompt styles: plain, "quote your source", "think step by step", and "you may decline".
- Every answer is saved with its prompt and passages, so it can be replayed exactly.
- A judge model (Claude Sonnet 5) splits the answer into short claims.
- It checks each claim against the 5 passages: supported, unsupported or contradicted.
- The groundedness score is the share of claims that are supported. An answer is fully grounded when every claim is supported.
Three other checks run alongside:
- Correctness against the known answer. Numbers count as right within 1%, and unit mix-ups (thousands vs millions) are flagged separately.
- Citation quality: do the passages the model cites actually support it?
- Declining: did the model decline to answer, and should it have?
- Human labels: 50 answers (202 claims) were labelled by hand, without seeing the judge's verdicts.
- Other raters: the judge was compared with those labels, with a second AI model (Haiku 4.5), with itself on a re-run, and with RAGAS.
- Predictions first: they were written down before the analysis was run. 10 of the 20 checkable predictions turned out wrong.
- Repeated runs: the two leading prompts were run three times each to measure run-to-run variation. Everything else is a single run.
| Search method (best setup) | Evidence found in top 5 | Right report in top 5 |
|---|---|---|
| Meaning-based (dense) | 18% | 77% |
| Mixed, then re-ranked | 10% | 72% |
| Mixed (meaning + keywords) | 10% | 61% |
| Keywords only (BM25) | 3% | 34% |
Source: results/metrics/retrieval_grid.json (126 questions)
- "Evidence found" is recall@5: the average share of a question's evidence that lands in the top 5 results.
- Meaning-based search wins clearly.
- Keyword search often lands on the wrong company. Many questions share the same wording, so keywords match other companies' reports.
- The search method matters most, explaining 78.9% of the differences between setups.
| Model | Pages given | Correct | Fully grounded (of scored answers) | Declined |
|---|---|---|---|---|
| Claude Sonnet 5 | Handed the right pages | 70% | 65% | 16% |
| Qwen2.5 3B | Handed the right pages | 18% | 27% | 23% |
| Claude Sonnet 5 | Its own search results | 10% | 59% | 78% |
| Qwen2.5 3B | Its own search results | 6% | 10% | 47% |
Source: results/metrics/eval_main.json (four prompts pooled; single run)
- With the right pages, Claude Sonnet 5 answers 70% of questions correctly.
- With its own search results, only 10%. It declines most questions rather than guess.
- When it does answer from its own search results, its answers are nearly as well grounded as with the right pages.
- Qwen2.5 3B is much weaker in both settings. Most of its answers are both wrong and unsupported.
| Who is compared (50 answers) | κ on "fully grounded" | Same verdict |
|---|---|---|
| Judge vs human | 0.52 [0.27, 0.75] | 76% |
| Second judge (Haiku 4.5) vs human | 0.40 [0.14, 0.64] | 70% |
| RAGAS vs human | 0.56 [0.32, 0.76] | 78% |
| Judge vs second judge | 0.72 [0.52, 0.88] | 86% |
| Judge vs itself (re-run) | 0.92 [0.80, 1.00] | 96% |
Source: results/metrics/judge_agreement.json, results/metrics/ragas_faithfulness.json (95% intervals in brackets)
- The AI raters agree with each other far more than with the human labels.
- RAGAS lands in the same place as this project's judge. Two independent tools hitting the same limit suggests the limit is the definition of "supported", not the tool.
- This comparison even favours the project's judge, because the human labels were given on its own claims.
Each dot is one pair of raters. Blue pairs include the human; grey pairs are AI against AI. Every grey estimate sits to the right of every blue one.
| Share of all 150 questions | Plain prompt | "Quote your source" prompt |
|---|---|---|
| Correct and fully grounded | 6.7% ± 1.3 pp | 6.2% ± 1.7 pp |
| Correct | 9.6% ± 1.0 pp | 9.1% ± 1.0 pp |
| Declined | 74.0% ± 0.7 pp | 76.0% ± 0.7 pp |
Source: results/metrics/replicates.json (mean ± standard deviation over three runs)
- The final system was chosen by a rule fixed in advance: the highest share of answers that are both correct and fully grounded.
- The plain prompt won that rule in 3 of 3 runs, but the two prompts cannot really be told apart.
- The first run was the luckiest of the three (8.0% against 6.7% on average). One run of this system is just one sample.
- Two reports needed: with the right pages, Claude Sonnet 5 answers all of these correctly. With its own search results, only 1 of 10 questions finds any evidence, so it mostly declines.
- Not answerable: Claude Sonnet 5 declines 95% of the time. It never answered from its own memory; Qwen2.5 3B did in 20% of answers.
- False premise: Claude Sonnet 5 points out the false premise in 30% of answers, Qwen2.5 3B in 5%. Qwen2.5 3B used the made-up figure as if it were real in 13 of 40 answers.
- Two opposite biases. The human was more lenient on company and year matching. The human was stricter on figures that had to be worked out by arithmetic or read across several passages.
- A human review of 30 "correct but ungrounded" answers found 14 wrong verdicts:
- 9 were the judge breaking its own written rules, such as treating a statement about the passages as a fact about the company.
- 5 were the judge applying a rule correctly where the rule itself is debatable, such as calling "not a high-growth company" unsupported because no passage says so.
- The first kind is a fixable judge error. The second is a definition choice that no better judge would remove.
The selected system runs as a web service that returns every answer with its groundedness score.
A grounded answer; a correct answer that scores zero because the passages don't support it; and the agreement figure that travels with every score.
- Every score comes with its reliability: the judge's measured agreement with the human labels (κ = 0.52).
- Faithful to the evaluation: the service reproduces all 150 evaluated answers and their scores exactly.
- Speed: a live answer takes 8.4 s at the median and costs $0.016.
- Checking is the expensive part: the groundedness check is 75% of the cost and 75% of the time. Checking an answer costs about 3 times as much as writing it.
- Load test: with cached answers, the service handles about 26 requests per second.
- Kubernetes: tested on a local cluster, including automatic scaling to more copies under load.
- Setup time: from a fresh download to the first answer took 277.4 s (one try, one home connection), within a five-minute target.
The design is drawn in docs/architecture.md.
- Install: dependencies are managed with
uv. Development used an RTX 3050 laptop GPU with 4 GB. - Data: the reports and every saved model response are published as a GitHub release, so the whole pipeline replays without calling any AI service.
- Time and cost: a full rebuild takes about 1.7 h on that GPU and needs no API spending. The original runs cost $27.15 in total.
uv sync
uv run python scripts/19_data_release.py restore # the reports and every saved model response
docker compose up -d qdrant # the second vector database
dvc repro --force # rebuild every result, no AI calls
uv run python scripts/17_readme.py --check # confirm these documents match the resultsThe technical report lists every step with its measured run time.
- One human, 50 answers. Every groundedness figure rests on one person's labels, and the AI checkers only partly agree with them.
- Mostly single runs. Only the two leading prompts were repeated. The repeats show a single run can move by several questions.
- Small and narrow. 150 questions, one domain (US company filings), one document type. Results may not carry over to contracts, medical records or other text.
- Metric blind spots. An answer can be "fully grounded" and still wrong, if it faithfully repeats a passage about the wrong year or the wrong company.
- Not production-ready. The service has no login, no rate limiting beyond a spending cap, and no process for adding new filings.
| Package | Purpose |
|---|---|
| anthropic | Claude Sonnet 5 answers and the groundedness judge |
| ollama | Running Qwen2.5 3B locally |
| sentence-transformers | The embedding models and the re-ranking model |
| faiss-cpu, qdrant-client | The two vector databases, behind one interface |
| rank-bm25 | Keyword search |
| lxml | Parsing the HTML filings, telling layout tables from data tables |
| ragas | The off-the-shelf checker compared with this project's judge |
| mlflow, dvc | Experiment tracking and the versioned data pipeline |
| fastapi, uvicorn, prometheus-client | The web service and its metrics |
| locust | The load test |
| matplotlib | Every plot |
Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement.
Es, S., James, J., Espinosa-Anke, L., & Schockaert, S. (2023). RAGAS: Automated Evaluation of Retrieval Augmented Generation. arXiv:2309.15217.
Islam, P., Kannappan, A., Kiela, D., Qian, R., Scherrer, N., & Vidgen, B. (2023). FinanceBench: A New Benchmark for Financial Question Answering. arXiv:2311.11944.
Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W. (2020). Dense Passage Retrieval for Open-Domain Question Answering. EMNLP.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS.
Robertson, S., & Zaragoza, H. (2009). The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval.
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks.