← Back to home

FinDocRAG: Grounded Q&A over Annual Reports

Oct 2026 · 7 min read · personal project

Python RAG Claude API FAISS BM25 LLM evaluation Gradio

Language models answer confidently whether or not the documents support the answer. In finance and advisory work that is the wrong way round: a wrong figure with no source is worse than no answer at all. The question I wanted to answer was whether a small system could read 977 pages of annual reports from Equinor, DNB and Norsk Hydro, cite the page behind every claim, and refuse when the reports do not contain the answer.

The second question mattered more to me: could I prove it works with measurement, rather than with a handful of good-looking demo answers?

I built the evaluation before trusting any result. Every question has a reference answer that I checked by hand against a highlighted crop of the source page. One question was dropped and four were corrected during that check. Every decision is recorded in the repository.

I kept two question sets. A development set of 39 questions (29 answerable, 10 that must be refused) was used to find failures and choose fixes. A held-out set of 20 new questions, several of them deliberate traps, was written and verified only after all changes and then run once with the system frozen. That second number is the honest one.

When the system refuses, it does not just say no. It says the answer is not in the reports and then explains, with citations, what related information they do contain, for example an annual average instead of a closing price. Those explanations are checked for grounding like any other answer, so a made-up explanation costs points.

Ingestion: The three PDFs are parsed page by page and split into 1,377 chunks of about 800 tokens, each tied to its company and page so that every citation points to an exact page.

Hybrid retrieval: Local bge-small embeddings in FAISS combined with BM25 keyword search through reciprocal rank fusion. If a question names one company, only that report is searched. Claude Haiku then reranks the top 30 candidates down to the 8 most useful passages.

Grounded answering: Claude answers only from those passages, with an inline [Company, p.X] citation for every claim, or abstains with a cited explanation.

Evaluation harness: An LLM judge scores correctness and groundedness. Citation validity, figure-on-cited-page and retrieval hit rate are checked programmatically with no model involved. Abstention is scored in both directions: refusing what it should, and only that. All calls are cached and keyed on the prompts, so a changed prompt can never reuse a stale result.

Live app: A Gradio app with an Ask tab (answers with page citations, a clear abstention badge, and the retrieved passages so anyone can check the grounding) and an Evaluation tab with the scorecards and the known limitations.

Report pages

977

Verified questions

59

Held-out correctness

93%

The first honest baseline, using embedding search only, scored 81% correctness and refused three questions it should have answered. The triage was the most useful part of the project: five of the six failures were retrieval misses, not answering mistakes. The right page never reached the model. All five were table pages, an income statement and headcount and production tables, which embed poorly because they are mostly numbers but contain the exact words of the question. That pointed straight at keyword search.

After one improvement round (hybrid search, the company filter and reranking) the development set reached 100%. Because the fixes were chosen on that set, I treat that number as optimistic. On the 20 unseen held-out questions the system scored 93.3% correctness (14 of 15), refused all 5 trap questions, and all 24 of its citations pointed to pages it had actually retrieved.

The one held-out failure is instructive. DNB's 2018 dividend exists only in a bar chart. The right page was retrieved, but converting the chart to text loses the link between each year and its value, so the system abstained rather than guess. That is the failure mode I would rather have.

Measure retrieval on its own. Without a retrieval hit-rate metric I would have spent time on prompts when the real problem was that the model never saw the right page.

The answer key can be wrong too. One "wrong" answer turned out to be a figure the report states on a different basis on another page. Fixing the key, openly and after the baseline, was more honest than tuning the system to match it.

LLM judges are lenient in places. On one question the correctness judge accepted an answer that treated two different figures as the same, and only the groundedness check caught it. Several independent checks, some with no model at all, are worth more than one clever judge. And a held-out set run once is what turns "it works on my questions" into a number I can stand behind.