Skip to content

Repository files navigation

RAG Evaluation Workbench

A standard-library workbench for evaluating retrieval-augmented generation outputs.

This project focuses on the pieces that usually get skipped in RAG demos: retrieval quality, citation coverage, required-fact coverage, groundedness checks, failure tags, and reports that explain what went wrong.

What It Demonstrates

  • Markdown corpus ingestion
  • Token-window chunking with overlap
  • BM25-style lexical retrieval
  • Keyword-coverage and hybrid retrieval strategies
  • Retriever comparison reports with metric deltas
  • JSONL evaluation dataset format
  • Retrieval recall@k against expected sources
  • Citation precision and citation recall
  • Required-fact coverage checks
  • Groundedness and failure-mode tagging
  • Markdown and JSON report generation
  • Unit-tested CLI workflow with no required API keys
  • GitHub Actions CI workflow for tests, validation, and sample report generation

Project Structure

.
|-- corpus/
|   `-- tenant_guide/
|       |-- access.md
|       |-- deposits.md
|       |-- rent_notices.md
|       `-- repairs.md
|-- data/
|   `-- eval_cases.jsonl
|-- docs/
|   |-- EVAL_SCHEMA.md
|   `-- RETRIEVER_COMPARISON.md
|-- reports/
|   |-- retriever_comparison.json
|   |-- retriever_comparison.md
|   |-- sample_report.json
|   `-- sample_report.md
|-- src/
|   `-- rag_eval_workbench/
|       |-- cli.py
|       |-- dataset.py
|       |-- evaluation.py
|       |-- ingest.py
|       |-- models.py
|       |-- reporting.py
|       |-- retrieval.py
|       `-- text.py
|-- tests/
|   `-- test_rag_eval_workbench.py
|-- pyproject.toml
|-- README.md
`-- requirements.txt

Quick Start

This project uses only the Python standard library.

Validate the corpus and eval cases:

PYTHONPATH=src python -m rag_eval_workbench.cli --corpus corpus validate data/eval_cases.jsonl

Inspect retrieval for one query:

PYTHONPATH=src python -m rag_eval_workbench.cli --corpus corpus retrieve "What should I do about mold near a window?" --top-k 3

Inspect retrieval with the hybrid strategy:

PYTHONPATH=src python -m rag_eval_workbench.cli --corpus corpus retrieve "What should I do about mold near a window?" --retriever hybrid --top-k 3

Generate reports:

PYTHONPATH=src python -m rag_eval_workbench.cli --corpus corpus evaluate data/eval_cases.jsonl --markdown reports/sample_report.md --json reports/sample_report.json

Compare retrievers:

PYTHONPATH=src python -m rag_eval_workbench.cli --corpus corpus compare data/eval_cases.jsonl --retrievers bm25 keyword hybrid --markdown reports/retriever_comparison.md --json reports/retriever_comparison.json

Run tests:

PYTHONPATH=src python -m unittest discover -s tests

Install as a local CLI:

pip install -e .
rag-eval --corpus corpus evaluate data/eval_cases.jsonl

Evaluation Case Format

{
  "id": "rent_notice",
  "question": "What details matter when checking whether a rent increase notice is valid?",
  "expected_sources": ["tenant_guide/rent_notices.md"],
  "answer": "The answer with citations like [tenant_guide/rent_notices.md#chunk-0].",
  "required_facts": ["check the written notice date"]
}

See docs/EVAL_SCHEMA.md for the full schema.

See docs/RETRIEVER_COMPARISON.md for comparison mode.

What The Report Shows

  • Retrieval recall@k
  • Citation precision and recall
  • Required-fact coverage
  • Grounded answer rate
  • Failure-tag counts
  • Top retrieved evidence per case
  • Retriever comparison deltas
  • Top-source ranking differences

Sample reports:

Why This Matters

A RAG app can look good in a demo while still retrieving weak evidence, citing the wrong document, or answering without support. This workbench makes those failures visible with a small, repeatable evaluation loop.

Repository Quality

  • ROADMAP.md explains how the project can mature toward a production-style RAG evaluation workflow.
  • CONTRIBUTING.md documents local setup and contribution expectations.
  • SECURITY.md documents data-handling rules for corpora, eval cases, and credentials.
  • .github/workflows/ci.yml runs tests and sample CLI checks on GitHub Actions.

Next Extensions

  • Add embedding retriever adapters
  • Add FAISS or Chroma vector-store backends
  • Add PDF ingestion
  • Add LLM-as-judge groundedness checks
  • Track prompt and retrieval-regression runs over time

About

RAG evaluation workbench for retrieval recall, citation coverage, groundedness checks, and failure analysis

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages