A standard-library workbench for evaluating retrieval-augmented generation outputs.
This project focuses on the pieces that usually get skipped in RAG demos: retrieval quality, citation coverage, required-fact coverage, groundedness checks, failure tags, and reports that explain what went wrong.
- Markdown corpus ingestion
- Token-window chunking with overlap
- BM25-style lexical retrieval
- Keyword-coverage and hybrid retrieval strategies
- Retriever comparison reports with metric deltas
- JSONL evaluation dataset format
- Retrieval recall@k against expected sources
- Citation precision and citation recall
- Required-fact coverage checks
- Groundedness and failure-mode tagging
- Markdown and JSON report generation
- Unit-tested CLI workflow with no required API keys
- GitHub Actions CI workflow for tests, validation, and sample report generation
.
|-- corpus/
| `-- tenant_guide/
| |-- access.md
| |-- deposits.md
| |-- rent_notices.md
| `-- repairs.md
|-- data/
| `-- eval_cases.jsonl
|-- docs/
| |-- EVAL_SCHEMA.md
| `-- RETRIEVER_COMPARISON.md
|-- reports/
| |-- retriever_comparison.json
| |-- retriever_comparison.md
| |-- sample_report.json
| `-- sample_report.md
|-- src/
| `-- rag_eval_workbench/
| |-- cli.py
| |-- dataset.py
| |-- evaluation.py
| |-- ingest.py
| |-- models.py
| |-- reporting.py
| |-- retrieval.py
| `-- text.py
|-- tests/
| `-- test_rag_eval_workbench.py
|-- pyproject.toml
|-- README.md
`-- requirements.txt
This project uses only the Python standard library.
Validate the corpus and eval cases:
PYTHONPATH=src python -m rag_eval_workbench.cli --corpus corpus validate data/eval_cases.jsonlInspect retrieval for one query:
PYTHONPATH=src python -m rag_eval_workbench.cli --corpus corpus retrieve "What should I do about mold near a window?" --top-k 3Inspect retrieval with the hybrid strategy:
PYTHONPATH=src python -m rag_eval_workbench.cli --corpus corpus retrieve "What should I do about mold near a window?" --retriever hybrid --top-k 3Generate reports:
PYTHONPATH=src python -m rag_eval_workbench.cli --corpus corpus evaluate data/eval_cases.jsonl --markdown reports/sample_report.md --json reports/sample_report.jsonCompare retrievers:
PYTHONPATH=src python -m rag_eval_workbench.cli --corpus corpus compare data/eval_cases.jsonl --retrievers bm25 keyword hybrid --markdown reports/retriever_comparison.md --json reports/retriever_comparison.jsonRun tests:
PYTHONPATH=src python -m unittest discover -s testsInstall as a local CLI:
pip install -e .
rag-eval --corpus corpus evaluate data/eval_cases.jsonl{
"id": "rent_notice",
"question": "What details matter when checking whether a rent increase notice is valid?",
"expected_sources": ["tenant_guide/rent_notices.md"],
"answer": "The answer with citations like [tenant_guide/rent_notices.md#chunk-0].",
"required_facts": ["check the written notice date"]
}See docs/EVAL_SCHEMA.md for the full schema.
See docs/RETRIEVER_COMPARISON.md for comparison mode.
- Retrieval recall@k
- Citation precision and recall
- Required-fact coverage
- Grounded answer rate
- Failure-tag counts
- Top retrieved evidence per case
- Retriever comparison deltas
- Top-source ranking differences
Sample reports:
- reports/sample_report.md
- reports/sample_report.json
- reports/retriever_comparison.md
- reports/retriever_comparison.json
A RAG app can look good in a demo while still retrieving weak evidence, citing the wrong document, or answering without support. This workbench makes those failures visible with a small, repeatable evaluation loop.
- ROADMAP.md explains how the project can mature toward a production-style RAG evaluation workflow.
- CONTRIBUTING.md documents local setup and contribution expectations.
- SECURITY.md documents data-handling rules for corpora, eval cases, and credentials.
.github/workflows/ci.ymlruns tests and sample CLI checks on GitHub Actions.
- Add embedding retriever adapters
- Add FAISS or Chroma vector-store backends
- Add PDF ingestion
- Add LLM-as-judge groundedness checks
- Track prompt and retrieval-regression runs over time