Thank you for your interest in contributing.
git clone <repo>
cd AgentEvalOps
python -m pip install -e ".[dev]"Requires Python 3.10+.
make test # pytest
make test-cov # pytest with coverage reportOr directly:
pytest
pytest --cov=agentevalops --cov-report=term-missingmake lint # ruff check src/ tests/
make typecheck # mypy src
make check # lint + typecheck + testOr directly:
ruff check src/ tests/
mypy srcAfter touching runtime code, run the full end-to-end smoke:
make smokeThis runs:
agentevalops run --config configs/toy_smoke.yaml --output runs/make-smoke
agentevalops validate-bundle --bundle runs/make-smoke
agentevalops replay --bundle runs/make-smokeand then cleans up runs/make-smoke.
pre-commit install
pre-commit run --all-filesHooks: trailing whitespace, EOF newlines, YAML/TOML syntax, large-file guard, ruff.
- Ruff enforces
E,F,Irules withline-length = 88. - Mypy
strict = true— all functions must be fully typed. - No
# type: ignorewithout a comment explaining why. - Prefer explicit over clever.
AgentEvalOps is a local-first evaluation framework. The following are explicitly out of scope until a design discussion occurs:
- Cloud backends (AWS Bedrock, any managed API)
- Model provider SDKs (OpenAI, Anthropic, LangGraph, Ollama)
- HTTP clients (
requests,httpx,urllib) in runtime code subprocesscalls in runtime code- External benchmark downloads (SWE-bench, HumanEval, etc.)
- FastAPI dashboard or web UI
- Database persistence
- Plugin or extension framework
If your change requires any of these, open a feature request first.
runs/is git-ignored. Do not commit generated run outputs.dist/andbuild/are git-ignored. Do not commit package build outputs.- Run
make cleanto remove caches and build artefacts.
- Add a new scenario key to
ToyBenchmarkAdapterinsrc/agentevalops/benchmarks/toy.py. - Add tasks to the scenario's task list.
- Add a YAML config under
configs/(use an existing config as a template). - Add tests under
tests/test_toy_scenarios.py. - Document the scenario in
README.mdanddocs/local-demo.md.
- Fork and create a branch.
- Run
make checkandmake smoke. - Open a PR using the pull request template.
- Ensure no generated
runs/,dist/, orbuild/files are staged.
The bundle format is declared in BundleWriter and validated by
BundleValidator. Both classes share a BUNDLE_FORMAT_VERSION constant.
Before changing the format:
- Check whether existing fixture bundles in
tests/need updating. - Bump
BUNDLE_FORMAT_VERSIONif the change is not backward-compatible. - Update
BundleValidator's required-file list if files are added or removed. - Update
README.mdbundle anatomy table anddocs/local-demo.mdif the file set changes. - Add or update tests in
tests/test_bundle_manifest.pyandtests/test_bundle_validator.py.
- New CLI command: update
README.mdQuickstart,docs/local-demo.md, and add a test intests/test_cli.py. - New toy scenario: follow the steps in "Adding a new toy scenario" above;
also update
README.mdscenario reference table anddocs/local-demo.md. - New config key: update
README.mdconfig reference,CONTRIBUTING.mdscope note if relevant, and add coverage intests/test_config_loader.py. - Changed public API (schemas, protocols): mypy strict will catch callers; update affected tests and docs.