Skip to content

Latest commit

 

History

History
655 lines (488 loc) · 44.1 KB

File metadata and controls

655 lines (488 loc) · 44.1 KB

ARCHITECTURE.md — AgentEvalOps Platform

Status: Draft v0.1 · Owner: (your name) · Last updated: 2026-07-05 This document specifies how AgentEvalOps is built. The what and why live in DESIGN.md. When this document and DESIGN.md disagree, DESIGN.md wins and this one is wrong.


1. Overview

AgentEvalOps is a Python package (agentevalops) with nine core protocols, a set of pluggable implementations behind each, a portable result-bundle format, and a planned AWS cloud backend. Local execution is the default; AWS is an optional extension behind the CloudBackend protocol.

The architecture has three layers:

The core layer (src/agentevalops/core/) defines protocols, schemas, and errors. It has no dependencies on any cloud provider, model provider, or agent framework — it is plain Python @dataclass models, not Pydantic (see core/schemas.py's own docstring: "no Pydantic, no ORM," to keep this layer dependency-free). OpenTelemetry wiring is planned (ROADMAP W9) but not yet implemented; nothing in core/ depends on it today.

The adapter layer contains concrete implementations of the core protocols: agents/ (AgentRunner), benchmarks/ (BenchmarkAdapter), evaluators/ (Evaluator), scorers/ (Scorer), stores/ (TraceStore), reports/ (ReportGenerator), policy/ (PolicyChecker). cloud/ (CloudBackend) is planned but not yet created — nothing submits cloud jobs yet, so there is no real CloudBackend implementation, only a test-only mock. observability/ exists as of the W9 scaffold (see §8) but instruments nothing yet. Adapters depend on the core; the core never depends on adapters. New benchmarks, agents, clouds, and evaluators are added by writing new adapters.

Three more packages sit alongside the adapters and don't map to a single protocol: orchestration/ (the eval loop; wires all the protocols together for a run), bundles/ (result-bundle read/write/validate), and replay/ (bundle verification — narrower than §7 describes; see the callout there). registry.py and config/ are top-level, not nested under core/, despite earlier drafts of this document describing them that way (see §4).

The edge layer (src/agentevalops/cli.py) is what users interact with — a CLI that composes adapters. The edge layer never reaches into adapter internals. A FastAPI service and a dashboard are out of scope for v0.1 and explicitly deferred.

This three-layer separation is the load-bearing architectural commitment. Everything else in this document is detail.


2. The nine protocols

These are defined in src/agentevalops/core/protocols.py. They are Python Protocol classes (PEP 544), not abstract base classes, because we want structural subtyping — an implementation does not need to inherit from the protocol, only match its shape. This makes it trivial to wrap third-party agent frameworks without modifying them.

All protocols are typed in strict mode. mypy --strict passes on all of src/ as a CI gate (mypy src, see pyproject.toml's [tool.mypy]). Earlier drafts of this document also named pyright --strict; that tool was never adopted — mypy alone is the actual, running gate.

2.1 AgentRunner

Executes an agent against a prepared environment and emits a trace.

from typing import Protocol, AsyncIterator
from agentevalops.core.schemas import (
    AgentConfig, EnvHandle, TaskSpec, TraceEvent, AgentResult, ResourceLimits,
)

class AgentRunner(Protocol):
    """Executes an agent against a single task in a prepared environment.

    Implementations are responsible for invoking the agent (LangGraph, Bedrock,
    OpenAI Agents SDK, replay from a recorded bundle, etc.) and emitting trace
    events as they occur. The runner does NOT score the result; that is the
    Evaluator's job.

    Implementations MUST:
      - Emit at least one TraceEvent per step the agent takes.
      - Respect resource_limits.max_tokens, .max_wall_seconds, .max_cost_usd.
      - Stop the agent and emit a terminal TraceEvent if any limit is reached.
      - Never raise on agent-level failure; capture it in AgentResult.error.
      - Raise only on infrastructure failure (sandbox died, model API 5xx).
    """

    runner_id: str  # stable identifier, e.g. "langgraph-claude-4-7"

    async def run(
        self,
        task: TaskSpec,
        env: EnvHandle,
        agent_config: AgentConfig,
        resource_limits: ResourceLimits,
    ) -> AsyncIterator[TraceEvent]:
        """Yield trace events as the agent runs. The final event MUST have
        kind='agent.terminal' and carry the AgentResult in its payload."""
        ...

The AsyncIterator return type is deliberate: it lets the platform stream traces to disk and to OpenTelemetry as they happen, rather than buffering an entire run before persisting anything. For long-running ML-engineering tasks this is the difference between losing a four-hour run on a sandbox crash and losing the last thirty seconds.

2.2 BenchmarkAdapter

Translates a benchmark (SWE-bench, MLE-bench, an internal workflow suite) into the platform's task model.

from collections.abc import Iterable
from typing import Any, Protocol
from agentevalops.core.schemas import AgentOutput, EnvHandle, EvaluationResult, TaskSpec

class BenchmarkAdapter(Protocol):
    """Bridges an external benchmark into the platform.

    One adapter per benchmark family. The adapter owns the mapping from the
    benchmark's native format (SWE-bench instances, MLE-bench competitions,
    Terminal-Bench tasks) into TaskSpecs, AND owns the upstream-blessed
    grading logic. The platform does NOT reimplement benchmark scoring.
    """

    benchmark_id: str  # e.g. "swebench-verified", "mlebench-lite"
    benchmark_version: str  # pinned upstream version

    def list_tasks(
        self, filter_spec: dict[str, Any] | None = None
    ) -> Iterable[TaskSpec]:
        """Enumerate tasks. filter_spec is benchmark-specific (e.g. difficulty,
        repo subset, competition list)."""
        ...

    def grade(self, task: TaskSpec, output: AgentOutput) -> EvaluationResult:
        """Apply the benchmark's official grading. This is deterministic
        scoring only — soft criteria (LLM-judge, trace quality) are separate
        Evaluators, not the adapter's job."""
        ...

    def prepare_environment(self, task: TaskSpec) -> EnvHandle:
        """Set up the sandbox/repo/dataset state the agent will operate on.
        Returns a handle teardown() consumes. On failure, remove any partial
        work created, then raise InfrastructureError (docs/benchmark_adapters.md D5)."""
        ...

    def teardown(self, handle: EnvHandle) -> None:
        """Release resources prepare_environment() set up. MUST be idempotent."""
        ...

Both prepare_environment and teardown are synchronous, not async — matching list_tasks/grade (docs/benchmark_adapters.md D5: a plain git clone, not an async container API; LocalOrchestrator runs them via asyncio.to_thread).

The split between grade (deterministic, benchmark-native) and Evaluator (everything else) is deliberate. SWE-bench's "did the patch pass the hidden test suite?" is a BenchmarkAdapter.grade concern. "Did the agent take 47 unnecessary steps before getting there?" is a TraceQualityEvaluator concern. They run independently and produce independent scores. As of row 15 (D5), no production code calls grade at all — LocalOrchestrator routes scoring through Evaluator.evaluate() exclusively; wiring a real grade() caller is row 16's job, and D5's open questions flag a real design gap it will need to resolve (whether grade needs the workspace EnvHandle before teardown deletes it).

2.3 Evaluator

Scores a completed run on a single dimension. Five evaluator types are first-class; the protocol is uniform across all five.

from typing import Protocol
from agentevalops.core.schemas import (
    TaskSpec, AgentResult, Trace, EvaluatorScore, EvaluatorContext,
)

class Evaluator(Protocol):
    """Scores one dimension of a completed run.

    Five canonical kinds, all behind this same interface:
      - DeterministicEvaluator: unit tests, exact match, state assertions
      - LLMJudgeEvaluator: quality, relevance, policy compliance via a model
        (optional; not required for v0.1)
      - ToolUseEvaluator: tool-call validity, argument correctness
      - StateBasedEvaluator: final filesystem/db/cloud state vs target
      - TraceQualityEvaluator: loops, hallucinated tool outputs, dead-ends

    Evaluators are pure functions of (task, result, trace). They MUST NOT
    mutate the environment. They MAY make outbound LLM calls (for judges)
    but MUST report those calls' cost in the EvaluatorScore.
    """

    evaluator_id: str  # e.g. "swebench-pytest", "claude-judge-relevance"
    evaluator_kind: str  # one of: "deterministic", "llm_judge", "tool_use",
                         #         "state_based", "trace_quality"

    async def evaluate(
        self,
        task: TaskSpec,
        result: AgentResult,
        trace: Trace,
        context: EvaluatorContext,
    ) -> EvaluatorScore:
        """Return a score on this evaluator's dimension. Does not raise on
        agent failure — a failed agent produces a low score, not an exception."""
        ...

A run is graded by all configured evaluators in parallel, and their scores are reported separately. The platform never collapses them into a single composite score by default — that decision belongs to the consumer of the bundle.

2.4 Scorer

Aggregates per-task scores into per-run summaries. Distinct from Evaluator because it operates over collections.

from typing import Protocol, Sequence
from agentevalops.core.schemas import EvaluatorScore, RunSummary

class Scorer(Protocol):
    """Aggregates scores across tasks within a run.

    Examples: pass@1, pass@k, mean cost per task, p95 latency, regression
    delta vs a baseline run.
    """

    scorer_id: str

    def summarize(
        self,
        scores: Sequence[EvaluatorScore],
        baseline: RunSummary | None = None,
    ) -> RunSummary:
        """Aggregate. If baseline is provided, include comparison fields."""
        ...

2.5 TraceStore

Persists trace events. The platform writes traces locally to JSONL by default. In future AWS-enabled runs, traces may additionally be exported to CloudWatch via OpenTelemetry; the local JSONL file always exists in the bundle regardless.

Note on TraceEvent.kind vs. OTel span names. TraceEvent.kind is a semantic event type from a closed enum (e.g. agent.tool_call, agent.final_answer). OpenTelemetry span names (e.g. agent.execute, benchmark.prepare) are observability labels used for dashboards and traces in external systems. They are parallel concepts, not the same field.

from typing import Protocol, AsyncIterator
from agentevalops.core.schemas import TraceEvent, RunId

class TraceStore(Protocol):
    """Persists and retrieves trace events."""

    async def append(self, run_id: RunId, event: TraceEvent) -> None:
        """Append a single event. MUST be safe to call concurrently for the
        same run_id."""
        ...

    async def stream(self, run_id: RunId) -> AsyncIterator[TraceEvent]:
        """Yield all events for a run in append order."""
        ...

    async def finalize(self, run_id: RunId) -> None:
        """Close the trace. Subsequent appends MUST raise. Idempotent."""
        ...

2.6 ArtifactStore

Persists files the agent produced (patches, model checkpoints, generated reports).

Status as of v0.1.0: defined, mock-only. No production implementation exists — nothing in the codebase produces binary artifacts yet (that lands with SWE-bench patch handling). The real signature (core/protocols.py) is simpler than earlier drafts of this section described — plain bytes/str, not BinaryIO/ArtifactRef (neither of those types exists):

from typing import Protocol
from agentevalops.core.types import RunId

class ArtifactStore(Protocol):
    """Stores binary artifacts produced during a run (patches, reports, etc.)."""

    async def put(self, run_id: RunId, path: str, content: bytes) -> str:
        """Store an artifact.  Returns an opaque reference string."""
        ...

    async def get(self, ref: str) -> bytes:
        """Retrieve an artifact by its reference string."""
        ...

    async def list(self, run_id: RunId) -> list[str]:
        """Enumerate artifact reference strings for a run."""
        ...

ArtifactStore is separate from TraceStore because artifact lifecycle is different — traces are append-only event logs, artifacts are immutable blobs that may be large (multi-GB ML model checkpoints). Local implementation writes to disk; AWS implementation writes to S3 with lifecycle rules.

2.7 ReportGenerator

Produces the human-readable report in the result bundle (report.md).

Status as of v0.1.0 (2026-07-05): real, in production. reports/markdown.py's MarkdownReportGenerator is what BundleWriter._write_report actually calls. The signature differs from earlier drafts of this section in two ways: it takes RunConfig/RunSummary (not a bare ScoreResult — a report needs cost, tokens, per-task results, and the policy verdict, none of which fit in a score aggregate), and it returns a plain str, not a ReportArtifact wrapper type (which was never built — the bundle writer handles the filename/content-type concerns itself):

from typing import Protocol
from agentevalops.core.schemas import RunConfig, RunSummary, TraceEvent

class ReportGenerator(Protocol):
    """Renders human-readable reports from a completed run."""

    report_id: str  # e.g. "markdown-v1"

    async def render(
        self,
        run_config: RunConfig,
        summary: RunSummary,
        trace: list[TraceEvent],
        config_name: str = "",
    ) -> str:
        """Produce a report.  Returns the report content as a string."""
        ...

A future ReportGenerator implementation that uses an LLM (e.g. a failure-analysis generator) would capture its cost the same way evaluator costs are captured — no such implementation exists yet.

2.8 CloudBackend

The cloud-neutrality boundary. All cloud-specific operations go through this protocol; everything else in the codebase is cloud-agnostic.

Status as of v0.1.0: defined, mock-only — no LocalBackend either. Despite what earlier drafts of this section said, there is no LocalBackend implementation today, not even for local execution: LocalOrchestrator runs tasks directly against AgentRunner/BenchmarkAdapter/etc. without going through any CloudBackend.submit_job-style call at all. backend_id on RunConfig is carried through into bundle metadata and the markdown report, but nothing dispatches on it — grep -rn "backend_id" src/ shows exactly those two write/print sites and no branch. The five job-related schema types this section's code block shows below (JobSpec, JobHandle, JobStatus, ContainerSpec, SecretRef) do not exist anywhere in src/ or tests/; the real protocol signatures (core/protocols.py) use plain dict[str, Any]/str throughout. The AwsBackend (and a real LocalBackend to match it) are planned for the AWS phase (ROADMAP W17+); until then only a test-only mock satisfies this protocol. See docs/adr/0004-cloud-backend-two-axes.md for the Azure-vs-AWS primitive mapping recorded ahead of AwsBackend's design, and for why "Azure support" splits into two independent axes (a model-provider runner concern, separate from this backend concern).

Note on ResourceLimits vs. PolicySpec. ResourceLimits (on AgentConfig) are runtime controls — the runner enforces them during execution and will stop the agent if a limit is reached. PolicySpec is a post-run compliance check — the PolicyChecker evaluates the completed trace after the fact and produces a verdict. They are separate concerns with separate enforcement points.

from typing import Protocol
from agentevalops.core.schemas import (
    JobSpec, JobHandle, JobStatus, ContainerSpec, SecretRef,
)

class CloudBackend(Protocol):
    """Abstracts the cloud primitives the platform actually needs.

    Deliberately small. The five primitives below are all that distinguishes
    a 'local laptop run' from a 'production AWS run' from a future cloud run.
    Anything cloud-specific beyond these (IAM policies, VPC config, IaC) is
    out of scope for the runtime and lives in infra/.
    """

    backend_id: str  # "local", "aws" (future: "azure", "gcp")

    async def submit_job(self, spec: JobSpec) -> JobHandle:
        """Launch a containerized job. Returns a handle for status/logs/cancel."""
        ...

    async def job_status(self, handle: JobHandle) -> JobStatus:
        """Poll job status. Includes resource usage and cost-to-date."""
        ...

    async def cancel_job(self, handle: JobHandle) -> None:
        """Stop the job. Idempotent."""
        ...

    async def fetch_secret(self, ref: SecretRef) -> str:
        """Read a secret. Implementations MUST NOT log secret values, MUST NOT
        write them to traces, MUST NOT include them in artifacts."""
        ...

    async def export_telemetry(self, run_id: str) -> None:
        """Forward this run's traces and metrics to the cloud's observability
        surface (e.g. CloudWatch for AwsBackend). Local backend is a no-op."""
        ...

The protocol is deliberately small. It is the answer to the question "what does AgentEvalOps actually need from a cloud?" and the answer is: launch containers, read secrets, ship telemetry. Anything more (IAM, VPC, autoscaling policies) belongs in IaC, not in the runtime.

2.9 PolicyChecker

Evaluates a completed run against organizational policy. Distinct from Evaluator because policy is a binary compliance question, not a graded dimension, and because the consequences of policy violation (block publication, alert security team) are different from low scores.

from typing import Protocol
from agentevalops.core.schemas import (
    PolicySpec, Trace, AgentResult, PolicyVerdict,
)

class PolicyChecker(Protocol):
    """Evaluates a run against a PolicySpec.

    Operates post-hoc on the completed trace. Does NOT intercept tool calls
    in flight — runtime safety is a different concern with different latency
    requirements (see DESIGN.md non-goals).
    """

    checker_id: str

    async def check(
        self,
        policy: PolicySpec,
        result: AgentResult,
        trace: Trace,
    ) -> PolicyVerdict:
        """Return PASS / FAIL / WARN with citations to specific trace events."""
        ...

3. Schemas

All schemas are Pydantic v2 models defined in src/agentevalops/core/schemas.py. They are versioned via a top-level schema_version field on the result bundle's metadata.json, and a compatibility-test suite ensures bundles produced by older platform versions remain readable.

The schema set is deliberately flat — TaskSpec, AgentConfig, RunConfig, EnvHandle, TraceEvent, AgentResult, GradeReport, EvaluatorScore, EvaluatorContext, RunSummary, ArtifactRef, ReportArtifact, JobSpec, JobHandle, JobStatus, ContainerSpec, SecretRef, PolicySpec, PolicyVerdict, Trace, RunId, ResourceLimits. Anything more nested becomes hard to evolve.

Correction, verified by grep -n "^class " src/agentevalops/core/schemas.py: seven of the names above are not defined in core/schemas.py, or anywhere in src//tests/ — ArtifactRef, ReportArtifact, JobSpec, JobHandle, JobStatus, ContainerSpec, SecretRef. The first two are ArtifactStore's target types (§2.6 already notes the real signature is plain bytes/str); the last five are CloudBackend's target types (§2.8, and docs/adr/0004-cloud-backend-two-axes.md). Every other name in the list above is real, though sometimes under a different name than shown in this document's protocol code blocks (e.g. the real dataclass is AgentOutput, not AgentResult) — that broader naming drift is out of this correction's scope.

A few schemas worth calling out specifically:

TraceEvent carries run_id, step_index, timestamp, kind (one of a closed enum: agent.plan, agent.tool_call, agent.tool_result, agent.observation, agent.final_answer, agent.terminal, evaluator.score, policy.verdict, cost.tick), payload (kind-specific), cost_delta_usd, tokens_delta, and an OpenTelemetry-compatible span_context. The closed enum is critical — open-ended kind strings make trace analysis a regex problem instead of a typed problem.

AgentResult carries success: bool, final_answer: str | None, artifacts: list[ArtifactRef], error: ErrorInfo | None, total_cost_usd, total_tokens, wall_seconds, and terminated_by (one of completed, limit_tokens, limit_time, limit_cost, infra_failure, agent_error). The terminated_by field is what makes regression analysis tractable — knowing whether a run failed because the agent gave up vs. ran out of budget is more useful than the success boolean alone.

EvaluatorScore carries the evaluator id and kind, a score (float in [0,1] for graded evaluators, bool for deterministic), a confidence (for LLM judges), cost_usd, latency_ms, and citations: list[TraceEventRef] pointing at the specific events the score is based on. Citations are what let consumers of the bundle understand why a score was what it was, not just what the score was.


4. Directory layout

This section describes the actual, as-built v0.1.0 layout (updated 2026-07-05; earlier drafts described an aspirational structure — orchestrator.py/bundle.py/replay.py/registry.py nested under core/, uv as the package manager, separate replay.yml/release.yml workflows, an examples/ tree — none of which is what actually got built. If you're picking this up in a new session, trust this section and the real src/ tree over memory of older drafts.

AgentEvalOps/
├── README.md
├── DESIGN.md                       # what & why
├── ARCHITECTURE.md                 # this file
├── ROADMAP.md
├── CHANGELOG.md                    # actual build history, by WBS number
├── SECURITY.md
├── CONTRIBUTING.md
├── Makefile                        # install, lint, typecheck, test, test-cov, check, smoke
├── pyproject.toml                  # pip/hatchling — no uv, no uv.lock
├── .github/
│   └── workflows/
│       ├── ci.yml                  # lint (ruff) + typecheck (mypy) + test+coverage
│       │                           #   + smoke (run / validate-bundle / replay), matrix 3.10-3.12
│       └── package.yml             # build check
│
├── src/
│   └── agentevalops/
│       ├── __init__.py
│       ├── cli.py                  # entry point: `agentevalops` (run, report, replay,
│       │                           #   validate-bundle, version, doctor)
│       ├── registry.py             # entry-point discovery (top-level, not under core/)
│       │
│       ├── core/                   # LEAF: no dependency on any adapter or engine package
│       │   ├── __init__.py
│       │   ├── protocols.py        # the nine Protocol definitions
│       │   ├── schemas.py          # plain @dataclass models (no Pydantic) incl. RunSummary
│       │   ├── types.py            # RunId/TaskId/AgentId/BackendId newtypes
│       │   └── errors.py
│       │
│       ├── orchestration/
│       │   └── local.py            # LocalOrchestrator: the eval loop, concurrency-bounded
│       │
│       ├── bundles/                # result-bundle read/write/validate (not under core/)
│       │   ├── constants.py
│       │   ├── manifest.py         # SHA-256 checksums, format version
│       │   ├── reader.py           # BundleReader (strict + recovery modes)
│       │   ├── serializers.py
│       │   ├── validator.py        # BundleValidator, 12 integrity checks
│       │   └── writer.py           # BundleWriter
│       │
│       ├── replay/
│       │   └── local.py            # LocalReplayVerifier — bundle *verification*,
│       │                           #   not re-execution; see the callout in §7
│       │
│       ├── config/
│       │   └── loader.py           # YAML RunConfig loader
│       │
│       ├── agents/                 # AgentRunner adapters
│       │   ├── __init__.py
│       │   ├── mock_runner.py      # deterministic runner for tests/demo
│       │   └── langgraph_runner.py # real Anthropic-backed runner (W8); anthropic SDK is an optional `llm` extra, imported lazily
│       │
│       ├── benchmarks/             # BenchmarkAdapter adapters
│       │   ├── __init__.py
│       │   └── toy/
│       │       ├── adapter.py
│       │       ├── scenarios.py
│       │       └── tasks.py
│       │   # Future: swebench/, mlebench/, etc. (not started)
│       │
│       ├── evaluators/             # Evaluator adapters
│       │   ├── __init__.py
│       │   └── deterministic.py
│       │   # Future: llm_judge.py (optional), tool_use.py, etc. (not started)
│       │
│       ├── scorers/                # Scorer adapters
│       │   ├── __init__.py
│       │   └── simple.py
│       │
│       ├── stores/                 # TraceStore adapters
│       │   ├── __init__.py
│       │   ├── memory.py           # InMemoryTraceStore (tests, and CLI when no --output)
│       │   └── local.py            # LocalTraceStore: disk JSONL, flush+fsync per event,
│       │                           #   used by `run --output` (survives a hard kill)
│       │   # ArtifactStore has no implementation yet — nothing produces binary
│       │   # artifacts; that lands with SWE-bench patch handling
│       │
│       ├── reports/                # ReportGenerator adapters
│       │   ├── __init__.py
│       │   └── markdown.py         # render_report() + MarkdownReportGenerator
│       │
│       ├── policy/                 # PolicyChecker adapters
│       │   ├── __init__.py
│       │   └── basic_checker.py    # cost ceiling + tool allow/deny list
│       │
│       └── observability/          # W9 scaffold only — see §8
│           ├── __init__.py
│           └── otel_setup.py       # span-name/attribute constants, NoOpTracer;
│                                    #   configure_tracer() raises NotImplementedError
│
│       # cloud/ does not exist yet (ROADMAP W17+). CloudBackend has no
│       # implementation, not even a LocalBackend — the orchestrator runs
│       # tasks directly, without going through a CloudBackend abstraction.
│
├── configs/                        # toy_smoke, toy_failure, toy_policy_violation,
│                                    #   toy_trace_limit, toy_mixed
│
├── tests/                          # flat, not split into unit/integration/smoke/replay/;
│                                    # one test file per source module, mostly
│
├── docs/
│   ├── local-demo.md
│   └── release-checklist.md
│
└── runs/                           # gitignored — local bundle output only

A few rules about this layout, and their actual enforcement status:

src/agentevalops/core/ imports nothing from any adapter, engine, or bundle package — it's the one hard rule everyone should keep true. Not currently enforced by tooling: no import-linter config exists yet (ROADMAP W1/W13 are both still open). mypy --strict will catch some violations incidentally (an import cycle fails to type-check) but that's not the same as a real layer-boundary lint.

Adapter modules import from the core, optionally from each other within the same adapter family, and from third-party libraries. No adapter module imports from a different adapter family. (Also unenforced today — same caveat as above.)

grep -r "boto3" src/ returns no hits at all today, since cloud/ doesn't exist yet. The intended check ("boto3 only inside cloud/aws/") isn't wired into CI.

Tests are mostly one file per source module, but flat under tests/ rather than split into unit//integration//smoke//replay/ subdirectories.


5. The orchestration loop

The orchestrator (src/agentevalops/core/orchestrator.py) is the only place that knows about all nine protocols simultaneously. It is the conductor; everything else is an instrument.

For a single run, the orchestrator does the following, in this order:

It loads the RunConfig and resolves it into concrete protocol implementations via a registry. It generates a run_id (UUIDv7, so it sorts by time) and a result-bundle directory. It captures the replay_command before executing anything — including the platform version, the git SHA if available, and the resolved config — so that even runs that crash partway through produce a partial bundle that can be inspected.

Correction, verified 2026-08-04 (docs/adr/0003-replay-semantics.md, Q3). run_id is not auto-generated. config/loader.py takes it verbatim from the YAML config's run_id field, falling back to the static default "unnamed-run" when omitted — there is no UUIDv7 generation, or any generation at all, anywhere in this codebase. Implementing the documented behavior would introduce a nondeterministic field into precisely the bundle-comparison work replay_command.txt exists to support, so this is being left as a doc fix, not a future TODO. Separately, replay_command.txt is now real (docs/adr/0003, Q4, implemented 2026-08-04 — see §7) — but BundleWriter.write() generates it alongside the bundle's other content files, not "before executing anything": a run that crashes before any bundle file is written has no replay_command.txt either, the same as every other bundle file (docs/adr/0002's best-effort-finalize discussion covers this failure mode already).

It calls BenchmarkAdapter.list_tasks() to enumerate tasks, applying any filters from the config. For each task in parallel (bounded by the configured concurrency limit), it calls BenchmarkAdapter.prepare_environment(), then AgentRunner.run(), streaming TraceEvents into the TraceStore as they arrive. The AgentRunner is responsible for enforcing ResourceLimits (tokens, wall-clock seconds, cost) during execution — if a limit is reached, the runner stops the agent and emits a terminal TraceEvent with the appropriate terminated_by reason. The orchestrator does not intercept mid-run. When the agent emits a terminal event, the orchestrator captures the AgentResult.

It then scores the run: in the target design this calls BenchmarkAdapter.grade() for the benchmark's official deterministic score and runs all configured Evaluators in parallel against the trace (see the implementation note below for how the current LocalOrchestrator differs). It then runs all configured PolicyCheckers against the completed trace and AgentResult. PolicyCheckers evaluate post-run compliance against a PolicySpec; they do not affect execution. All results are written to scores.json. It calls BenchmarkAdapter.teardown() regardless of outcome.

Note (prepare_environment/teardown failure handling, docs/benchmark_adapters.md D5, implemented row 15). A prepare_environment failure (InfrastructureError/ConfigurationError) never reaches the runner at all — the task gets the same not_evaluated result the runner-infrastructure-failure path produces, and teardown is not called, since no EnvHandle was ever returned for it to act on. A teardown failure, by contrast, is only logged as a warning — it never changes the task's already-computed result, because a cleanup failure after a successful (or already-scored) run is not evidence the run itself failed.

Note (current LocalOrchestrator implementation): The current LocalOrchestrator routes per-task scoring through Evaluator (concretely DeterministicEvaluator) rather than calling BenchmarkAdapter.grade() directly. BenchmarkAdapter.grade() is the benchmark's own upstream-blessed deterministic grading — for example, running an external test-suite — and is intentionally reserved for a future phase when such upstream verification is needed. Evaluator is the platform's flexible scoring layer and is appropriate for the toy benchmark and any scenario where the platform supplies the criteria. When a real benchmark (e.g. SWE-bench) is integrated, both will run: BenchmarkAdapter.grade() for the official pass/fail and Evaluators for additional dimensions.

After all tasks complete, it runs the configured Scorers to produce a RunSummary, then runs the configured ReportGenerators to produce the human-readable artifacts. Finally it finalizes the TraceStore, writes metadata.json, and the bundle is sealed.

The whole loop is async. Concurrency is bounded at the task level by a semaphore whose size comes from RunConfig.max_concurrent_tasks. The default is 1 (run tasks serially) because most failure modes show up with concurrency, and we want the unsurprising default. Production AWS runs use higher values.

If the orchestrator is interrupted (Ctrl-C, OOM, sandbox died), it makes a best effort to finalize the partial bundle so what did complete is still inspectable. Bundles produced this way have metadata.json.status = "interrupted" and are excluded from regression baselines by default.


6. AWS topology (planned — not required for v0.1)

AWS is the planned cloud deployment. The topology described here guides the runtime architecture so the interface design and future IaC stay aligned. Local execution requires none of this infrastructure.

The runtime communicates with AWS exclusively through the CloudBackend protocol. The AwsBackend implementation (planned for a post-v0.1 phase) will use:

  • ECS Fargate for eval runner jobs
  • S3 for result bundles and artifacts
  • DynamoDB as a run index
  • Secrets Manager for credentials
  • CloudWatch via ADOT for observability

Specific choices and the reasons for them:

ECS Fargate over Lambda for the eval runner. Eval runs commonly exceed Lambda's 15-minute ceiling. Fargate also gives a real Linux environment for sandboxed code execution.

S3 as the artifact and bundle store. Bundles benefit from S3's lifecycle rules. Bundles are stored at s3://{bucket}/runs/{run_id}/ so a run is a single prefix and IAM policies can be scoped to a single run.

DynamoDB as the run index, not the bundle store. DynamoDB holds the metadata that needs to be queried (run id, agent id, benchmark, timestamp, summary stats, S3 pointer). The bundles themselves live in S3.

OpenTelemetry → CloudWatch via ADOT. Spans become CloudWatch Logs entries; selected span attributes (cost, latency, tokens) become CloudWatch Metrics via the embedded metric format.

Secrets Manager, never task-definition env vars. API keys are fetched at runtime by the runner via CloudBackend.fetch_secret(). Task definitions reference the secret ARN, not the secret value.

Bedrock AgentCore is future work beyond the initial AWS backend phase and is not part of the AwsBackend planned implementation. It will be added as a separate AgentRunner adapter when relevant.

Azure and GCP as CloudBackend implementations are future roadmap items. The CloudBackend protocol is the extension point; they are not active development and have no placeholder modules in the current codebase. Azure specifically splits into two independent axes that should not be conflated — see docs/adr/0004-cloud-backend-two-axes.md for the full split, the primitive-mapping table (AWS's planned ECS/S3/DynamoDB/Secrets Manager/CloudWatch topology above, against Azure's Container Apps Jobs/Blob Storage/Table Storage/Key Vault/Azure Monitor equivalents), and the asymmetries that constrain AwsBackend's eventual design.


7. Replay determinism

Status as of 2026-08-04 (docs/adr/0003-replay-semantics.md, ROADMAP row 12b): the structural half is real and has a real CLI surface; this section's reproducibility story is still target design, not what's built. agentevalops validate-bundle is now the single structural entry point (bundles/validator.py's validate_bundle()): required-file presence (version-aware — a bundle declaring format "0.4" must have replay_command.txt, older bundles are not penalized for lacking it), manifest checksums, and cross-file consistency (replay/local.py's LocalReplayVerifier, now an internal helper validate_bundle() calls and folds findings into, not a class exposed with its own summary type). agentevalops replay still exists but is a deprecated alias delegating to the identical code path, with a stderr deprecation notice — not removed, not renamed to something else. replay_command.txt does now exist, written by BundleWriter alongside every other bundle file: the reconstructable invocation, git_sha/git_dirty, platform_version, run_id, the resolved model id, and (azure-openai runs only) the Azure resource name — human-readable text, not JSON, since it is read by a person deciding whether a bundle is still comparable, not parsed by a machine. Still not built: any ReplayRunner class, behavioral (recorded-output-substitution) comparison of any kind, and tests/replay/fixtures/ — tests/test_replay.py/tests/test_replay_command.py build bundles on the fly with tmp_path, they don't replay a checked-in recording. The reproducibility story below (recorded-output substitution, exact tool-call match, the MATCH/DIFFER/CODE_CHANGED outcome vocabulary) remains unbuilt; closing that gap is real, unstarted work (W12c/W29), not a renaming exercise — validate-bundle performing genuine structural checks doesn't change that behavioral replay is a separate, larger, unstarted feature.

Replay is a tested property of the platform, not a hope. The CI pipeline includes a job that takes a checked-in bundle from tests/replay/fixtures/, runs replay_command.txt, and asserts the new bundle matches the original on a defined schema:

The new bundle must produce the same AgentResult.success, the same set of tool calls in the same order, the same terminated_by reason, and total_cost_usd within ±5% (tolerating provider-side pricing drift). It need not produce identical model outputs token-for-token — that would require provider cooperation we don't have — but the behavioral trace must match.

This is achieved by capturing, in the trace, every input that determined behavior: the resolved prompt, the tool definitions, the temperature and seed (where supported), the model version, the random state at each sampling point. The ReplayRunner consumes a recorded bundle and substitutes recorded model outputs for live calls, which produces an exact match against the original behavior and is the strongest form of replay.

Live replay (re-running with the real model) is supported but not required for the regression test, because provider-side non-determinism is outside our control. Live replay is what users do to test "does this still work against the new model version?"; recorded replay is what CI does to test "did the platform itself break?"


8. Observability

Status as of 2026-08-02: scaffold only, ~10% of this section. observability/otel_setup.py exists and matches the span-name/attribute-key constants below (eval.run, agent.tool_call, cost_usd, etc. — see tests/test_observability.py's constants-match-spec tests). It also provides a NoOpTracer/NoOpSpan pair so future instrumented call sites are safe whether or not the otel extra is installed. That is the entire built surface. Not built: a real TracerProvider, a span exporter, any instrumentation call site (verified — grep for otel_setup/get_tracer/configure_tracer outside observability/ itself returns nothing; no orchestrator, runner, or evaluator emits a span), and the spans.jsonl file described below. get_tracer() always returns the no-op tracer; configure_tracer() raises NotImplementedError — both are explicit TODO(W9) stubs, not bugs. The only structured run data that exists today is traces.jsonl (application-level TraceEvents written by the orchestrator), which is a separate file and a separate data model from the OpenTelemetry spans this section describes — see the note in §2.5.

The platform emits OpenTelemetry spans for every step of every run. Span hierarchy:

eval.run                        (root span, run_id)
├── eval.task                   (one per task)
│   ├── benchmark.prepare
│   ├── agent.execute           (the AgentRunner's span)
│   │   ├── agent.plan
│   │   ├── agent.tool_call     (one per tool call)
│   │   └── agent.final_answer
│   ├── benchmark.grade
│   ├── evaluator.deterministic
│   ├── evaluator.llm_judge
│   ├── evaluator.tool_use
│   ├── evaluator.state_based
│   ├── evaluator.trace_quality
│   ├── policy.check
│   └── benchmark.teardown
├── scorer.summarize
└── report.render

Span attributes carry run_id, task_id, agent_id, model_id, cost_usd, tokens_in, tokens_out, latency_ms. Cost and tokens are emitted as both span attributes and CloudWatch Metrics (via the embedded metric format), so dashboards can graph them without parsing logs.

Local runs export to spans.jsonl in the bundle — a distinct file from traces.jsonl (the application-level TraceEvent log, which exists today; see §2.5). AWS runs export to CloudWatch via ADOT, with spans.jsonl still produced inside the bundle for portability — the bundle is the contract, the cloud surface is augmentation.


9. Configuration and the registry

RunConfig is YAML, loaded by config/loader.py. String references in the config (e.g. runner: mock) are resolved against a registry defined in src/agentevalops/registry.py — top-level, not under core/ (an earlier draft of this document put it there). Built-in adapters register themselves at import time; third-party plugins can publish additional benchmark adapters under the agentevalops.adapters entry-point group, discovered via importlib.metadata.entry_points.

Users can register their own adapters by installing a package that exposes the agentevalops.adapters entry point group in its pyproject.toml. The platform discovers third-party adapters via importlib.metadata.entry_points() at startup and adds them to the registry. This is how external users contribute new benchmarks and runners without forking the platform.

A sample v0.1 RunConfig (local execution, toy benchmark):

run_name: toy-mock-agent-local
benchmark:
  id: toy
agent:
  runner: mock
  resource_limits:
    max_tokens: 10_000
    max_wall_seconds: 60
    max_cost_usd: 0.0
evaluators:
  - id: deterministic-pytest     # deterministic (required for v0.1)
policy:
  spec_id: cost-ceiling-default
backend:
  id: local
storage:
  bundles: ./runs/
  trace_store: local
max_concurrent_tasks: 1

In a future phase, the backend.id can be changed to aws, runner to a real agent runner, and benchmark.id to a real benchmark — no orchestrator changes needed.

Every field maps to a protocol implementation registered in the system. Adding a new runner is one entry in the registry plus a YAML field, no orchestrator changes.


10. What this document does not specify

This document specifies interfaces, layout, topology, and the orchestration loop. It deliberately does not specify:

The exact implementation of any one adapter — those live in their respective modules and are documented in docs/benchmark_adapters.md, docs/writing_an_evaluator.md, etc. The IaC implementation details — those live in infra/aws/. The CI pipeline configuration beyond the rules it must enforce — that is in .github/workflows/.

If a future change touches interfaces, layout, or the orchestration contract, it requires an update to this document. If it only touches an adapter's internals, it does not.