Chopin's background system runs durable, versioned work outside the shared Planner conversation. A job may execute ordinary code or open one or more isolated harness worker sessions. The code calls the durable unit a background job; this document uses worker for the disposable execution session. A worker is not a child Planner turn, coding agent, or runtime plugin.
Important
The registry is a closed, code-owned allowlist. Persisted rows name a registered job type and version; they never contain executable code. Adding a job requires a source change and deployment.
This guide covers the generic job framework, how to register a new definition, and the worker boundaries used by generated document descriptions and research requests. See Hosted agent for the Planner conversation and Storage for the broader durable model.
| Term | Meaning |
|---|---|
| Definition | The code-owned JobDefinition for one type@version. |
| Job | One durable request, including normalized input and lifecycle state. |
| Target | The semantic resource a job updates, such as one document or one internal research attempt. |
| Target generation | The newest request for a target. A newer generation supersedes older active work. |
| Attempt | One job claim. Credential or capacity failure may prevent its executor from starting. |
| Failure | An attempt that consumes retry budget. failures, not attempts, is compared with maxAttempts. |
| Claim generation | A fencing counter that prevents an expired or cancelled worker from publishing late output. |
| Artifact | The immutable, validated JSON result written only when a job completes. |
| Worker | A disposable execution context, optionally backed by an isolated harness session. |
| Background-job channel revision | The invalidation counter for job mutations in one channel, separate from each job's revision. |
Keep these counters separate from the Yjs epoch, document sequence, plan revision, storage revision, research request revision, and implementation graph counters.
flowchart LR
E[Scheduler, user route, or domain service] --> S[JobService]
S -->|validate, normalize, fingerprint| P[(PostgreSQL)]
P -->|fenced claim| R[JobRunner]
R --> C[Credential resolver]
C --> X[Registered executor]
X -->|optional disposable session| A[Harness worker]
X -->|progress| P
X -->|candidate artifact| S
S -->|validate and settle| P
P --> I[job:changed invalidation]
P --> D[Optional domain reconciliation]
The main boundaries are:
JobRegistrystores every supported definition version and selects the highest version for new work.JobServiceowns origin checks, canonical JSON, byte limits, fingerprints, enqueueing, control mutations, artifact validation, and publication hooks.BackgroundJobStoreowns durable state transitions, target generations, claims, progress, and artifacts.JobRunnerpolls and wakes for work, resolves credentials, enforces concurrency, heartbeats claims, executes definitions, and handles retries and shutdown.- A definition's executor owns the actual work and any worker-specific capability boundary.
Product code should enqueue through JobService, not call
storage.jobs.enqueue() directly. The raw storage port enforces persistence
invariants but does not know registry versions, origins, codecs, or byte limits.
- Choose a stable lowercase type and start at version 1.
- Define strict, deterministic input and artifact codecs.
- Select only the origins an actual caller needs; enforce authorization at the route, socket, scheduler, or tool boundary.
- Select
active-plannerornonedeliberately. For a user-triggeredactive-plannerjob, establish or validate ownership before enqueueing. - Set timeout, failure budget, aggregate AI-credit metadata, and byte limits.
- Declare fixed progress stages without private or model-generated text.
- Implement execution with cancellation, deadlines, and bounded cleanup.
- For model work, use fresh audited workers and structured result tools.
- Add a publication hook only when completion needs a freshness gate.
- Register every supported version in
apps/server/src/main.ts. - Add an explicit scheduler, route, domain service, or Planner tool that
enqueues through the matching
JobService.enqueue*()method. - Choose stable target and idempotency keys; do not pre-prefix the target.
- Add idempotent domain reconciliation if completion creates related state.
- Add explicit browser or Planner projection only when required.
- For new domain persistence, add and register a migration, update both adapters, and extend the shared storage contract.
- Test codecs, origins, authorization, retries, cancellation, credentials, publication, storage contracts, and every exposed UI path.
A registered definition has this shape:
type JobDefinition<Input extends JsonValue, Artifact extends JsonValue> = {
type: string;
version: number;
label: string;
description: string;
origins: readonly ("scheduler" | "planner" | "user")[];
credential: "active-planner" | "none";
limits: {
timeoutMs: number;
maxAttempts: number;
maxAiCredits: number;
maxInputBytes: number;
maxArtifactBytes: number;
};
progress?: Readonly<Record<string, string>>;
input: { parse(value: JsonValue): Input };
artifact: { parse(value: JsonValue): Artifact };
execute(execution: JobExecution<Input>): Promise<Artifact>;
publish?(publication: JobPublication): Promise<void>;
};JobRegistry validates definitions when the process starts:
- types and progress stages use lowercase hyphenated identifiers;
- versions and limits are positive safe integers;
- labels and descriptions are nonblank;
- at least one origin is declared, with distinct members of
scheduler,planner, anduser; - credential mode is
active-plannerornone; - progress, when present, has 1 through 32 code-owned labels, each at most 160 characters;
- only one definition may claim a particular
type@version.
label and description are catalog metadata. They do not create or label a
browser surface automatically.
Input and artifact codecs are security and compatibility boundaries. They should:
- accept exact fields and reject unknown ones;
- normalize deterministically;
- enforce semantic, collection, and string bounds;
- return plain JSON only;
- reject
undefined, dates, class instances, accessors, sparse arrays, non-finite numbers, and cyclic values; - validate provenance and relationships rather than trusting model output.
JobService canonicalizes JSON before and after the input codec. Object keys are
sorted, nesting is bounded, UTF-8 bytes are measured, and the normalized form is
used in the idempotency fingerprint. Settlement applies the same process to the
artifact codec.
Changing a persisted contract normally requires a new definition version. Keep old versions registered while durable pending, paused, or running rows can still reference them. New enqueue requests use the highest registered version; the runner and settlement always resolve the exact persisted version.
The following skeleton shows a non-model job. Real codecs should usually be more specific than these examples.
import { JobExecutionError } from "./registry";
import type { JsonValue } from "../storage/model";
import type { JobDefinition } from "./registry";
type ExampleInput = {
resourceId: string;
};
type ExampleArtifact = {
resourceId: string;
result: string;
};
async function analyze(
resourceId: string,
signal: AbortSignal,
deadline: Date,
): Promise<string> {
if (signal.aborted || Date.now() >= deadline.getTime()) {
throw new Error("analysis stopped");
}
return `Analyzed ${resourceId}`;
}
function record(value: JsonValue): Record<string, JsonValue> {
if (!value || typeof value !== "object" || Array.isArray(value)) {
throw new Error("expected object");
}
return value;
}
function input(value: JsonValue): ExampleInput {
let item = record(value);
let resourceId = typeof item.resourceId === "string" ? item.resourceId.trim() : "";
if (Object.keys(item).length !== 1 || !resourceId || resourceId.length > 128) {
throw new Error("invalid example input");
}
return { resourceId };
}
function artifact(value: JsonValue): ExampleArtifact {
let item = record(value);
if (
Object.keys(item).length !== 2
|| typeof item.resourceId !== "string"
|| typeof item.result !== "string"
|| item.result.length > 4_000
) throw new Error("invalid example artifact");
return { resourceId: item.resourceId, result: item.result };
}
export function exampleDefinition(): JobDefinition<ExampleInput, ExampleArtifact> {
return {
type: "example-analysis",
version: 1,
label: "Example analysis",
description: "Analyzes one example resource.",
origins: ["user"],
credential: "none",
progress: { analyze: "Analyzing resource" },
limits: {
timeoutMs: 60_000,
maxAttempts: 2,
maxAiCredits: 1,
maxInputBytes: 8 * 1024,
maxArtifactBytes: 16 * 1024,
},
input: { parse: input },
artifact: { parse: artifact },
async execute(execution) {
await execution.progress("analyze", "started");
if (execution.signal.aborted) {
throw new JobExecutionError("analysis-cancelled");
}
let result = await analyze(
execution.input.resourceId,
execution.signal,
execution.deadline,
);
await execution.progress("analyze", "completed");
return artifact({ resourceId: execution.input.resourceId, result });
},
};
}Even a credential-free definition currently needs positive maxAiCredits.
That field is catalog metadata, not proof that AI is used.
Definitions are assembled explicitly in apps/server/src/main.ts:
let definitions: JobDefinition[] = [];
if (config.backgroundJobs) {
definitions.push(exampleDefinition());
}
let jobRegistry = new JobRegistry(definitions);There is no directory scan or plugin discovery. Register every version that may still have durable work:
definitions.push(exampleDefinitionV1(), exampleDefinitionV2());Registration alone does not schedule work or create an API. Add an explicit enqueue path using the intended origin method:
let result = await jobs.enqueueUser({
channelId,
type: "example-analysis",
targetKey: resourceId,
idempotencyKey: `example-analysis:${requestId}`,
input: { resourceId },
});Choose the origin method deliberately:
| Method | Intended caller |
|---|---|
enqueueScheduler() |
A code-owned scheduler or coordinator. |
enqueuePlanner() |
A Planner tool or Planner-origin domain action. |
enqueueUser() |
An explicit authenticated user action. |
Callers do not supply an origin field. The service stamps it and checks that the definition permits it.
Origins record how work entered the system; they are not authentication or
authorization. A browser route must still authenticate the session, enforce
Origin, admission, repository node identity, and role. Planner tools and
schedulers must enforce their own trusted ingress before calling JobService.
The caller supplies the semantic target without the job type prefix.
JobService changes resource-id into
example-analysis:resource-id. Supplying an already-prefixed target would
produce example-analysis:example-analysis:resource-id.
An idempotency key is unique within a channel. Reusing it with the same
fingerprint returns the original job with repeated: true; reusing it with
different normalized input, origin, target, version, or availability conflicts.
A new idempotency key for the same target increments the target generation and supersedes older pending, paused, or running generations. Older terminal rows remain durable, but consumers must reject artifacts whose target generation is no longer current.
Use request identity for idempotency and resource identity for the target. Do not generate both from the current time.
pending -> running
pending/running -> paused
paused -> pending
running -> pending retry or process recovery
running -> completed
running -> failed
pending/paused/running -> cancelled
pending/paused/running -> superseded
The lifecycle is:
JobServiceresolves the current definition, validates origin and input, computes the fingerprint, and enqueues under the writer fence.- Storage advances the target generation, supersedes older active work, and
inserts a pending row before
job:changedis published. JobRunnerpolls globally and wakes immediately for local pending changes.- Storage claims current pending work with
FOR UPDATE ... SKIP LOCKED, or reclaims an expired running claim. - The runner resolves the exact persisted definition version and parses the persisted input again.
- The runner resolves any required credential, stores only token-free binding metadata, and applies global and per-owner concurrency.
- The executor receives cloned input, credential, abort signal, deadline, and progress callback.
JobServicevalidates the candidate artifact and runs any publication hook.- Fenced storage atomically marks the job completed and inserts its immutable artifact.
- Generic and domain-specific invalidations happen only after persistence.
The default runner uses global concurrency 2, per-owner concurrency 1, a 30-second claim TTL, a 10-second heartbeat, one-second polling, exponential retry up to five minutes, and a 10-second shutdown grace period.
Every claimed mutation must match the channel, job ID, process claim owner, claim generation, unexpired claim, and current target generation. Cancellation, supersession, owner revocation, timeout, or shutdown clears or invalidates that claim. A worker may continue briefly if it ignores cancellation, but its late progress and artifact cannot commit.
claim_owner identifies the runner process, not the Planner owner or an
authorization principal.
attempts increments on every claim. failures increments only when a retry
consumes failure budget. Despite its name, JobLimits.maxAttempts is currently
compared with failures.
Credential rotation, owner capacity, shutdown, and similar recovery paths may
requeue without increasing failures. A JobExecutionError supplies a safe,
code-owned durable reason; an unclassified executor failure becomes
attempt-error.
Definitions must honor both execution.signal and execution.deadline, and
must clean up disposable resources in finally. Fencing protects persistence;
it does not stop leaked external work.
Progress stages must be declared by the definition. The executor can append
started and completed. The runner makes a best-effort interrupted append
while it still owns the claim; cancellation, supersession, or claim loss may
leave no interruption entry. Storage retains at most 32 entries.
Stage IDs, labels, and reason codes are durable and browser-visible. Structured diagnostics are bounded but log-only. Neither category may contain:
- prompts or model prose;
- query text or private document content;
- source URLs or provider payloads;
- credentials, paths, or raw errors.
Use JobExecutionError for a bounded durable machine reason and optional
operator diagnostic. It accepts at most 16 boolean, nonnegative integer, or
token-like string diagnostic fields. Unknown browser reasons are deliberately
rendered as a generic interruption.
Executors can also report bounded operational snapshots through
execution.diagnostic(). The runner retains a validated copy for the current
attempt and reports it on timeout as well as ordinary failure. Late snapshots
after cancellation or timeout are ignored. These snapshots are log-only; they
must not contain source text, search queries, URLs, or credentials.
Public research reports setup, session opening, model waiting, and search-call
boundaries. Search-call counts distinguish a model that has not invoked search
from an outstanding search or a model that has received results. Each web-search
tool invocation has a 60-second limit, including authorization, and passes its
abort signal to the MCP client. A timed-out search ends the research attempt with
web-search-timeout; a stalled authorization check produces
web-search-authorization-timeout. Search calls are counted after authorization.
The overall evidence-stage limit remains five minutes. Use the host wrapper's
search counters for live progress: the harness can buffer tool lifecycle
callbacks until the completed step is published.
A definition may delay completion with publish():
async publish({ job, artifact, commit }) {
if (!await stillCurrent(job, artifact)) {
throw new Error("target changed before publication");
}
await commit();
}The hook runs after artifact validation but before settlement. It must call and
normally await commit() exactly once while the hook is active. A hook that
does not commit leaves the attempt uncompleted; a late or second commit is
rejected. If the hook commits and then throws, the durable completion remains
authoritative and the service logs the hook error.
Use a publication hook only when completion needs an external freshness or
serialization guard. document-summary@1 uses one to hold the document mutation
gate and recheck revision and source hash. Ordinary jobs settle directly.
The generic service publication callback sends only a background-job channel revision invalidation. A separate domain service may observe a completed job and persist its own projection. That reconciliation must be idempotent and recoverable on a later read because observer delivery is not a durable queue.
credential: "active-planner" means the job uses the channel's existing
Planner owner and Copilot entitlement. The runner resolves ownership through
ActiveOwnerBindings; it never claims ownership itself. An explicit product
action may establish ownership before enqueueing. Without an available owner,
the job pauses as owner-unavailable.
Resolution and later revalidation check the process-local session, credential revision, ownership generation, instance admission, repository node identity, GitHub App installation, repository write role, and expiry. Definitions must never copy the GitHub token into input, artifacts, progress, diagnostics, or PostgreSQL. A durable claim binding contains only the owner session ID, owner generation, credential revision, and repository ID.
credential: "none" skips owner resolution. The current server still disables
the entire runner when AGENT=off, so a credential-free definition does not run
under that flag without a deliberate change to the main runner gate.
Do not reuse the Planner conversation session. Every worker stage runs its own
HarnessAgent with a fresh, disposable session. A HarnessAgent's output
schema is fixed per agent, not per turn, so the four research stages — public
evidence, private document analysis, private report synthesis, and private
answer synthesis — each get their own named agent with its own result schema;
summaryAgent is a fifth, for document descriptions, and researchBriefAgent
prepares Chat research offers in its own no-tool session. Generated document
descriptions reuse one disposable session across multiple bounded chunk and
reduction turns; every other stage submits one bounded structured result.
A private stage's agent declares an empty tool set and empty activeTools, so
its session has no MCP server, web search, URL fetch, repository tool, shell,
filesystem, host Git, skill, or plugin. The only way it can finish is by
returning output that matches its output schema.
Only the public evidence agent binds a host web_search tool, reached through
createMCPClient against GitHub MCP's web_search toolset, and lists it as
its only activeTools entry. Never give one agent both private document
context and public web capability.
The public evidence worker receives only the submitted brief. An isolated no-web document-analysis worker receives the brief and parent-document snapshot. A separate isolated no-web report-synthesis worker receives the brief, normalized public evidence, and the private findings from document analysis.
Accepted public sources must be canonical public HTTPS URLs observed in successful web-search citation or typed-resource metadata. Bare URLs in model prose do not establish provenance. Chopin never follows, resolves, or fetches a source URL, and generated reports may cite only URLs retained in normalized evidence. This validates provenance, not factual accuracy or page safety.
JobLimits.maxAiCredits is not centrally enforced by JobRunner. It describes
the maximum aggregate AI credits for one job attempt, so a model-backed
executor must register per-harness-session limits whose possible total does
not exceed it. The Copilot SDK adapter keeps a credit-limited harness session
on one SDK session across its turns and refuses a later turn that would require
different native settings or a new credit window. All document-description
chunks and reductions share its 64-credit limit. Each research stage uses a
fresh 30-credit session; the two initial answer stages fit the 60-credit job
limit. Keep definition metadata and adapter settings synchronized.
Always recheck credential.authorize, observe credential and job abort
signals, and destroy the harness session in finally.
| Definition | Production trigger | Persisted input | Worker boundary |
|---|---|---|---|
document-summary@1 |
Open, edit, restore, or MCP persistence | Revision, source hash, generator version, and output:"description"; not the source |
Private worker; source loaded at execution and publication rechecks current revision/hash |
research-evidence@1 |
Immediate research request | Internal workspace, initial turn, exact submitted brief | Public evidence agent bound to a host web_search tool, no other tools |
research-answer@1 |
Completed evidence | Parent document snapshot, evidence, and internal compatibility history | Two no-web private agents with structured output for analysis and report synthesis |
research-brief@1 |
New or updated Chat research offer | Offer generation, selected Chat sources, decision snapshots, and prior brief | Private structured worker; no tools; source IDs and offer generation checked before projection |
document-summary@1 remains the only durable definition version; there is no
document-summary@2. New V1 requests carry the output marker and use the
idempotency key description-v1:<plan revision>:<source hash>. The distinct
request identity means an unchanged document whose existing V1 artifact is a
legacy summary is eligible for lazy description generation.
The V1 codecs still read markerless legacy inputs and their summary artifacts.
Those artifacts remain available through Background Work, but the catalogue
projector ignores them. A new marked artifact instead carries
output:"description" and one physical line identifying the document's type,
purpose, and subject, for example PRD for ..., RFC about ..., or Plan for .... The description retains the broad 4,000-codepoint and 16 KiB artifact
safety limits. A blank canonical source produces Empty document without a
model call.
Scheduling is lazy rather than a database-wide backfill. Canonical edits debounce new work; opening or restoring a document ensures work for its current source; MCP creation and replay schedule after persistence. Execution requires that channel's active Planner owner, so work pauses without an available owner. There is no unattended scan that establishes owners or regenerates every stored document.
After a marked job completes, an idempotent projector writes the description to
channel catalogue metadata with the source plan revision and hash, generator
version, source job ID, projection timestamp, and an independent description
revision. The previous description remains visible and searchable while newer
work is pending or failed. Projection does not advance collaboration revision or
channel updatedAt, so it does not change list recency.
Allowed origins are broader than current callers: the document-summary definition permits scheduler and planner; research definitions permit user and planner. Current production enqueue paths use scheduler for descriptions and user for research execution.
Generated descriptions and legacy document-summary artifacts are not the
reserved Planner transcript summary or an MCP creation brief. They are not
injected into a later Planner turn automatically. Descriptions are untrusted
model output. The Planner may explicitly call list_background_jobs and
read_background_job, and must treat the returned artifact as untrusted
evidence.
The public product has one immediate research request; the database and service
retain research_workspaces, turns, and messages as internal staging and
historical-reference compatibility:
- An inline card or current Planner turn persists the exact brief, one initial turn, attribution, and stable request identity.
- The same action establishes the channel's Planner owner when needed and
enqueues
research-evidence; there is no draft or confirmation step. - The evidence worker receives only that exact brief, with no private document context, and may derive or refine the queries it sends to web search.
- Completed normalized evidence causes an idempotent
research-answerenqueue. - The answer job captures current canonical parent-document context. It first runs an isolated no-web document-analysis worker, then a separate isolated no-web report-synthesis worker with the resulting private findings and normalized public evidence.
- A complete report artifact is validated and converted to canonical MDX.
- One fenced storage transaction creates an initialized ordinary child channel and links it to the request. Only then does the card become ready and the child enter navigation.
Failure or cancellation publishes no child. Explicit retry clears only terminal initial job links under optimistic guards, preserves the request and brief, and starts a new evidence generation. Reconciliation runs after completion and on a later request read, so interruption cannot require a second child or a manual repair.
Here, private means not disclosed to public web search. Internal staging rows, selected compatibility projections, and artifacts are readable through authorized server paths. Normalized job input is persisted but omitted from the inline card projection. Private worker material is still sent to the hosted model inference service, through the configured harness, under the active owner's credential.
Archiving a document suspends its description coordinator, cancels any pending description debounce, and prevents new description scheduling. It does not blanket-cancel background jobs already admitted for that channel. New research requests and explicit retries are blocked while archived, but already-started evidence and answer jobs may settle, and their idempotent request reconciliation may still persist. Publication rechecks that the parent is active, so an archived parent does not gain a child.
Permanent deletion has a stronger boundary. The runner first blocks new claims for the channel while description scheduling remains suspended, aborts its active attempts, cancels every pending, paused, or running job, and waits a bounded grace period. The server then closes the live plan before atomically deleting the archived channel. Claim fencing rejects any late worker progress or artifact, and an eligible channel delete cascades job targets, jobs, artifacts, and internal research staging. A parent with a child and a child linked from a published request are deliberately not eligible for deletion.
The generic schema supports a new standalone job type without a migration:
background_job_channelsstores the independent background-job channel revision;background_job_targetsstores current target generations;background_jobsstores requests, state, claims, normalized input, progress, and sanitized failure reason;background_job_artifactsstores one immutable completed result per job.
A migration is needed only when the job adds domain-specific durable records or
links, as research request staging does. When changing StorageAdapter, update the
model, port, memory adapter, PostgreSQL adapter, migration, migration registry in
apps/server/src/storage/postgres/migrations.ts, migration tests, and shared
contract suite together. Never edit an applied migration; checksums make that a
startup error. See Storage for the table map.
Job summaries omit input, fingerprint, idempotency key, and claim binding. This does not make input ephemeral: normalized input is stored in PostgreSQL. For example, document-description input omits document source, while research-answer input deliberately persists the captured source snapshot. Store only material the workflow needs and document its retention boundary.
There is no current pruning policy for jobs, artifacts, or internal research staging while their parent document remains stored. Cancellation prevents child publication; it does not erase input or history. Full eligible document deletion is the exception and removes that channel-owned state.
The current browser has no generic Background Work destination. Research status
is deliberately projected through the inline request card: queued, searching,
analyzing, writing, publishing, ready, failed, or cancelled, plus discovered
sources and safe errors where available. The card uses request create, detail,
cancel, and retry routes and refreshes after research:changed.
The server retains job:list, job:get, job:cancel, and job:changed
for internal and compatibility consumers. Repository readers may list and read.
Cancellation requires current write access and is restricted to active,
user-origin research evidence and answer jobs.
A newly registered type receives no browser surface automatically. Add or update these when a domain projection is needed:
- subject extraction in
apps/server/src/jobs/service.tsand both job storage adapters; - safe browser serialization and cancellation rules in
apps/server/src/jobs/browser.ts; - protocol declarations in
packages/protocol/job.d.ts; - domain HTTP state and invalidations when the job belongs to a larger workflow.
The Planner can list and read jobs, but has no generic enqueue, pause, resume, cancel, or artifact-trust capability. A domain-specific Planner tool must still respect the job origin, authorization, disclosure, and persistence boundaries.
Only the exact value off disables these flags:
AGENT |
BACKGROUND_JOBS |
WEB_RESEARCH |
Result |
|---|---|---|---|
| on | on | on | Document description, research evidence, and research answer definitions; runner and coordinator active |
| on | on | off | Document description and private research answer definitions; no new public evidence |
| off | on | forced off | Persisted jobs and descriptions remain stored; the runner and description coordinator are off |
| any | off | forced off | No definitions execute |
Configuration is read at startup. Restoring a disabled capability requires a restart. Existing durable rows and artifacts are not deleted by a flag.
AGENT=off disables the entire runner, including credential: "none" jobs.
Gating a definition itself can leave durable rows paused as unregistered-type.
Prefer keeping old definitions registered while gating new enqueue paths. If a
definition must be removed conditionally, provide and test an explicit resume
path; automatic owner recovery only revives compatible active-planner work.
If a new definition needs its own flag, update configuration parsing and tests,
.env.example, Compose metadata and tests, self-hosting documentation, and the
E2E server environment where relevant.
- Registering a definition without adding an enqueue path.
- Removing an old version while durable work still references it.
- Treating caller input as the job origin.
- Treating a stamped origin as authentication or authorization.
- Pre-prefixing the target key before calling
JobService. - Reusing an idempotency key with changed semantic input.
- Persisting tokens, raw provider events, prompts, or private prose in progress and diagnostics.
- Assuming
maxAiCreditsis runner-enforced. - Assuming
maxAttemptsis compared with claim count rather than failures. - Ignoring the abort signal because claim fencing rejects late output.
- Reusing a Planner session for background work.
- Giving a private worker ambient MCP, URL, repository, shell, or filesystem capability.
- Giving one worker both private context and public web search.
- Calling an undeclared progress stage.
- Calling a publication
commit()twice, too late, or not at all. - Broadcasting or projecting state before its durable commit.
- Assuming process-local completion callbacks are durable delivery.
- Bypassing
JobServiceand therefore bypassing codecs, origins, versions, and byte limits. - Assuming Playwright executes live harness work; E2E runs with
AGENT=offand seeds validated job state.
Use the narrowest tests that exercise the new behavior:
# Definition, service, runner, and domain tests
bun test apps/server/src/jobs apps/server/src/research
# Memory adapter and the rest of the unit/domain suite
bun test
# Real PostgreSQL transitions, fencing, and migrations
bun run test:postgres
# Type and static validation
bun run types
bun run ci
# Browser integration when the job has a user-facing surface
bun run e2eDefinition tests should inject engines rather than call a live harness. Cover strict codecs, exact context separation, progress, citation or provenance rules, and artifact rejection. Runner tests should cover success, retries, timeout, owner loss, credential rotation, cancellation, heartbeat loss, late output, and shutdown. Storage behavior belongs in the shared adapter contract so memory and PostgreSQL run the same cases.
- Definition and registry contract:
apps/server/src/jobs/registry.ts - Validation, enqueue, settlement, and publication:
apps/server/src/jobs/service.ts - Claims, execution, retries, and shutdown:
apps/server/src/jobs/runner.ts - Current definitions:
apps/server/src/jobs/document-summary.tsandapps/server/src/jobs/research-workspace.ts - Document-description projection:
apps/server/src/jobs/document-description.ts - Definition registration and runtime wiring:
apps/server/src/main.ts - Active owner resolution:
apps/server/src/agent/active-owner.ts - Worker agents (
summaryAgent,researchAgent, and the private research stage agents):apps/server/src/harness/agents.ts - Durable job model and port:
apps/server/src/storage/model.tsandapps/server/src/storage/port.ts - PostgreSQL job store:
apps/server/src/storage/postgres/jobs.ts - Generic protocol compatibility:
packages/protocol/job.d.ts - Research orchestration:
apps/server/src/research/service.ts