Skip to content

[FEAT]: Wire durable xdist shard transport into the pytest plugin - #3

Closed
nina-msft wants to merge 1 commit into
dev/nina-msft/xdist-shard-primitivesfrom
dev/nina-msft/xdist-durable-transport
Closed

nina-msft wants to merge 1 commit into
dev/nina-msft/xdist-shard-primitivesfrom
dev/nina-msft/xdist-durable-transport

Conversation

@nina-msft

Copy link
Copy Markdown
Owner

Description

Stacked on microsoft#113 (dev/nina-msft/xdist-shard-primitives).
This intra-fork draft targets the PR-A branch so the diff shows only the wiring
added on top of the shard primitives. Once microsoft#113 merges into
main, this branch will be rebased onto main and reopened against
microsoft:main.

This PR wires the durable shard-transport primitives added in
microsoft#113 into the pytest plugin, closing the two durability gaps
@bashirpartovi asked us to "call out clearly" in the review of
microsoft#73:

What changed

For local popen topologies (plain pytest -n N, or --tx where every gateway
is popen), each worker now streams every finished result to its own on-disk
JSONL shard, flushing per write
. The controller reads the shards back in
pytest_testnodedown. workeroutput carries only a small completion sentinel
(schema, expected record count, trial specs).

  • Durable: a worker killed mid-run (crash, OOM, timeout, -x) keeps every
    result it had already flushed; the controller recovers them from the shard and
    marks the run incomplete rather than losing the worker's whole contribution.
  • Granular size cap: the cap now applies per record — one oversized result
    is dropped as a truncation marker while its siblings survive.
  • Remote fallback: any non-popen gateway (ssh, socket, via proxy)
    transparently falls back to the existing inline transport, since the controller
    cannot read a worker-local shard file. Eligibility is gated by shard_eligible.

Controller/worker wiring: pytest_configure provisions the shared shard dir on an
eligible controller; pytest_configure_node pushes it into each worker's input;
the worker opens a ShardWriter; the _rampart_collect fixture appends per
result; pytest_sessionfinish closes the shard and writes the sentinel;
pytest_testnodedown recovers and merges. Shard dirs are removed at run end
unless --rampart-keep-shards is passed. Single-process and remote runs are
unchanged.

Breaking changes

None. Behavior is identical for single-process runs and for remote/proxied xdist
topologies. Local -n runs gain durability transparently; the only observable
difference is more granular incomplete-reason strings (per-record rather than
whole-payload).

Checklist

  • pre-commit run --all-files passes
  • Tests added or updated for changes
  • Documentation updated

Note: Keep this PR in draft until microsoft#113 merges. Once it
lands on main, rebase this branch onto main and reopen against
microsoft:main, then mark ready for review.

Stream each finished result to a per-worker on-disk JSONL shard (flushing
per write) for local popen topologies, so a worker killed mid-run keeps
its already-finished results and the size cap applies per record. The
controller recovers shards in pytest_testnodedown; workeroutput carries
only a completion sentinel. Remote/proxied gateways fall back to the
inline transport. Closes the durability (#3) and granular size-cap (microsoft#4)
gaps from the microsoft#73 review.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant