Problem: The reporting CI job takes 41m (PR #98), far longer than query_engine. All ~42 reporting tests run serially in one Spark session.
Context: CI already parallelises across components (matrix over query_engine/reporting in acceptance.yml), but within the reporting job pytest is single-process. pytest-xdist is not a dependency. The 15 integration tests likely dominate runtime.
Catch: Tests share global catalog state under fixed names (setup_basic_db rebuilds spark_catalog.silver; cleanup_gold drops all gold tables after every test). Any process-level parallelism needs per-worker isolation of the Spark warehouse/metastore or per-worker schema names.
Proposed steps:
- Profile:
pytest tests/impulse_reporting --durations=25 to confirm the hot spots.
- Quick win: split the
reporting matrix entry into sub-path shards (integration vs unit). No code change.
- Bigger win: add
pytest-xdist (-n auto) and make the spark fixture worker-aware so shared-state fixtures stop colliding.
Done when: reporting CI is well under 20m, with no coverage loss and no cross-worker state bleed.
Problem: The
reportingCI job takes 41m (PR #98), far longer thanquery_engine. All ~42 reporting tests run serially in one Spark session.Context: CI already parallelises across components (matrix over
query_engine/reportinginacceptance.yml), but within the reporting job pytest is single-process.pytest-xdistis not a dependency. The 15 integration tests likely dominate runtime.Catch: Tests share global catalog state under fixed names (
setup_basic_dbrebuildsspark_catalog.silver;cleanup_golddrops all gold tables after every test). Any process-level parallelism needs per-worker isolation of the Spark warehouse/metastore or per-worker schema names.Proposed steps:
pytest tests/impulse_reporting --durations=25to confirm the hot spots.reportingmatrix entry into sub-path shards (integration vs unit). No code change.pytest-xdist(-n auto) and make thesparkfixture worker-aware so shared-state fixtures stop colliding.Done when: reporting CI is well under 20m, with no coverage loss and no cross-worker state bleed.