Skip to content

fix(compliance-checks): don't require __file__ at module load on serverless - #795

Draft
surojitchowdhury wants to merge 3 commits into
developmentfrom
fix/compliance-checks-serverless-file-685
Draft

surojitchowdhury wants to merge 3 commits into
developmentfrom
fix/compliance-checks-serverless-file-685

Conversation

@surojitchowdhury

Copy link
Copy Markdown
Contributor

Fixes #685

Summary (plain language)

Ontos runs its compliance checks as a scheduled Databricks job on serverless compute. The entry script tried, at its very first executable line, to work out its own folder location using Python's __file__ variable. On serverless the runtime executes the script via exec(compile(...)) in a namespace where __file__ is not bound, so evaluating Path(__file__) raised NameError at module load — before main() ever ran. The scheduled compliance run therefore died on startup every time on serverless.

This change makes the module-level sys.path bootstrap no longer depend on __file__ being present, so the job starts on serverless, while behaving exactly as before on a normal cluster / locally.

Change

src/backend/src/workflows/compliance_checks/compliance_checks.py — the single unguarded line

sys.path.insert(0, str(Path(__file__).parent.parent.parent))

is replaced with a guarded bootstrap:

_entry_file = globals().get("__file__")
if _entry_file is not None:
    try:
        sys.path.insert(0, str(Path(_entry_file).parent.parent.parent))
    except Exception as _exc:  # pragma: no cover - defensive; must never abort import
        print(f"WARNING: compliance_checks sys.path bootstrap failed: {_exc}", file=sys.stderr)

Approach

A key fact drives the design: the workflow's from src.* imports are lazy (inside load_policies / run_policy / main), never at module top level. So importing this module never itself needs src on sys.path — the entry only matters at runtime on a normal cluster / locally.

  • With __file__ present (local dev / normal cluster): behaviour is unchanged — the backend source root is Path(__file__).parent.parent.parent, prepended to sys.path.
  • Without __file__ (serverless exec(compile(...))): insert nothing. There is no reliable signal to derive the real source root here — the deployer uploads only this workflow folder (not the src tree), and serverless job environments cannot carry env vars. Guessing from the current working directory could prepend an unrelated directory and mask similarly named packages, so the bootstrap adds nothing and lets the runtime environment (installed package / PYTHONPATH) resolve the lazy imports. If that is ever insufficient, the failure surfaces later as a clear ImportError instead of a cryptic NameError at module load.
  • The __file__-present insert is narrowly guarded: an unexpected failure is reported to stderr (diagnosable) but can never abort module import.

Scope is strictly the path-bootstrap; no unrelated logic was refactored, and nothing under .github/workflows/ is touched.

Tests

Adds src/backend/tests/test_compliance_checks_workflow.py with two smoke tests that go beyond "no crash" and assert path correctness:

  • Serverless case — compiles the module source and execs it in a globals dict with no __file__ ({"__name__": "not_main"}), after chdir-ing into a simulated deployed workflow folder. Asserts: no NameError; main() did not run; and — crucially — sys.path is left exactly unchanged (nothing prepended), with the cwd-derived grandparent specifically absent. This closes the false-PASS gap where a wrong path would otherwise slip through.
  • Normal caseexecs with __file__ present and asserts the correct source root is inserted at sys.path[0].

Top-level third-party imports (sqlalchemy, databricks.sdk) are stubbed when not importable, so the tests isolate the __file__ bootstrap; because the assertions inspect sys.path directly, stubbing cannot mask an incorrect entry.

Gate output

$ hatch -e dev run pytest backend/tests/test_compliance_checks_workflow.py \
    backend/tests/test_compliance_dsl.py backend/tests/test_compliance_actions.py -q
======================= 57 passed =======================

Only 2 files changed (the workflow script + the new test); nothing under .github/workflows/.


Re-raised from a branch on this repository (previously #725, opened from a fork) so the required CI workflows can run — fork PRs do not receive the internal JFrog OIDC token, which failed every check at setup before any test ran. Same commits and diff; #725 is closed in favour of this PR.

…erless

The compliance-checks workflow is submitted as a spark_python_task that runs
on Databricks serverless, where the entry script is executed via
exec(compile(...)) in a namespace with no __file__ bound. Evaluating
Path(__file__) at module import time raised NameError before main() ever ran,
so the scheduled compliance job died on startup every time on serverless.

Guard the sys.path bootstrap with globals().get("__file__"): use the
__file__-derived source root when present (unchanged local behaviour), and fall
back to the working directory (the deployed workflow folder on serverless) when
absent. The whole bootstrap is wrapped in try/except so it can never abort
module load.

Adds a smoke test that execs the module source with no __file__ in globals
(faithfully simulating serverless) and asserts no NameError, plus a case
proving normal import with __file__ present still works.

Fixes #685

Co-authored-by: Isaac
…erless

Address cross-review of #685:

- Serverless (no __file__): do NOT fabricate a sys.path entry. There is no
  reliable signal to derive the real source root -- the deployer uploads only
  the workflow folder (not the src tree) and serverless job environments cannot
  carry env vars ("compute.Environment doesn't support env_vars directly"). A
  cwd-based guess could prepend an unrelated dir and mask similarly named
  packages. The module's app imports are lazy (inside functions), so module
  import needs no sys.path entry; runtime from-src.* resolution relies on the
  environment/PYTHONPATH/installed package. If insufficient it fails later with
  a clear ImportError, not a cryptic NameError at load.
- __file__ present (local/normal cluster): behaviour unchanged -- source root is
  Path(__file__).parent.parent.parent, prepended to sys.path.
- Narrowed the exception handling: only the __file__-branch insert is guarded,
  and it warns to stderr instead of silently swallowing, so a real bootstrap
  failure is diagnosable while module import can never crash.

Strengthen the smoke test to close the false-PASS gap: the serverless
simulation now chdirs into a deployed-workflow-shaped temp dir and asserts
sys.path is unchanged (no fabricated/cwd-derived entry), and the __file__ case
asserts the correct source root is prepended. sys.path is snapshotted/restored
per test; dependency stubs cannot mask a wrong entry because assertions inspect
sys.path directly.

Fixes #685

Co-authored-by: Isaac
@surojitchowdhury
surojitchowdhury requested a review from a team September 10, 2026 13:33
@surojitchowdhury surojitchowdhury added type/bug Something isn't working scope/compliance Compliance check related feature tech/python Pull requests that update python code labels Sep 10, 2026

@mvkonchits-db mvkonchits-db left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed with the code-review skill. The fix correctly removes the import-time NameError from referencing __file__ at module load, but I don't think it makes the workflow actually runnable on serverless — leaving as a comment rather than approving so it can be verified.

  1. The crash may just move later. On serverless the bootstrap now inserts nothing on sys.path. Module load then succeeds, but main() later hits lazy imports like from src.controller.compliance_manager import ComplianceManager and from src.db_models.compliance import CompliancePolicyDb. The serverless env spec installs only databricks-sdk/sqlalchemy/psycopg2-binary (not the ontos backend), so unless the runtime already has the backend src parent on PYTHONPATH, those raise ModuleNotFoundError — the same job failure, just later.

  2. Possible off-by-one in the non-serverless path. The insert uses Path(__file__).parent.parent.parent = src/backend/src (the src package dir itself). But from src.* resolves only when src/backend (the package's parent) is on sys.path. Inserting the package dir makes import src unresolvable in any env that relies solely on this insert.

  3. Tests don't cover the failure point. Both new tests set __name__ != '__main__', so main() never runs and the lazy from src.* imports are never exercised — the suite passes without touching the thing that actually fails on serverless. Also assert sys.path == sys_path_before is order-dependent and can flake if a real dep import mutates sys.path first.

Suggestion: validate against a real serverless run (or a test that invokes main()/the lazy imports under a simulated serverless sys.path) before merging, and double-check the insert targets src/backend, not src/backend/src.

@surojitchowdhury
surojitchowdhury marked this pull request as draft September 11, 2026 14:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

scope/compliance Compliance check related feature tech/python Pull requests that update python code type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Compliance-checks scheduled job crashes on serverless: NameError on __file__ at import time

2 participants