Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 8 additions & 2 deletions SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ V17ab: ∀ `.py` file in repo ! `ruff check` + `ruff format --check` clean (CI-e
V18rs: ∀ `.github/workflows/**` change ! at least one validation job executes the affected workflow path; shared CI workflow changes ⊥ leave the validation matrix entirely skipped
V19cn: public API ! use VOCAB nouns (sample/variable=parameter/response=QoI); ⊥ "job" in any itis_sumo signature, docstring, or result field
V20dm: consumer-facing entrypoints ! accept plain tabular data; ⊥ `FunctionJob`/`JobVariableSelection`/oSPARC status strings inside itis_sumo (test-enforced, sibling of V4ty)
V21pf: preprocessing ! auto-defaulted + overridable in domain vocabulary only; ⊥ transform vocabulary reachable from a public signature; effective config inspectable in the result
V21pf: preprocessing ! auto-defaulted + overridable in domain vocabulary only; ⊥ transform vocabulary reachable from a public signature; effective config inspectable in the result; ∀ value-producing entry point ! accept `scale` (default linear — ⊥ caller must pass) ∧ ! apply it (V45ls)
V22rs: results ! typed dataclass, original units + original names, `dataclasses.asdict()`-serializable; ⊥ `_hat`/`_std_hat` suffix keys in the public shape
V23er: ∀ error escaping `itis_sumo.api` ! be a `SumoError` subclass; ⊥ raw `KeyError`/`IndexError`/`ValueError` crossing the boundary
V24af: run dir ! discarded on success, PRESERVED on failure w/ its path + stderr tail attached to the raised `SumoEngineError`
Expand All @@ -109,6 +109,8 @@ V40gz: commit-time hooks (`prek` tool, run via `uvx prek run --all-files`; mirro
V41hp: `ty check` ! zero diagnostics repo-wide (blocking in both the local `prek` hook and CI); version-gated stdlib/typing symbols (e.g. `datetime.UTC`, `typing.Self`, `tomllib` — all 3.11+-only, caught mid-stack once actually run, T42jn) ! resolved against `requires-python`'s floor, not the CI runner's own interpreter — this is strictly cheaper than the 3.10-3.13 test matrix (V18rs) for that bug class, though the matrix stays as the ground-truth runtime check
V19nd: ∀ Dependabot ecosystem entry in `.github/dependabot.yml` → weekly Monday 03:00 Europe/Zurich schedule ∧ exactly one wildcard dependency group; ⊥ unspecified default schedule or one PR per dependency
V20qx: ∀ ignored Dependabot dependency → ignore reason names active compatibility constraint ∧ references tracked resolution task; `itis-dakota` remains ignored until T16mo resolves the Dakota 6.23+ interface-cache regression
V44ls: ∀ scale-consuming `itis_sumo.api` workflow (surrogate fit ∧ UQ/Sobol input sampling ∧ MOGA search domain ∧ CV accuracy metric) → a `scale="log"` override ! be honoured in that workflow's own computation (log-space surrogate + log-space sampling/search where a distribution or domain is involved, positivity-guarded); ⊥ a `scale` override accepted but silently ignored outside the surrogate fit
V45ls: ∀ value-producing `itis_sumo.api` entry point → `scale` ! be threaded into the computation that yields those values ∧ observable in the result; ⊥ a value the api produced while ignoring the caller's scale; enforced: structural (internal value-producers `scale_distribution`/`resolve_log_scale` take scale/flag as a required arg — unwired code breaks at call w/ `TypeError`, never silent drift) ∧ behavioural (linear↔log flip-test over ∀ entry points (10 table-mode workflows + `evaluate_correlations`); rank-correlation exempt — monotone-invariant, asserted unchanged)

## §R
R1: `export_model`/`import_model` child keywords; formats `text_archive`(.sps)/`binary_archive`(.bsps)/`algebraic_file`(.alg); naming `{prefix}.{resp}.{ext}` | branch R2
Expand Down Expand Up @@ -149,7 +151,7 @@ T23bn|✓|`itis_sumo.api`: `cross_validate()` + `evaluate_along_axes()` — inte
T24cm|✓|remove `preprocess/models.py::{FunctionJob,JobVariableSelection}`; re-express the Dakota sufficiency rule over tabular data; jobs→table adapter moves to flaskapi (this port DELETES itis-sumo code)|V20dm
T25dp|.|`itis_sumo.api`: remaining 6 workflows (UQ-w-uncertainty incl. the ~120-line erfinv/histogram block, correlation, Sobol, grid, MOGA, cv-accuracy-metrics) + E1 `export_model`/`evaluate_stored_model` facade|V22rs,R1-R4
T26eq|.|POST-PORT: extract the fitted-model handle (`fit()` → methods → `save()`/`load()`); carries the fitted preprocessing config ⇒ closes the E1 gap (model store persists archive+metadata+training copy but ⊥ preprocessor config, so a reloaded model cannot inverse-transform to original units)|V27fq,V10jk
T27fr|.|POST-PORT: split `domain` vs `distribution` config + consumer migration; absorb the mmux_vite `jgo/fullstack-logscale` work|V26dd
T27fr|~|POST-PORT: split `domain` vs `distribution` config + consumer migration; absorb the mmux_vite `jgo/fullstack-logscale` work — backend log-scale now honoured end-to-end (surrogate + along-axes + grid + CV + UQ sampler + CV-metrics + MOGA + Sobol, V44ls); remaining = the `domain`⊥`distribution` config split (V26dd) ∧ the mmux_vite frontend consumer migration|V26dd,V44ls
T28gs|.|POST-PORT `?`: decide whether `distribution` gets an auto-generated default — explicit discussion required, ⊥ silently defaulted|§C `?`
T29hw|~|`publish.yml` tag trigger accepts `v`-prefixed PEP 440 prereleases ✓; tagging `v0.1.0a1` BLOCKED on the one-time PyPI Trusted Publisher config (user action), then clean-venv install + `itis-sumo validate` + headless smoke|T1pw
T30qa|✓|release/CI workflow refresh: build→TestPyPI automatic on tag push (✓), PyPI+Release gated manual (✓, T35cc); auto-tag-on-branch redesigned to PR-time check + tag-only-at-merge (no bot commit) — see T31xx/T32yy; keep git-cliff release notes, dependency-review + concurrency (✓); skip weekly cron/healthchecks for now|§C,V17rt
Expand All @@ -159,6 +161,8 @@ T33zz|✓|swap itis-sumo LICENSE + `pyproject.toml` `license`/classifiers to mat
T34aa|✓|`make publish-testpypi-dev` ! source `TESTPYPI_TOKEN` from local `.env` (gitignored) instead of requiring a pre-exported shell var|§C
T35cc|✓|gate `publish`/`release` jobs behind explicit `workflow_dispatch`; `build`+`verify` stay CI-only for release tags; feature `.devN` uploads move to local Make target|V31vp,V32bb
T36dd|✓|add `.env` to `.gitignore`; make `publish-testpypi` source `.env` without printing token, then build/check/publish|V37bb
T46ls|x|scale first-class across the WHOLE api surface: `generate_lhs_samples` ∧ `generate_grid_samples` ∧ `compute_correlations` take `preprocessing` ∧ honour scale; one `scale_distribution` unit→value helper absorbs the log10-power hand-roll ∧ the duplicate log guards ∧ the `_is_log` closure; V45ls linear↔log flip-tests over ∀ 10 entry points + 5 gap tests (MOGA log·maximize, log-variable axis x-restoration, log-response UQ spread, Sobol mixed partition, repo-wide `ty`); rows land in the V&V report Category I|V45ls,V44ls,V21pf,V22rs
T47mt|✓|`api.evaluate_correlations`: move the #470 MC-through-surrogate correlation workflow (draw → surrogate predict → Pearson+Spearman over the SHARED sample set) out of the flaskapi endpoint into the api — scale-native from birth (engine producer `correlate_manual_uq_samples` takes `input_scales`/`output_scale` as required kwargs); `CorrelationResult.seed` records the draw; flip matrix extended over ∀ 11 entry points|V45ls,V22rs
T37ef|✓|validate manual publish tag + artifact version before PyPI upload|V35wk
T38dd|✓|make `publish-testpypi-dev` auto-compute/write `.devN`, build/check/upload directly to TestPyPI, then restore project version; CI verifies alpha/beta/rc before real PyPI; no dev tag cascade|V38cc,V39qf
T39pk|.|POST-PORT/deepen candidate: `itis_sumo.api.workflows` repeats the same build-session → translate-config → call → translate-result → wrap-errors skeleton per function; pull the shared shape down into `_session.py` so each `workflows.py` function shrinks to signature+one call — behavior-preserving refactor, tests green before/after; natural precursor to T26eq's handle extraction|V16qf,V27fq,T26eq
Expand Down Expand Up @@ -191,3 +195,5 @@ B14zt|2026-08-27|`itis-dakota==1.5.9` was deleted from PyPI upstream (release li
B15tg|2026-08-27|the `v0.1.0a2` tag (created by auto-tag.yml on the develop merge) exists, but no TestPyPI publish followed — auto-tag's trailing `gh workflow run publish.yml` did not fire. Root cause: the `tag-merged-version` step invoked `gh` from an inline Python script that read `GITHUB_TOKEN` out of the step environment, but the step's `env:` block only exported `GH_REF_NAME`, so `gh` got no usable token and the dispatch silently did nothing|add `GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}` to the `tag-merged-version` step's `env:` block so the inline `gh` call inherits a valid token; one-off `gh workflow run publish.yml --ref develop --field tag=v0.1.0a2 --field registry=testpypi` published a2 manually to recover the missed alpha; V29yn
B16mo|2026-08-28|T16mo rung 2 bumped the engine `itis-dakota==1.5.11` (Dakota 6.20) → `==6.24.7` (Dakota 6.24). 6.24.7 still carries the 6.23+ `Interface::interface_cache()` regression (R5): `dakota.environment.study()`'s ctor unconditionally calls `interface_cache(problem_db)` and throws `IndexError: map::at` ("called with nonexistent study!") for itis-sumo's pure data-fit surrogate confs, which construct no Interface (only `import_build_points_file`, no `interface` block) — the cache static map is never populated. All 45 real-Dakota surrogate tests failed on it (training, cross-validation, UQ propagation, MOGA, import/export, evaluation). Upstream fix (get-or-create `operator[]`) is still pending|shim in itis-sumo, not waiting on upstream: `config/funs_create_dakota_conf.py::add_r5_interface_cache_workaround()` (appended to every conf via `start_dakota_file()`) declares an unused `single` model `R5_WA_MODEL` wired to a no-op `fork` interface (`analysis_drivers='true'`) and points each surrogate `model` at it via `truth_model_pointer='R5_WA_MODEL'`, forcing the Interface to be constructed (populating the cache) during study setup; the dummy model is never referenced by any method/iterator so the interface is never executed. 318-test suite green on 6.24.7. Took 6.24.7 (not 6.24.9) because 6.24.8+ dropped cp311 wheels, which would have forced a Python 3.12 migration of both repos; R5
B8hm|2026-09-11|Dependabot entries used unspecified weekly timing, and `itis-dakota` was intentionally ignored without a durable policy invariant|V19nd,V20qx
B17sc|2026-09-28|`scale="log"` on `PreprocessingSpec.overrides` reached the surrogate fit but was a near-no-op for UQ/MOGA/cv-metrics/Sobol — the logscale port (T27fr) landed the preprocessor transform but not per-workflow log-space sampling/search, so `evaluate_uncertainty` with a log input returned linear-space statistics (probe: mean 12.057 linear vs 12.058 log). Symptom never seen in flaskapi, whose frontend drove log at the request-payload layer|wire log-space into every scale-consuming path: `create_manual_uq_samples` log_scale branch (log-uniform, uniform-only + positivity), `_uq_engine_distributions` flags ∧ validates log vars for UQ + Sobol, `evaluate_sobol_indices` samples `scipy.stats.loguniform`, `optimize_pareto_front` maps a log variable's search domain into ln space ∧ fits log objectives; `evaluate_cv_metrics` inherits log via `cross_validate`. Guarded by V44ls ∧ tests across the four workflows|V44ls,T27fr
B18mt|2026-09-29|T25dp's "correlation" port landed only table-mode `api.compute_correlations`, leaving the #470 MC→surrogate→correlate endpoint workflow (sample `distributions` ∧ predict ∧ correlate over the SHARED set) inside the flaskapi blueprint — surfaced while re-landing the mmux_vite consumption migration: either compute stays in flaskapi (⊥ §G) or the endpoint silently degrades to table correlation (the superseded branch #537 chose the latter, ignoring `distributions`/`num_samples`/`seed` unnoticed)|`api.evaluate_correlations` + `SumoSession.correlations` (log guards via existing `_uq_engine_distributions`) + engine `correlate_manual_uq_samples` w/ REQUIRED `input_scales`/`output_scale`; flip matrix extended over ∀ 11 entry points|V45ls,V22rs,T47mt
29 changes: 29 additions & 0 deletions docs/TIER1_TIER2_UNIT_TESTS_PLAN.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,6 +109,35 @@ Supporting utility used across the workflow above, not a workflow stage on its o

**`DataPreprocessor`** — roundtrip fidelity (see Tier 2, P3/P6)

### 7. Scale semantics helpers (`funs_data_processing.py`, `api/_session.py`)

The `unit→value` map shared by every value producer (SPEC V44ls/V45ls).

**`resolve_log_scale(var, dist_info) -> bool`**
- `log_scale=True` on a non-`uniform` entry → `ValueError` naming the variable
- absent flag → `False` (linear default)

**`scale_distribution(minimum, maximum, *, scale)`**
- `scale` is a REQUIRED keyword — omitting it raises `TypeError` (structural tripwire)
- `"linear"` → `uniform(loc=min, scale=max-min)` (bit-identical to the pre-scale sampler math)
- `"log"` → `loguniform(a=min, b=hi)`: geometric ppf/rvs, geometric mean at the log-midpoint
- `min <= 0` or `max <= min` under log → `ValueError` ("strictly positive and increasing")
- unknown scale string → `ValueError`

**`scale_values(values, *, scale) -> np.ndarray`**
- `"log"` of any `<= 0` → `ValueError("log scale is undefined...")`; otherwise `np.log`
- `"linear"` → unchanged array

**`compute_correlation_indices(..., *, input_scales, output_scale)`**
- scales required (tripwire test); log-reparametrized input shifts Pearson, leaves Spearman bit-identical

**`correlate_manual_uq_samples(...)` → `api.evaluate_correlations`**
- MC→surrogate→correlate workflow (#470) with REQUIRED `input_scales`/`output_scale` (tripwire test); distributions must cover variables exactly; log variable requires uniform ∧ strictly-positive lower bound

**`create_manual_uq_samples`** — `log_scale` uniform drawn log-uniform in original units; log+normal and log+min≤0 refused

**`DataPreprocessor` log transform** — `setup_log_transform`/fit/transform/inverse round-trip; delta-method `inverse_transform_output_std` (`tests/test_data_preprocessor.py::TestLogTransform`)

---

## Tier 2: Property-Based / Invariant Tests
Expand Down
37 changes: 32 additions & 5 deletions docs/verification-validation.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,11 +17,11 @@ Live results as of the last run of the ported V&V suite:

| Suite | Tests | Result |
|---|---|---|
| Full standalone suite (`uv run pytest`) | 245 | **all passing** |
| Analytical/integration tier (`-m analytical`, real Dakota subprocess, no mocking) | 26 | **all passing** |
| Full standalone suite (`uv run pytest`) | 365 | **all passing** |
| Analytical/integration tier (`-m analytical`, real Dakota subprocess, no mocking) | 30 | **all passing** |
| Sobol' / Ishigami acceptance gate (`test_sobol_indices.py`) | 4 | **all passing** |

The analytical tier spawns a real `itis-dakota` (1.5.9 / Dakota 6.20)
The analytical tier spawns a real `itis-dakota` (6.24.7)
process per test — nothing here is mocked. Re-run locally with:

```sh
Expand Down Expand Up @@ -131,6 +131,33 @@ See [Data preprocessing](reference/preprocess.md).
| H1-H3 | Z-score / min-max / sign-switch round-trip to 1e-10 | ✅ |
| H4 | Round-trip on 1000×20 dataset | ✅ |
| H5 | Normalization improves accuracy on badly-scaled `f(x) = 1000x+1` | ✅ |
| H6 | `log_transform` round-trip (`np.log`/`np.exp`), non-positive values refused, delta-method std inverse `std_orig ≈ |y_hat|·std_log` | ✅ |

## Category I: Scale semantics (log)

`scale` is a first-class, un-forgettable axis (SPEC V44ls/V45ls): every public
value-producing entry point accepts it (defaulting to linear, so existing calls
are unchanged) and every value the api yields is computed under it. The one
`unit→value` map behind all of it is `scale_distribution` (`uniform` vs
`scipy.stats.loguniform`), and it requires the scale argument — unwired code
breaks with `TypeError`, it can never silently default.

| Test | What | Status |
|---|---|---|
| I1 | LHS + grid samplers: log domain ⇒ log-uniform/geometric fill; linear default bit-identical; non-positive log domain ⇒ `SumoInputError` | ✅ |
| I2 | Correlations: Pearson moves under log, Spearman provably unchanged (monotone-invariant), untouched columns bit-identical; correlator's scale args are required (`TypeError` tripwire) | ✅ |
| I3 | Surrogate / CV / along-axes / grid eval: log response exp-restored to original units; non-positive training outputs rejected pre-Dakota | ✅ |
| I4 | UQ propagation: log inputs drawn log-uniform (response skews low vs linear, directionally asserted); log+normal / log+min≤0 rejected; log response ⇒ multiplicative (not additive) spread | ✅ |
| I5 | CV accuracy metrics: inherit log through `cross_validate` (metrics differ from linear); reject non-positive log responses | ✅ |
| I6 | MOGA: log variable explored in ln-space (domain mapped, positivity-guarded); log objective exp-restored for **both** minimize and maximize (sign-after-log inverse order verified) | ✅ |
| I7 | Sobol: log input shifts the variance decomposition in the expected direction (compressed variable explains less); mixed log+constant partition | ✅ |
| I8 | Flip matrix: all 11 public value-producing entry points' outputs move when a column turns log — the V45ls machine guard against any silent scale-ignore, shipped or future | ✅ |
| I9 | MC-through-surrogate correlation (`evaluate_correlations`, #470 workflow): dominant variable recovered over the shared sample set, seed-reproducible, log-scale coefficients move, log+non-uniform / log+non-positive / non-covering distributions rejected, engine producer requires its scales (`TypeError` tripwire) | ✅ |

Tests: `tests/test_api_workflows.py` (`TestLogScale*`, `TestScaleGapCoverage`,
`TestScaleAwareSamplers`, `TestScaleFlipMatrix`),
`tests/test_correlation_indices.py`, `tests/test_dakota_funs_data_processing.py`,
`tests/test_data_preprocessor.py`.

## Ishigami analytical acceptance gate

Expand Down Expand Up @@ -161,7 +188,7 @@ first-order-only check and fail this one.
## Known limitations

Carried over from the original test plan, still true of the pinned engine
(`itis-dakota==1.5.9`, Dakota 6.20):
(`itis-dakota==6.24.7`):

- **Built-in Dakota CV parsing is unreliable** — `log_output` comes back
hardcoded empty on some study configurations, which is why
Expand All @@ -185,7 +212,7 @@ Carried over from the original test plan, still true of the pinned engine

## Summary

Every category in the original V&V plan (A through H, plus the Ishigami
Every category in the original V&V plan (A through I, plus the Ishigami
acceptance gate) currently passes against its analytical reference, with
real (unmocked) Dakota execution for every category that requires the
engine. The known limitations above are pre-existing engine/pipeline
Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "uv_build"

[project]
name = "itis-sumo"
version = "0.1.0a5"
version = "0.1.0a6"
description = "Surrogate Modeling functionality for IT'IS Foundation / ZMT Modeling Intelligence suite: build, evaluate, cross-validate surrogates + UQ + sampling, headless or embedded"
readme = "README.md"
# T16mo rung 1 (1.5.9->1.5.11): 1.5.11 ships cp313 wheels, so the <3.13
Expand Down
2 changes: 2 additions & 0 deletions src/itis_sumo/api/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,7 @@
compute_correlations,
cross_validate,
evaluate_along_axes,
evaluate_correlations,
evaluate_cv_metrics,
evaluate_grid,
evaluate_sobol,
Expand Down Expand Up @@ -75,6 +76,7 @@
"compute_correlations",
"cross_validate",
"evaluate_along_axes",
"evaluate_correlations",
"evaluate_cv_metrics",
"evaluate_grid",
"evaluate_sobol",
Expand Down
Loading
Loading