Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 12 additions & 2 deletions SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,8 @@ Consumer contract (grill 2026-08-18): itis-sumo owns ALL surrogate machinery end
- CLI: `itis-sumo validate`
- python (consumer-facing) → `itis_sumo.api` = THE module flaskapi/mmux_vite imports; nothing else. Tabular in, typed results out, taxonomy errors (T22ax)
- python (in-package/headless) → `itis_sumo.{core,config,data,sampling,evaluate,preprocess,utils}` public funcs; `core` now re-exports the E1 model store, `evaluate` now re-exports `export_sumo_model`/`import_sumo_model` (were reachable only via `funs_evaluate`, breaking V16qf)
- python (forthcoming) → `analyze_dataset` narrow diagnostics entrypoint (T18ry) — its output IS the override payload for the preprocessing defaults AND the auto-`domain` source, ⊥ a side feature
- python (diagnostics) → `analyze_dataset` narrow entrypoint (T18ry — unblocked by the T48np/T49np promotion, ships next tag) — its output IS the override payload for the preprocessing defaults AND the auto-`domain` source, ⊥ a side feature
- python (report authoring) → `itis_sumo.report` = THE flat surface automated report notebooks import (V49rp, T54rp): diagnostics ∧ CV validation/calibration ∧ sensitivity/correlation ∧ manual-UQ propagation; re-exports only, ⊥ logic lives there
- artifacts: run_dir `{dakota_stdout.txt,dakota_stderr.txt,*.dat,*.sps/.alg}`; models dir `{id}.metadata.json` (E1)

## §V
Expand Down Expand Up @@ -114,6 +115,12 @@ V45ls: ∀ value-producing `itis_sumo.api` entry point → `scale` ! be threaded
V46rn: log scale ! compose w/ ∀ uncertainty distribution it is defined on: `uniform` → log-uniform draw ∧ `normal` → μ/σ parameterize the ln-space distribution (raw draw = `exp(N(μ,σ))`, lognormal in original units, positive by construction; the surrogate sees exactly `N(μ,σ)` in its training space — same "distribution describes what the model sees" contract as log-uniform); ⊥ a sampler path rejecting log⊗normal or drawing it linear; `constant` ⊥ composes w/ log (rejection stays)
V47st: scale/spec surface input guards ! fail loud at the boundary, never mid-engine: `preprocessing` override maps ! be exact-cover at ∀ entry points (⊥ unknown-column override silently ignored — session path's rule extends to `compute_correlations`/`generate_lhs_samples`/`generate_grid_samples`); required scale maps ! be indexed by required key (⊥ per-variable `.get(..., "linear")` silent default); log-uniform bounds ! validated present ∧ ordered at the api boundary (→ `SumoInputError`, not `_run_engine`'s `SumoEngineError`); ∀ bounded spec (`DomainSpec`/`DistributionSpec`) ! reject `minimum ≥ maximum` at construction
V48tr: Sobol' 2nd-order pairs ! come from the exact joint-pair estimator `S_ij = Var(E[Y|X_i,X_j])/V − S_i − S_j` (independent C stream ∧ U^ij/V^ij mixed designs, cost `n·(2+d+d(d−1))`, EXACT ∀ d); ⊥ the retired B26nc identity that inferred `O(d²)` pairs from `d` first/total gaps (underdetermined d≥4: pair sums collapse→0/negative); ∀ estimator quantity ! use pooled-A/B mean removal ∧ pooled variance ∧ centered products → translation-invariant under `f→f+c` ∧ identical to displayed scipy `saltelli_2010` point estimator; order masses `M1=ΣᵢSᵢ`/`M2=Σ_{i<j}S_ij` (unordered pair once)/`R=1−M1−M2` ⇒ `M1+M2+R=1` by construction (⊥ clamp raw estimates; `ΣᵢS_Ti` ⊥ expected to close); ∀ CIs (first/total/second/masses) ! come from ONE shared-row bootstrap resample recomputing ∀ index+mass per replicate (preserves estimator covariance ∧ closure; ⊥ extra surrogate calls); `V̂=0` (all-constant ∨ degenerate surrogate) ⇒ mass fractions undefined ⇒ `SobolResult.order_contributions=null` (⊥ silent `(0,0,0)`); ∀ finite index/mass value ! validated (`ValueError` on non-finite)
V16wq: convergence diagnostics ! one `compute_cv_diagnostics(actual,predicted)` call per CV subset/draw (rmse+mae+paired-ttest+cohens_d+tukey-outlier-flags); ⊥ second full CV rerun to derive a different metric
V17kb: convergence "done" signal ! MAE (bootstrap CI) + paired-diff effect size (Cohen's d) stabilizing near baseline; p-value evolution reported as diagnostic only, ⊥ used as pass/fail (confounded — power grows w/ N regardless of practical bias)
V18wp: surrogate error/dispersion convergence plots (fig_convergence*) ! physical units (or normalized by data's own std/mean); ⊥ normalized against the final/converged-N value — anchoring to an unknown future answer inflates early-N points, unavailable in a live (non-retrospective) check
V19cz: coverage/calibration ! per-point predictive std (`{output}_std_hat`, V8df) propagated through convergence pipeline; empirical coverage @ nominal 68.27/95/99.7% (1σ/2σ/3σ) computed pooling ALL bootstrap draws per N (⊥ per-draw — too few held-out points for reliable tail-coverage below N≈50); tracked across full N-sweep so flat (non-shrinking) miscalibration is visible, distinct from a converging bias
V46np: raw-value Tukey fence ! computed in the variable's selected scale — log-selected (multiplicative) variables get the GEOMETRIC fence (`exp`-mapped ln-space fence), ⊥ silently applying the arithmetic linear fence to them (it over-flags the benign right tail of lognormal data); the affine log-basis change (ln↔log10) ! move zero flags ∧ fences stay reported in the values' own units — flip-asserted, mirroring V45ls' rank-correlation exemption
V49rp: automated report notebooks ! authored against the flat `itis_sumo.report` surface only; ⊥ reaching into `itis_sumo.{api,data,evaluate}` internals directly — the report-authoring mirror of V16qf's consumer rule (façade completeness test-enforced, `tests/test_report_surface.py`)

## §R
R1: `export_model`/`import_model` child keywords; formats `text_archive`(.sps)/`binary_archive`(.bsps)/`algebraic_file`(.alg); naming `{prefix}.{resp}.{ext}` | branch R2
Expand Down Expand Up @@ -145,7 +152,7 @@ T15mn|~|wire mmux/vite flaskapi to `itis_sumo.api` (export/import + all workflow
T16mo|✓|stepwise engine modernization: rung 1 `1.5.9→1.5.11` ✓ (Dakota 6.20); rung 2 `1.5.11→6.24.7` ✓ (Dakota 6.24 engine bump). Took 6.24.7 — the latest 6.24.x still shipping cp311 wheels — to keep Python 3.11 parity with mmux/vite; 6.24.8+ dropped cp311 (would force a 3.12 migration, deferred). The 6.23+ `interface_cache` regression (R5) was NOT fixed upstream in 6.24.7, so itis-sumo shims it via `add_r5_interface_cache_workaround()` + `truth_model_pointer='R5_WA_MODEL'` in config/funs_create_dakota_conf.py: a no-op `fork` interface (`analysis_drivers='true'`) on an otherwise-unused `single` model, so the `study()` ctor's cache lookup succeeds; the interface is never executed. JSON input seam (R4) deferred per scope. 318-test suite green|R4,R5
T17bq|✓|`add_surrogate_model` (`config/funs_create_dakota_conf.py`) ! replaced filename-substring sniffing w/ explicit `has_eval_id_column` param (infer-or-override idiom); all 5 `create_sumo_*_conffile` call sites updated (landed via `feat/vv-port-tests-and-docs`, `a6d694e`, ahead of this merge)|V15,B2
T17bc|✓|E1: persist real training data on export for reference (`{id}.processed_training.dat` sidecar copy, user preference over the leaner archive-only design); `stage_model_for_import` stages it back, falling back to a synthesized header-only placeholder + loud warning log only when that stored copy is missing (legacy model / deleted out-of-band) — fallback verified safe vs Dakota source + fake-points empirical tests (header reorder, wrong/missing descriptors, 0/1/2-row placeholders)|R8,R9,V10,V11
T18ry|.|design+implement `analyze_dataset(df, input_cols, output_cols, alpha=0.05, include_detail=False) -> DatasetDiagnostics` narrow entrypoint (scale/distribution auto-selection + outlier surfacing, plain dataclasses, JSON-serializable via `dataclasses.asdict()`) as the sole flaskapi-facing dataset-diagnostics API, replacing any per-function flaskapi orchestration; BLOCKED — building blocks `select_variable_scale`/`auto_select_distributions` (+ a new raw-value outlier detector, reusing `_tukey_outlier_mask`'s IQR technique) currently exist only on confidential incubator branch `feat/nih-in-silico-example`, not `develop` — needs individual promotion first, same promotion rule as T15mn/E1|§C,V16qf,I
T18ry|x|design+implement `analyze_dataset(df, input_cols, output_cols, alpha=0.05, include_detail=False) -> DatasetDiagnostics` narrow entrypoint (scale/distribution auto-selection + outlier surfacing, plain dataclasses, JSON-serializable via `dataclasses.asdict()`) as the sole flaskapi-facing dataset-diagnostics API, replacing any per-function flaskapi orchestration; unblocked by T48np's promotion of `select_variable_scale`/`auto_select_distributions` + T49np's new raw-value outlier detector|§C,V16qf,I
T19kp|.|headless notebook, post-alpha — add runnable notebook counterpart to docs getting-started flow once alpha docs publish is stable|T10le
T20hm|.|CI workflow changes ! activate shared validation jobs and regression-check detector classification|V18rs
T21vk|✓|VOCAB section (above) + `docs/reference/glossary.md` (nav-registered, strict build green); vocabulary enforced on the api layer by `tests/test_api_contract.py::TestPublicSurface`. Sweep of the OLDER modules' docstrings still outstanding → folded into T24cm|V19cn
Expand Down Expand Up @@ -181,6 +188,9 @@ T50vb|x|`DomainSpec`/`DistributionSpec` ! reject `minimum ≥ maximum` at constr
T51pq|x|port mmux_vite T31rb (PR#649): exact arbitrary-d 2nd-order Sobol' pair estimator + order masses — `_saltelli_abc` 3-stream design builder ∧ `_saltelli_pair_designs` (U/V mixed designs), `_sobol_algebra` (scipy `saltelli_2010` parity pinned by test, translation-invariant), `_sobol_joint_bootstrap` (shared-row CIs over ∀ index+mass), zero-variance → null masses; replaces the V9gh closed-form identity + the `d==1` special case + the runtime `scipy.stats.sobol_indices` call; public surface gains `OrderMasses` dataclass ∧ `SobolResult.order_contributions`; analytic benchmarks additive d=8 / pair-interaction d=5 / Ishigami / pair-quadratic d=10 drive the PRODUCTION helpers with 3·CI-scaled tolerances; B26nc degeneracy regression kept|V48tr,V9gh,V22rs,B20qt
T52xx|x|land V26dd domain⊥distribution split (ahead of the handle transformation, per owner call 2026-09-30 — no defer): `evaluate_sobol`/`SumoSession.sobol` take optional `domains: Mapping[str, DomainSpec]` (unknown names → `SumoInputError`; boxes auto-inferred from observed bounds + detected scale; columns constant in the samples pinned ∧ echoed in `fixed`); Saltelli draws become uniform/log-uniform over the box (V44ls) — the `distributions` parameter, log⊗normal Sobol sampling ∧ the `mean±3σ` conflation die; engine `evaluate_sobol_indices` consumes a domain-`sampling` map (`{"minimum","maximum","log_scale"?}` | `{"value"}`); `SobolResult`: `distributions` → `domains` + `fixed` + `effective_config` (V21pf); NORMAL flip matrix ⊥ Sobol row (its flip rides DOMAIN boxes, uniform matrix); docs/examples/V&V aligned|V26dd,V44ls,V45ls,V21pf,V47st
T53xx|x|explicit factor pinning for Sobol: `evaluate_sobol`/`SumoSession.sobol` accept `fixed: Mapping[str, float]` (unknown names rejected ∧ boxed∧fixed overlap rejected ∧ finite ∧ log-positive); pins echo in `SobolResult.fixed` beside auto-inferred constants — the domain-vocabulary freeze a consumer needs to hold a factor without distribution parameters (mmux FE Sobol-panel migration, prelude); test: pinned factor zero indices ∧ surrogate keeps its real spread|V26dd,V47st
T48np|x|port the NIH incubator's still-unpromoted diagnostics tooling onto this line (`CONFIDENTIAL/nih-merck-validation` → here, clean re-port not rebase): data `select_variable_scale`/`auto_select_distributions`; evaluate `compute_coverage`/`compute_cv_diagnostics`/`fit_convergence_exponential{,_asymptotic}` + bootstrap/`draws`/`diagnostics` redesign of `compute_cv_convergence` + `mean_signed_error`; NIH test files ride along; NIH's duplicate `propagate_uq_with_uncertainty` ⊥ ported — `propagate_manual_uq_with_uncertainty` made its superset instead (optional-`preprocessor` None-branch), callers adapt; api `CVAccuracyMetrics` absorbs the new `mean_signed_error` field; `ty`-clean (heterogeneous result dicts `dict[str, Any]` per house precedent)|V16wq,V17kb,V18wp,V19cz
T49np|x|new raw-value outlier detector `detect_raw_outliers(values, scale=, k=)` in `data/funs_dataset_diagnostics.py` (Tukey fence computed in the variable's selected scale, flag-only, NaN-safe) + wire `analyze_dataset` for real (scale/distribution via `auto_select_distributions`, outliers in each variable's own scale, `include_detail` candidates, JSON round-trip) — closes T18ry and the "raw-value outlier detection does not exist yet, anywhere" gap in BRANCH_CONSOLIDATION §3; flip-tests: geometric-vs-arithmetic over-flagging guard + ln↔log10 zero-flag-move|V46np,V16qf,T18ry
T54rp|x|new `itis_sumo.report` façade (V49rp): flat curated surface for automated report notebooks — api workflow verbs (cross_validate/evaluate_sobol/evaluate_correlations/…) ∧ `analyze_dataset` ∧ the T48np diagnostics (coverage, CV bundle, convergence bootstrap/fits) ∧ `propagate_manual_uq_with_uncertainty`; re-exports only, ⊥ logic; contract test pins completeness + object identity so report code never reaches into engine internals|V49rp,V16qf

## §B
id|date|cause|fix
Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "uv_build"

[project]
name = "itis-sumo"
version = "0.1.0a10"
version = "0.1.0a11"
description = "Surrogate Modeling functionality for IT'IS Foundation / ZMT Modeling Intelligence suite: build, evaluate, cross-validate surrogates + UQ + sampling, headless or embedded"
readme = "README.md"
# T16mo rung 1 (1.5.9->1.5.11): 1.5.11 ships cp313 wheels, so the <3.13
Expand Down
1 change: 1 addition & 0 deletions src/itis_sumo/api/types.py
Original file line number Diff line number Diff line change
Expand Up @@ -228,6 +228,7 @@ class CVAccuracyMetrics:
sum_abs: float
mean_abs: float
max_abs: float
mean_signed_error: float
seed: int


Expand Down
6 changes: 6 additions & 0 deletions src/itis_sumo/data/__init__.py
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
"""Training-data parsing, filtering, sampling grids, and statistics."""

from itis_sumo.data.funs_data_processing import (
auto_select_distributions,
compute_correlation_indices,
create_grid_samples,
create_manual_uq_samples,
Expand All @@ -16,23 +17,27 @@
load_data,
process_input_file,
sanitize_varnames,
select_variable_scale,
)
from itis_sumo.data.funs_dataset_diagnostics import (
DatasetDiagnostics,
OutlierSummary,
VariableDiagnostics,
analyze_dataset,
detect_raw_outliers,
)

__all__ = [
"DatasetDiagnostics",
"OutlierSummary",
"VariableDiagnostics",
"analyze_dataset",
"auto_select_distributions",
"compute_correlation_indices",
"create_grid_samples",
"create_manual_uq_samples",
"create_samples_along_axes",
"detect_raw_outliers",
"extract_predictions_along_axes",
"extract_predictions_gridpoints",
"get_bounds_uniform_distribution",
Expand All @@ -44,4 +49,5 @@
"load_data",
"process_input_file",
"sanitize_varnames",
"select_variable_scale",
]
Loading
Loading