From f876b04252d257af92219b2c0ae0185074eae726 Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 22:33:00 +0200 Subject: [PATCH 01/13] feat(api): support log-scale variables in the itis-sumo backend (T27fr) Port the mmux_vite fullstack-logscale backend transforms into the library. DataPreprocessor (V21pf/V22rs): - add log_transform to VariableConfig + setup_log_transform/_configure_log_transform - natural-log applied in _fit/_transform_variable_group, exp restored in inverse_transform (both input and output groups) - inverse_transform_output_std: delta-method std_orig ~= |y_hat_orig| * std_log, with normalization-aware scaling; falls back unchanged + warns without points api._session: - honour PreprocessingSpec scale: drop the "log not supported yet" rejection, drive setup_log_transform from the effective per-column scale in fit() - route predicted std through the delta-method inverse in cross_validate and along_axes (linear path is a plain remap, unchanged) - reject non-positive log columns and non-positive held log values as SumoInputError before Dakota runs (V23er taxonomy, not the raw ValueError) Tests: preprocessor round-trip / mixed-scale / positivity / delta-method (unit) + real-Dakota cross_validate/along_axes/grid original-space units and the rejection guards (integration); replace the obsolete "unsupported scale" api-contract test with honour + reject cases. 332 passed, ruff/ty clean. T27fr stays ~: UQ/MOGA log paths inherit via the preprocessor but are untested here, and the domain/distribution config split is the other half. --- SPEC.md | 2 +- src/itis_sumo/api/_session.py | 105 ++++++++++++--- src/itis_sumo/preprocess/data_preprocessor.py | 125 ++++++++++++++++++ tests/test_api_contract.py | 18 ++- tests/test_api_workflows.py | 82 ++++++++++++ tests/test_data_preprocessor.py | 66 +++++++++ uv.lock | 2 +- 7 files changed, 375 insertions(+), 25 deletions(-) diff --git a/SPEC.md b/SPEC.md index d3a2ef3..e392723 100644 --- a/SPEC.md +++ b/SPEC.md @@ -149,7 +149,7 @@ T23bn|✓|`itis_sumo.api`: `cross_validate()` + `evaluate_along_axes()` — inte T24cm|✓|remove `preprocess/models.py::{FunctionJob,JobVariableSelection}`; re-express the Dakota sufficiency rule over tabular data; jobs→table adapter moves to flaskapi (this port DELETES itis-sumo code)|V20dm T25dp|.|`itis_sumo.api`: remaining 6 workflows (UQ-w-uncertainty incl. the ~120-line erfinv/histogram block, correlation, Sobol, grid, MOGA, cv-accuracy-metrics) + E1 `export_model`/`evaluate_stored_model` facade|V22rs,R1-R4 T26eq|.|POST-PORT: extract the fitted-model handle (`fit()` → methods → `save()`/`load()`); carries the fitted preprocessing config ⇒ closes the E1 gap (model store persists archive+metadata+training copy but ⊥ preprocessor config, so a reloaded model cannot inverse-transform to original units)|V27fq,V10jk -T27fr|.|POST-PORT: split `domain` vs `distribution` config + consumer migration; absorb the mmux_vite `jgo/fullstack-logscale` work|V26dd +T27fr|~|POST-PORT: split `domain` vs `distribution` config + consumer migration; absorb the mmux_vite `jgo/fullstack-logscale` work|V26dd T28gs|.|POST-PORT `?`: decide whether `distribution` gets an auto-generated default — explicit discussion required, ⊥ silently defaulted|§C `?` T29hw|~|`publish.yml` tag trigger accepts `v`-prefixed PEP 440 prereleases ✓; tagging `v0.1.0a1` BLOCKED on the one-time PyPI Trusted Publisher config (user action), then clean-venv install + `itis-sumo validate` + headless smoke|T1pw T30qa|✓|release/CI workflow refresh: build→TestPyPI automatic on tag push (✓), PyPI+Release gated manual (✓, T35cc); auto-tag-on-branch redesigned to PR-time check + tag-only-at-merge (no bot commit) — see T31xx/T32yy; keep git-cliff release notes, dependency-review + concurrency (✓); skip weekly cron/healthchecks for now|§C,V17rt diff --git a/src/itis_sumo/api/_session.py b/src/itis_sumo/api/_session.py index 2b85004..68f3646 100644 --- a/src/itis_sumo/api/_session.py +++ b/src/itis_sumo/api/_session.py @@ -127,14 +127,6 @@ def _validate_samples( f"Preprocessing overrides given for columns that are not in play: " f"{unknown_overrides}" ) - logarithmic = sorted( - name for name, override in spec.overrides.items() if override.scale == "log" - ) - if logarithmic: - raise SumoInputError( - f"Logarithmic scale is not supported yet (requested for {logarithmic}); " - "it arrives together with the domain/distribution split" - ) selected = samples.loc[:, columns].copy() try: @@ -155,6 +147,20 @@ def _validate_samples( "Incomplete samples must be filtered out before they are passed in" ) + # A log-scale column is trained as log(value); log is undefined at or below + # zero, so rejecting a non-positive sample here (as a SumoInputError the + # consumer can surface) preempts the preprocessor's raw ValueError deeper in. + non_positive = sorted( + name + for name, override in spec.overrides.items() + if override.scale == "log" and (selected[name].to_numpy() <= 0).any() + ) + if non_positive: + raise SumoInputError( + f"Columns {non_positive} are marked log-scale but hold values <= 0, " + "for which the logarithm is undefined" + ) + minimum = _minimum_samples(len(variables)) if len(selected) < minimum: raise SumoInputError( @@ -227,6 +233,18 @@ def fit(self) -> Self: preprocessor.setup_variables( input_vars=list(self._variables), output_vars=[self._response] ) + preprocessor.setup_log_transform( + input_log_vars=[ + variable + for variable in self._variables + if self._scale_of(variable) == "log" + ], + output_log_vars=[ + response + for response in [self._response] + if self._scale_of(response) == "log" + ], + ) transformed = preprocessor.fit_transform(self._samples) assert self._run_dir is not None @@ -265,12 +283,17 @@ def cross_validate(self, *, folds: int, seed: int) -> CrossValidationResult: "with uncertainty estimates" ) + predicted_original = self._to_original_units( + response, results[f"{response}_hat"] + ) return CrossValidationResult( response=self._response, observed=self._to_original_units(response, results[response]), - predicted=self._to_original_units(response, results[f"{response}_hat"]), - predicted_std=self._to_original_units( - response, results[f"{response}_std_hat"] + predicted=predicted_original, + predicted_std=self._to_original_std( + response, + results[f"{response}_std_hat"], + {self._response: predicted_original}, ), warnings=list(results.get("warnings", [])), seed=seed, @@ -304,17 +327,22 @@ def along_axes( sweeps: dict[str, AxisSweep] = {} for mapped_variable, axis in results.items(): variable = original_names.get(mapped_variable, mapped_variable) + predicted = self._to_original_units(response, axis["y_hat"]) + # A standard deviation is a width, not a position: it goes through the + # delta-method std inverse (a no-op remap unless this response is log + # scale), not the plain point inverse applied to ``predicted``. + predicted_std = ( + self._to_original_std( + response, axis["std_hat"], {self._response: predicted} + ) + if "std_hat" in axis + else None + ) sweeps[variable] = AxisSweep( variable=variable, x=self._to_original_units(mapped_variable, axis["x"]), - predicted=self._to_original_units(response, axis["y_hat"]), - # A standard deviation is a width, not a position: it is reported - # as produced rather than shifted back through the transform. - predicted_std=( - [float(value) for value in axis["std_hat"]] - if "std_hat" in axis - else None - ), + predicted=predicted, + predicted_std=predicted_std, ) return AlongAxesResult( @@ -472,6 +500,38 @@ def uncertainty( maximum=summary["max"], ) + def _scale_of(self, column: str) -> str: + """The scale the caller asked for on ``column`` (default: linear).""" + return self._spec.overrides.get(column, VariableSpec()).scale + + def _to_original_std( + self, + mapped_name: str, + std_values: Sequence[float], + point_estimates_original: Mapping[str, Sequence[float]], + ) -> list[float]: + """Restore a predicted *standard deviation* to original units. + + A std is a width, not a position, so it cannot go through the ordinary + point inverse-transform. A log-scale response in particular needs the + multiplicative delta-method rule, which is why the point estimates (already + back in original units) are threaded in alongside. For every other column + this is a plain name remap, matching the pre-log behaviour exactly. + """ + assert self._preprocessor is not None + original = self._preprocessor.get_inverse_mapping().get( + mapped_name, mapped_name + ) + points = { + name: [float(value) for value in values] + for name, values in point_estimates_original.items() + } + restored = self._preprocessor.inverse_transform_output_std( + {mapped_name: [float(value) for value in std_values]}, + point_estimates_original=points, + ) + return [float(value) for value in restored.get(original, list(std_values))] + def _mapped_name(self, variable: str) -> str: assert self._preprocessor is not None return self._preprocessor.input_variables[variable].mapped_name @@ -530,6 +590,13 @@ def _map_held_values( raise SumoInputError( f"Cannot hold {unknown} fixed: they are not variables of this model" ) + bad = sorted( + name + for name, value in at.items() + if self._scale_of(name) == "log" and float(value) <= 0 + ) + if bad: + raise SumoInputError(f"Cannot hold log-scale {bad} fixed at a value <= 0") assert self._preprocessor is not None held_row = { **self._samples.mean().to_dict(), diff --git a/src/itis_sumo/preprocess/data_preprocessor.py b/src/itis_sumo/preprocess/data_preprocessor.py index 9bf43c1..95dd4dc 100644 --- a/src/itis_sumo/preprocess/data_preprocessor.py +++ b/src/itis_sumo/preprocess/data_preprocessor.py @@ -33,6 +33,7 @@ class VariableConfig: mapped_name: str normalize: bool = False switch_sign: bool = False + log_transform: bool = False mean: float | None = None std: float | None = None min_val: float | None = None @@ -132,6 +133,47 @@ def setup_sign_switching( f"Configured sign switching for {len(input_sign_switches)} input and {len(output_sign_switches)} output variables" ) + def setup_log_transform( + self, + input_log_vars: list[str] | None = None, + output_log_vars: list[str] | None = None, + ) -> None: + """ + Configure natural-log transformation for variables. + + The surrogate then trains on ``log(value)`` and predictions are mapped + back through ``exp`` on the way out (with the delta-method rule applied + to any accompanying standard deviation -- see + :meth:`inverse_transform_output_std`). + + Args: + input_log_vars: Input vars the surrogate should train on log(value) + output_log_vars: Output vars the surrogate should train on log(value) + """ + input_log_vars = input_log_vars or [] + output_log_vars = output_log_vars or [] + + self._configure_log_transform(self.input_variables, input_log_vars, "input") + self._configure_log_transform(self.output_variables, output_log_vars, "output") + + _logger.info( + f"Configured log-transform for {len(input_log_vars)} input and {len(output_log_vars)} output variables" + ) + + def _configure_log_transform( + self, + variables: dict[str, VariableConfig], + log_vars: list[str], + var_type: str, + ) -> None: + for var_name in log_vars: + if var_name in variables: + variables[var_name].log_transform = True + else: + _logger.warning( + f"{var_type.capitalize()} variable {var_name} not found in setup variables" + ) + def _setup_variable_group( self, var_names: list[str], prefix: str ) -> dict[str, VariableConfig]: @@ -193,6 +235,15 @@ def _fit_variable_group( continue values = np.array(data[var_name].values, dtype=float) + + if config.log_transform: + if np.any(values <= 0): + raise ValueError( + f"Cannot apply log_transform to {var_type} variable " + f"'{var_name}': log is undefined for values <= 0" + ) + values = np.log(values) + if config.switch_sign: values = -values @@ -220,6 +271,14 @@ def _transform_variable_group( values = np.array(data[var_name].values, dtype=float).copy() + if config.log_transform: + if np.any(values <= 0): + raise ValueError( + f"Cannot apply log_transform to {var_type} variable " + f"'{var_name}': log is undefined for values <= 0" + ) + values = np.log(values) + if config.switch_sign: values = -values @@ -344,6 +403,9 @@ def inverse_transform( if not isinstance(value, list): value = [value] + if config.log_transform: + value = np.exp(np.array(value, dtype=float)).tolist() + result[var_name] = value for var_name, config in self.output_variables.items(): @@ -362,10 +424,71 @@ def inverse_transform( if not isinstance(value, list): value = [value] + if config.log_transform: + value = np.exp(np.array(value, dtype=float)).tolist() + result[var_name] = value return result + def inverse_transform_output_std( + self, + std_data: dict[str, list[float]], + point_estimates_original: dict[str, list[float]] | None = None, + ) -> dict[str, list[float]]: + """ + Inverse-transform output *uncertainty* (standard deviation) values. + + A standard deviation cannot be undone the way a point value can -- mean + shifting a std is meaningless. This applies the correct multiplicative + rule for each configured transform: + + - ``z_score``: ``std_orig = std_mapped * std`` (no mean shift) + - ``min_max``: ``std_orig = std_mapped * (max - min)`` + - ``log_transform``: ``std_orig ~= y_hat_orig * std_log`` (delta-method / + natural-log approximation ``d(exp(x))/dx = exp(x)``, evaluated at the + point estimate). Requires the already-inverse-transformed point estimate + for that variable in ``point_estimates_original``; when it is missing the + log-space std is returned unchanged (with a warning) rather than silently + producing a wrong value. + + ``switch_sign`` does not change the magnitude of a std. + """ + point_estimates_original = point_estimates_original or {} + result: dict[str, list[float]] = {} + + for var_name, config in self.output_variables.items(): + if config.mapped_name not in std_data: + continue + + values = np.array(std_data[config.mapped_name], dtype=float) + + if config.normalize and config.normalization_method == "z_score": + if config.std is not None: + values = values * config.std + elif ( + config.normalize + and config.normalization_method == "min_max" + and config.max_val is not None + and config.min_val is not None + ): + values = values * (config.max_val - config.min_val) + + if config.log_transform: + point_est = point_estimates_original.get(var_name) + if point_est is not None and len(point_est) == len(values): + values = values * np.abs(np.array(point_est, dtype=float)) + else: + _logger.warning( + f"Cannot delta-method inverse-transform std for log-scale " + f"output '{var_name}' without matching point estimates; " + "returning log-space std unchanged" + ) + + result[var_name] = values.tolist() + + return result + def _normalize_values( self, values: np.ndarray, config: VariableConfig ) -> np.ndarray: @@ -648,6 +771,7 @@ def get_summary(self) -> dict[str, Any]: "normalize": config.normalize, "normalization_method": config.normalization_method, "switch_sign": config.switch_sign, + "log_transform": config.log_transform, } for name, config in self.output_variables.items(): @@ -656,6 +780,7 @@ def get_summary(self) -> dict[str, Any]: "normalize": config.normalize, "normalization_method": config.normalization_method, "switch_sign": config.switch_sign, + "log_transform": config.log_transform, } return summary diff --git a/tests/test_api_contract.py b/tests/test_api_contract.py index 539daeb..2605937 100644 --- a/tests/test_api_contract.py +++ b/tests/test_api_contract.py @@ -218,11 +218,21 @@ def test_rejects_overrides_for_columns_not_in_play(self): with pytest.raises(SumoInputError, match="not in play"): cross_validate(make_samples(), VARIABLES, RESPONSE, preprocessing=spec) - def test_reports_unsupported_scale_instead_of_ignoring_it(self): - """An override that cannot be honoured must fail loudly (SPEC V21pf).""" + def test_honours_a_log_scale_override(self): + """A log-scale override is now applied, not rejected, and is surfaced as + the domain-level ``scale`` (never a transform name) in the result (V21pf).""" spec = PreprocessingSpec(overrides={"width": VariableSpec(scale="log")}) - with pytest.raises(SumoInputError, match="Logarithmic"): - cross_validate(make_samples(), VARIABLES, RESPONSE, preprocessing=spec) + result = cross_validate(make_samples(), VARIABLES, RESPONSE, preprocessing=spec) + assert result.effective_config["width"].scale == "log" + + def test_rejects_a_non_positive_log_scale_column(self): + """A log-scale column that holds a non-positive value fails loudly with a + SumoInputError, before any engine work (log is undefined at <= 0).""" + samples = make_samples() + samples.loc[samples.index[0], "width"] = 0.0 + spec = PreprocessingSpec(overrides={"width": VariableSpec(scale="log")}) + with pytest.raises(SumoInputError, match="log-scale"): + cross_validate(samples, VARIABLES, RESPONSE, preprocessing=spec) def test_rejects_a_single_fold(self): with pytest.raises(SumoInputError, match="at least 2 folds"): diff --git a/tests/test_api_workflows.py b/tests/test_api_workflows.py index 0fd6410..98a4d50 100644 --- a/tests/test_api_workflows.py +++ b/tests/test_api_workflows.py @@ -20,7 +20,9 @@ from itis_sumo.api import ( DistributionSpec, DomainSpec, + PreprocessingSpec, SumoInputError, + VariableSpec, compute_correlations, cross_validate, evaluate_along_axes, @@ -261,3 +263,83 @@ def test_requires_a_domain_for_each_variable(self, samples): domains={"width": DomainSpec(minimum=1.0, maximum=5.0)}, max_evaluations=200, ) + + +_LOG_SCALE = PreprocessingSpec(overrides={RESPONSE: VariableSpec(scale="log")}) + + +class TestLogScale: + """The domain-level ``scale`` flag must reach the surrogate and come back + in the caller's own units -- SPEC V21pf (scale in, no transform in the + signature) and the port of the mmux_vite log-scale backend (T27fr).""" + + def test_log_scale_response_comes_back_in_original_units(self, samples): + result = cross_validate(samples, VARIABLES, RESPONSE, preprocessing=_LOG_SCALE) + + assert result.effective_config[RESPONSE].scale == "log" + predicted = [v for v in result.predicted if not math.isnan(v)] + assert predicted, "no fold produced a prediction" + # Log-space predictions must be exp-restored: same order of magnitude as + # the raw stress (which is O(10)), not the O(1..3) log values. + assert min(predicted) > 0.0 + assert max(predicted) < 10 * samples[RESPONSE].max() + observed_mean = float(np.mean(samples[RESPONSE])) + assert min(predicted) < 3 * observed_mean < 10 * max(predicted) + + def test_log_scale_survives_from_request_to_engine(self, samples): + linear = cross_validate(samples, VARIABLES, RESPONSE) + log = cross_validate(samples, VARIABLES, RESPONSE, preprocessing=_LOG_SCALE) + # Both agree roughly with the observed stress -- the transform is internal. + assert np.mean(log.predicted) == pytest.approx( + np.mean(linear.predicted), rel=0.5, nan_ok=True + ) + + def test_along_axes_log_response_returns_original_units(self, samples): + result = evaluate_along_axes( + samples, + VARIABLES, + RESPONSE, + preprocessing=_LOG_SCALE, + points_per_variable=5, + ) + assert result.effective_config[RESPONSE].scale == "log" + for sweep in result.sweeps.values(): + assert len(sweep.x) == len(sweep.predicted) + assert min(sweep.predicted) > 0.0 + assert max(sweep.predicted) < 10 * samples[RESPONSE].max() + + def test_grid_log_response_returns_original_units(self, samples): + result = evaluate_grid( + samples, + VARIABLES, + RESPONSE, + grid_variables=["width", "height"], + preprocessing=_LOG_SCALE, + points_per_variable=5, + ) + flat = [v for row in result.data["stress"] for v in row] + assert min(flat) > 0.0 + assert max(flat) < 10 * samples[RESPONSE].max() + + def test_non_positive_log_response_is_rejected_before_dakota(self): + bad = pd.DataFrame( + { + "width": [1.0, 2.0, 3.0, 4.0, 5.0], + "height": [100.0, 200.0, 300.0, 400.0, 500.0], + "stress": [5.0, 8.0, 0.0, 12.0, 15.0], # <= 0 under log + } + ) + with pytest.raises(SumoInputError, match="log-scale but hold values"): + cross_validate(bad, VARIABLES, RESPONSE, preprocessing=_LOG_SCALE) + + def test_holding_a_log_variable_at_zero_is_rejected(self, samples): + log_input = PreprocessingSpec(overrides={"height": VariableSpec(scale="log")}) + with pytest.raises(SumoInputError, match="log-scale .* fixed at a value"): + evaluate_along_axes( + samples, + VARIABLES, + RESPONSE, + at={"height": 0.0}, + preprocessing=log_input, + points_per_variable=5, + ) diff --git a/tests/test_data_preprocessor.py b/tests/test_data_preprocessor.py index 723bd2b..e6a1faa 100644 --- a/tests/test_data_preprocessor.py +++ b/tests/test_data_preprocessor.py @@ -1,4 +1,6 @@ +import numpy as np import pandas as pd +import pytest from itis_sumo.preprocess.data_preprocessor import DataPreprocessor @@ -82,3 +84,67 @@ def test_save_and_load_config_round_trip(self, tmp_path): assert loaded.get_variable_mapping() == {"alpha": "x1", "beta": "y1"} assert loaded.output_variables["beta"].mean == 5.0 + + +class TestLogTransform: + def test_log_output_round_trips_to_original_space(self): + dataframe = pd.DataFrame( + {"length": [1.0, 2.0, 3.0], "stress": [10.0, 100.0, 1000.0]} + ) + preprocessor = DataPreprocessor() + preprocessor.setup_variables(["length"], ["stress"]) + preprocessor.setup_log_transform(output_log_vars=["stress"]) + + transformed = preprocessor.fit_transform(dataframe) + restored = preprocessor.inverse_transform(transformed) + + # The surrogate sees the natural log of the response, not the raw value. + assert transformed["y1"].tolist() == pytest.approx( + np.log([10.0, 100.0, 1000.0]) + ) + # And the caller gets their own units back. + assert restored["stress"] == pytest.approx([10.0, 100.0, 1000.0]) + + def test_mixed_scales_keep_each_column_in_its_own_space(self): + dataframe = pd.DataFrame( + {"lin": [2.0, 4.0], "log": [2.0, 4.0], "y": [1.0, 1.0]} + ) + preprocessor = DataPreprocessor() + preprocessor.setup_variables(["lin", "log"], ["y"]) + preprocessor.setup_log_transform(input_log_vars=["log"]) + + transformed = preprocessor.fit_transform(dataframe) + + assert transformed["x1"].tolist() == pytest.approx([2.0, 4.0]) + assert transformed["x2"].tolist() == pytest.approx(np.log([2.0, 4.0])) + + def test_fit_rejects_non_positive_values_for_a_log_column(self): + preprocessor = DataPreprocessor() + preprocessor.setup_variables(["x"], ["y"]) + preprocessor.setup_log_transform(input_log_vars=["x"]) + with pytest.raises(ValueError, match="log is undefined"): + preprocessor.fit(pd.DataFrame({"x": [1.0, 0.0, 3.0], "y": [1.0, 1.0, 1.0]})) + + def test_inverse_transform_output_std_uses_the_delta_method(self): + preprocessor = DataPreprocessor() + preprocessor.setup_variables(["x"], ["y"]) + preprocessor.setup_log_transform(output_log_vars=["y"]) + preprocessor.fit(pd.DataFrame({"x": [1.0, 2.0, 3.0], "y": [10.0, 20.0, 40.0]})) + + # std_orig ~= |point_orig| * std_log + restored = preprocessor.inverse_transform_output_std( + {"y1": [0.1, 0.2, 0.15]}, + point_estimates_original={"y": [10.0, 20.0, 40.0]}, + ) + assert restored["y"] == pytest.approx([1.0, 4.0, 6.0]) + + def test_inverse_transform_output_std_without_points_leaves_std_unchanged( + self, caplog + ): + preprocessor = DataPreprocessor() + preprocessor.setup_variables(["x"], ["y"]) + preprocessor.setup_log_transform(output_log_vars=["y"]) + preprocessor.fit(pd.DataFrame({"x": [1.0, 2.0], "y": [10.0, 20.0]})) + + restored = preprocessor.inverse_transform_output_std({"y1": [0.1, 0.2]}) + assert restored["y"] == pytest.approx([0.1, 0.2]) diff --git a/uv.lock b/uv.lock index 30b80a8..c84e42c 100644 --- a/uv.lock +++ b/uv.lock @@ -458,7 +458,7 @@ wheels = [ [[package]] name = "itis-sumo" -version = "0.1.0a4" +version = "0.1.0a5" source = { editable = "." } dependencies = [ { name = "itis-dakota" }, From fc93cbeb83a5cfb54c04513f150e55e590463d35 Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 22:41:24 +0200 Subject: [PATCH 02/13] feat(api): log-scale UQ manual sampling + CV metrics (T27fr) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Log now reaches the uncertainty sampler and accuracy metrics, not just the surrogate fit. - data.create_manual_uq_samples: add the log_scale branch — a log-scale uniform is drawn uniformly in log10 space and returned in the caller's original units (log-uniform); log_scale with a non-uniform distribution or a non-positive bound raises. Faithful to the mmux_vite source. - api._session._uq_engine_distributions: flag log-scale variables as log_scale for the sampler, and reject log-scale + non-uniform / non-positive-lower-bound at the boundary as SumoInputError (V23er) instead of the sampler's raw ValueError. - evaluate_cv_metrics already composes cross_validate (now log-aware), so a log-scale response flows to the metrics through the preprocessor. Tests: create_manual_uq_samples log-uniform / reject cases (unit) + real-Dakota evaluate_uncertainty log input skews response low, log+normal and log+min<=0 rejections, and cv_metrics honouring / rejecting log (integration). --- src/itis_sumo/api/_session.py | 32 +++++++- src/itis_sumo/data/funs_data_processing.py | 33 +++++++- tests/test_api_workflows.py | 90 ++++++++++++++++++++++ tests/test_dakota_funs_data_processing.py | 51 ++++++++++++ 4 files changed, 200 insertions(+), 6 deletions(-) diff --git a/src/itis_sumo/api/_session.py b/src/itis_sumo/api/_session.py index 68f3646..e8fd117 100644 --- a/src/itis_sumo/api/_session.py +++ b/src/itis_sumo/api/_session.py @@ -463,9 +463,7 @@ def uncertainty( f"Distributions must cover variables exactly; missing={missing}, " f"unknown={unknown}" ) - engine_distributions = { - variable: spec.as_engine_dict() for variable, spec in distributions.items() - } + engine_distributions = self._uq_engine_distributions(distributions) samples = self._run_engine( "propagating uncertainty", propagate_manual_uq_with_uncertainty, @@ -500,6 +498,34 @@ def uncertainty( maximum=summary["max"], ) + def _uq_engine_distributions( + self, distributions: Mapping[str, DistributionSpec] + ) -> dict[str, dict[str, float | str]]: + """Translate distributions for the sampler, flagging log-scale variables. + + A log-scale variable is sampled uniformly in log space (the surrogate + preprocessor re-applies the log). Sampling in log space is only defined for + a strictly-positive uniform, so anything else is rejected here -- at the + API boundary -- rather than surfacing as the sampler's raw ``ValueError``. + """ + engine: dict[str, dict[str, float | str]] = {} + for variable, spec in distributions.items(): + entry = spec.as_engine_dict() + if self._scale_of(variable) == "log": + if spec.distribution != "uniform": + raise SumoInputError( + f"'{variable}' is log-scale but its uncertainty is a " + f"'{spec.distribution}'; only a uniform supports log sampling" + ) + if spec.minimum is None or spec.minimum <= 0: + raise SumoInputError( + f"'{variable}' is log-scale but its distribution lower " + "bound is not strictly positive" + ) + entry["log_scale"] = True + engine[variable] = entry + return engine + def _scale_of(self, column: str) -> str: """The scale the caller asked for on ``column`` (default: linear).""" return self._spec.overrides.get(column, VariableSpec()).scale diff --git a/src/itis_sumo/data/funs_data_processing.py b/src/itis_sumo/data/funs_data_processing.py index 1d6c99b..e71226d 100644 --- a/src/itis_sumo/data/funs_data_processing.py +++ b/src/itis_sumo/data/funs_data_processing.py @@ -501,6 +501,11 @@ def create_manual_uq_samples( for var in input_vars: dist_info = distributions[var] dist_type = dist_info["distribution"] + log_scale = bool(dist_info.get("log_scale", False)) + if log_scale and dist_type != "uniform": + raise ValueError( + f"log_scale is only supported for uniform distributions: {var}" + ) if dist_type == "normal": mean = float(dist_info["mean"]) std = float(dist_info["std"]) @@ -510,9 +515,31 @@ def create_manual_uq_samples( elif dist_type == "uniform": min_val = float(dist_info["min"]) max_val = float(dist_info["max"]) - samples[var] = uniform.rvs( - size=num_samples, loc=min_val, scale=max_val - min_val, random_state=rng - ).tolist() + if log_scale: + if min_val <= 0: + raise ValueError( + f"Log-scale uniform bounds must be strictly positive: {var}" + ) + # Draw uniformly in log10 space, map back to the caller's original + # units -- the surrogate's own preprocessor re-applies the log, so + # what reaches it is a uniform-in-log sample (SPEC T27fr). + log_min, log_max = np.log10(min_val), np.log10(max_val) + samples[var] = np.power( + 10, + uniform.rvs( + size=num_samples, + loc=log_min, + scale=log_max - log_min, + random_state=rng, + ), + ).tolist() + else: + samples[var] = uniform.rvs( + size=num_samples, + loc=min_val, + scale=max_val - min_val, + random_state=rng, + ).tolist() elif dist_type == "constant": value = dist_info["value"] samples[var] = [float(value)] * num_samples diff --git a/tests/test_api_workflows.py b/tests/test_api_workflows.py index 98a4d50..422d38e 100644 --- a/tests/test_api_workflows.py +++ b/tests/test_api_workflows.py @@ -343,3 +343,93 @@ def test_holding_a_log_variable_at_zero_is_rejected(self, samples): preprocessing=log_input, points_per_variable=5, ) + + +_LOG_WIDTH = PreprocessingSpec(overrides={"width": VariableSpec(scale="log")}) +_UQ_DISTS = { + "width": DistributionSpec("uniform", minimum=1.0, maximum=5.0), + "height": DistributionSpec("uniform", minimum=100.0, maximum=500.0), +} + + +class TestLogScaleUncertaintyAndMetrics: + """Log must reach the UQ sampler and the CV metrics, not just the surrogate + (SPEC T27fr -- 'log applied everywhere').""" + + def test_log_input_samples_log_uniform_skewing_response_low(self, samples): + linear = evaluate_uncertainty( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + num_samples=400, + n_histograms=5, + seed=7, + ) + logw = evaluate_uncertainty( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + preprocessing=_LOG_WIDTH, + num_samples=400, + n_histograms=5, + seed=7, + ) + # width drawn log-uniform is skewed toward the low end, so the stress it + # drives must sit meaningfully below the linear-uniform case. + assert logw.mean > 0.0 + assert logw.mean < linear.mean - 0.5 + + def test_log_input_rejects_a_normal_uncertainty_distribution(self, samples): + normal_width = { + "width": DistributionSpec("normal", mean=3.0, std=0.5), + "height": _UQ_DISTS["height"], + } + with pytest.raises(SumoInputError, match="only a uniform supports log"): + evaluate_uncertainty( + samples, + VARIABLES, + RESPONSE, + distributions=normal_width, + preprocessing=_LOG_WIDTH, + num_samples=50, + n_histograms=3, + seed=7, + ) + + def test_log_input_rejects_a_non_positive_lower_bound(self, samples): + zero_min = { + "width": DistributionSpec("uniform", minimum=0.0, maximum=5.0), + "height": _UQ_DISTS["height"], + } + with pytest.raises(SumoInputError, match="not strictly positive"): + evaluate_uncertainty( + samples, + VARIABLES, + RESPONSE, + distributions=zero_min, + preprocessing=_LOG_WIDTH, + num_samples=50, + n_histograms=3, + seed=7, + ) + + def test_cv_metrics_honour_log_scale(self, samples): + linear = evaluate_cv_metrics(samples, VARIABLES, RESPONSE, seed=7) + log_response = evaluate_cv_metrics( + samples, VARIABLES, RESPONSE, preprocessing=_LOG_SCALE, seed=7 + ) + assert linear.root_mean_squared >= 0.0 + assert log_response.root_mean_squared >= 0.0 + # A log-trained surrogate is a different fit, so the metrics must not be + # byte-identical to the linear run (proving the flag reached the path). + assert log_response.root_mean_squared != pytest.approx( + linear.root_mean_squared, rel=1e-9 + ) + + def test_cv_metrics_reject_non_positive_log_response(self, samples): + bad = samples.copy() + bad.loc[bad.index[0], RESPONSE] = -1.0 + with pytest.raises(SumoInputError, match="log-scale but hold values"): + evaluate_cv_metrics(bad, VARIABLES, RESPONSE, preprocessing=_LOG_SCALE) diff --git a/tests/test_dakota_funs_data_processing.py b/tests/test_dakota_funs_data_processing.py index 407a839..6852d9e 100644 --- a/tests/test_dakota_funs_data_processing.py +++ b/tests/test_dakota_funs_data_processing.py @@ -189,6 +189,57 @@ def test_create_manual_uq_samples_unsupported_distribution_raises(): ) +def test_create_manual_uq_samples_log_scale_is_log_uniform_in_original_units(): + samples = create_manual_uq_samples( + ["x"], + {"x": {"distribution": "uniform", "min": 1.0, "max": 100.0, "log_scale": True}}, + num_samples=2000, + seed=3, + ) + values = np.asarray(samples["x"]) + # Returned in the caller's original units, but uniform in log10 space: the + # geometric mean sits at the midpoint of the log range (~10), well below the + # arithmetic midpoint (~50.5) a linear draw would centre on. + assert values.min() >= 1.0 and values.max() <= 100.0 + geometric_mean = float(np.exp(np.log(values).mean())) + assert geometric_mean == pytest.approx(10.0, rel=0.15) + assert float(values.mean()) < 50.0 # skewed toward the low end, unlike linear + + +def test_create_manual_uq_samples_log_scale_rejects_non_uniform(): + with pytest.raises(ValueError, match="only supported for uniform"): + create_manual_uq_samples( + ["x"], + { + "x": { + "distribution": "normal", + "mean": 1.0, + "std": 0.1, + "log_scale": True, + } + }, + num_samples=5, + seed=1, + ) + + +def test_create_manual_uq_samples_log_scale_rejects_non_positive_bound(): + with pytest.raises(ValueError, match="strictly positive"): + create_manual_uq_samples( + ["x"], + { + "x": { + "distribution": "uniform", + "min": 0.0, + "max": 5.0, + "log_scale": True, + } + }, + num_samples=5, + seed=1, + ) + + class TestCreateManualUqSamplesSeedReproducibility: """B12/V27: `seed` must actually control reproducibility of generated samples.""" From b07b3708b5e0b566c0d2b04b65d7c5e80c48eaa2 Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 22:42:29 +0200 Subject: [PATCH 03/13] feat(api): log-scale variables and objectives in MOGA (T27fr) Complete "log everywhere" for the optimizer. - api._session.optimize_pareto_front: accept a PreprocessingSpec; a log-scale variable is trained and explored in log space (its search domain is mapped into ln space and must be strictly positive -> else SumoInputError), a log-scale objective is fitted on ln(y) with the Pareto front exp-restored to original units via the preprocessor inverse. - api.optimize: forward the preprocessing spec. Tests: real-Dakota MOGA with a log objective returns an original-units front, and a log variable with a non-positive domain bound is rejected. --- src/itis_sumo/api/_session.py | 48 +++++++++++++++++++++++++++++----- src/itis_sumo/api/workflows.py | 4 +++ tests/test_api_workflows.py | 36 +++++++++++++++++++++++++ 3 files changed, 81 insertions(+), 7 deletions(-) diff --git a/src/itis_sumo/api/_session.py b/src/itis_sumo/api/_session.py index e8fd117..073d1a3 100644 --- a/src/itis_sumo/api/_session.py +++ b/src/itis_sumo/api/_session.py @@ -670,9 +670,17 @@ def optimize_pareto_front( domains: Mapping[str, DomainSpec], max_evaluations: int, workspace: Path | None, + preprocessing: PreprocessingSpec | None = None, ) -> ParetoFrontResult: - """Fit a surrogate per objective and find its Pareto-optimal trade-off front.""" + """Fit a surrogate per objective and find its Pareto-optimal trade-off front. + + A ``scale="log"`` override applies the same way everywhere else: a log-scale + *variable* trains and is explored in log space (so its search domain is mapped + into log space and must be strictly positive), and a log-scale *objective* is + fitted on ``ln(y)`` with the front restored to original units on the way out. + """ variables = tuple(variables) + spec = preprocessing or PreprocessingSpec() missing_domains = sorted(set(variables) - set(domains)) unknown_domains = sorted(set(domains) - set(variables)) if missing_domains or unknown_domains: @@ -681,9 +689,25 @@ def optimize_pareto_front( f"unknown={unknown_domains}" ) - validated = _validate_samples( - samples, variables, list(objectives), PreprocessingSpec() + def _is_log(column: str) -> bool: + return spec.overrides.get(column, VariableSpec()).scale == "log" + + non_positive_domain = sorted( + name + for name, dom in domains.items() + if _is_log(name) and (dom.minimum <= 0 or dom.maximum <= 0) ) + if non_positive_domain: + raise SumoInputError( + f"Log-scale {non_positive_domain} domains must be strictly positive" + ) + + # The training-data positivity guard covers any log-scale objective (log is + # undefined for <= 0 outputs) and reuses the shared validation path. + validated = _validate_samples(samples, variables, list(objectives), spec) + + log_inputs = [variable for variable in variables if _is_log(variable)] + log_objectives = [objective for objective in objectives if _is_log(objective)] run_dir = ( create_run_dir(Path(workspace), "sumo") @@ -702,6 +726,10 @@ def optimize_pareto_front( ] if maximize: preprocessor.setup_sign_switching(output_sign_switches=maximize) + if log_inputs or log_objectives: + preprocessor.setup_log_transform( + input_log_vars=log_inputs, output_log_vars=log_objectives + ) transformed = preprocessor.fit_transform(validated) training_file = run_dir / "processed_samples.dat" transformed.to_csv(training_file, sep=" ", index=False) @@ -713,10 +741,16 @@ def optimize_pareto_front( preprocessor.output_variables[response].mapped_name for response in objectives ] - mapped_domains = { - preprocessor.input_variables[variable].mapped_name: spec.as_engine_dict() - for variable, spec in domains.items() - } + mapped_domains: dict[str, dict[str, float | str]] = {} + for name, dom in domains.items(): + if _is_log(name): + dom = DomainSpec( + minimum=float(np.log(dom.minimum)), + maximum=float(np.log(dom.maximum)), + ) + mapped_domains[preprocessor.input_variables[name].mapped_name] = ( + dom.as_engine_dict() + ) try: results = perform_moga_optimization( diff --git a/src/itis_sumo/api/workflows.py b/src/itis_sumo/api/workflows.py index d5f8731..b57205e 100644 --- a/src/itis_sumo/api/workflows.py +++ b/src/itis_sumo/api/workflows.py @@ -255,6 +255,7 @@ def optimize( *, domains: Mapping[str, DomainSpec], max_evaluations: int = 1000, + preprocessing: PreprocessingSpec | None = None, workspace: Path | None = None, ) -> ParetoFrontResult: """Find the Pareto-optimal trade-off front across one or more objectives. @@ -262,6 +263,8 @@ def optimize( Unlike the other workflows, this fits one surrogate per objective over a domain (where exploration is allowed), not a real-world uncertainty distribution -- MOGA cannot use anything but a uniform domain (SPEC T27fr). + A ``scale="log"`` override on a variable explores it in log space; on an + objective it fits and reports the front in log space. """ return optimize_pareto_front( samples, @@ -269,6 +272,7 @@ def optimize( objectives, domains=domains, max_evaluations=max_evaluations, + preprocessing=preprocessing, workspace=workspace, ) diff --git a/tests/test_api_workflows.py b/tests/test_api_workflows.py index 422d38e..75326fd 100644 --- a/tests/test_api_workflows.py +++ b/tests/test_api_workflows.py @@ -264,6 +264,42 @@ def test_requires_a_domain_for_each_variable(self, samples): max_evaluations=200, ) + def test_log_objective_front_returns_original_units(self, samples): + domains = { + "width": DomainSpec(minimum=WIDTH_RANGE[0], maximum=WIDTH_RANGE[1]), + "height": DomainSpec(minimum=HEIGHT_RANGE[0], maximum=HEIGHT_RANGE[1]), + } + result = optimize( + samples, + VARIABLES, + {RESPONSE: "minimize"}, + domains=domains, + max_evaluations=200, + preprocessing=_LOG_SCALE, + ) + front = result.data[RESPONSE] + assert front + # A log-trained objective must come back exp-restored to original stress + # units (O(10)), not the O(1..3) ln values. + assert min(front) > 0.0 + assert max(front) < 10 * samples[RESPONSE].max() + + def test_log_variable_domain_rejects_non_positive_bounds(self, samples): + domains = { + "width": DomainSpec(minimum=0.0, maximum=5.0), # log => must be > 0 + "height": DomainSpec(minimum=HEIGHT_RANGE[0], maximum=HEIGHT_RANGE[1]), + } + log_width = PreprocessingSpec(overrides={"width": VariableSpec(scale="log")}) + with pytest.raises(SumoInputError, match="domains must be strictly positive"): + optimize( + samples, + VARIABLES, + {RESPONSE: "minimize"}, + domains=domains, + max_evaluations=100, + preprocessing=log_width, + ) + _LOG_SCALE = PreprocessingSpec(overrides={RESPONSE: VariableSpec(scale="log")}) From 041a70dfa52f1fd55cddf7bb42d638f9c5877a49 Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 22:45:49 +0200 Subject: [PATCH 04/13] feat(api): log-scale sampling in Sobol indices (T27fr) Complete "log everywhere" for the sensitivity workflow. - api._session.sobol: route distributions through _uq_engine_distributions, so a log-scale variable is flagged log_scale and rejected (SumoInputError) unless it carries a strictly-positive uniform -- same contract as uncertainty. - evaluate.evaluate_sobol_indices: build the Saltelli ppf in log space for a log-scale variable (scipy.stats.loguniform -> log-uniform in the caller's original units); the surrogate's preprocessor re-applies the log downstream. - Reword the shared guard as "distribution" not "uncertainty" so it reads right for both UQ and Sobol. Test: real-Dakota Sobol with a log input shifts the variance decomposition in the expected direction (compressed width explains less, height's share rises), plus log+normal and log+min<=0 rejections. --- src/itis_sumo/api/_session.py | 6 +-- src/itis_sumo/evaluate/funs_evaluate.py | 26 +++++++++--- tests/test_api_workflows.py | 56 +++++++++++++++++++++++++ 3 files changed, 79 insertions(+), 9 deletions(-) diff --git a/src/itis_sumo/api/_session.py b/src/itis_sumo/api/_session.py index 073d1a3..5530b33 100644 --- a/src/itis_sumo/api/_session.py +++ b/src/itis_sumo/api/_session.py @@ -420,9 +420,7 @@ def sobol( f"Distributions must cover variables exactly; missing={missing}, " f"unknown={unknown}" ) - engine_distributions = { - variable: spec.as_engine_dict() for variable, spec in distributions.items() - } + engine_distributions = self._uq_engine_distributions(distributions) results = self._run_engine( "computing Sobol indices", evaluate_sobol_indices, @@ -514,7 +512,7 @@ def _uq_engine_distributions( if self._scale_of(variable) == "log": if spec.distribution != "uniform": raise SumoInputError( - f"'{variable}' is log-scale but its uncertainty is a " + f"'{variable}' is log-scale but its distribution is a " f"'{spec.distribution}'; only a uniform supports log sampling" ) if spec.minimum is None or spec.minimum <= 0: diff --git a/src/itis_sumo/evaluate/funs_evaluate.py b/src/itis_sumo/evaluate/funs_evaluate.py index c674eb7..8880dd4 100644 --- a/src/itis_sumo/evaluate/funs_evaluate.py +++ b/src/itis_sumo/evaluate/funs_evaluate.py @@ -1035,7 +1035,7 @@ def evaluate_sobol_indices( import math import pandas as pd - from scipy.stats import norm, sobol_indices, uniform + from scipy.stats import loguniform, norm, sobol_indices, uniform from scipy.stats.qmc import Sobol # NOTE: input_vars/distributions must stay in the caller's original @@ -1059,17 +1059,33 @@ def evaluate_sobol_indices( d_varying = len(varying_vars) - # Build frozen scipy distributions with .ppf for each varying variable + # Build frozen scipy distributions with .ppf for each varying variable. A + # log-scale variable is drawn uniformly in log space (log-uniform in the + # caller's original units), mirroring create_manual_uq_samples -- the + # surrogate's own preprocessor re-applies the log downstream. ppfs = {} for var in varying_vars: dist_info = distributions[var] dist_type = dist_info["distribution"] + log_scale = bool(dist_info.get("log_scale", False)) + if log_scale and dist_type != "uniform": + raise ValueError( + f"log_scale is only supported for uniform distributions: {var}" + ) if dist_type == "normal": ppfs[var] = norm(loc=dist_info["mean"], scale=dist_info["std"]) elif dist_type == "uniform": - ppfs[var] = uniform( - loc=dist_info["min"], scale=dist_info["max"] - dist_info["min"] - ) + if log_scale: + lo, hi = float(dist_info["min"]), float(dist_info["max"]) + if lo <= 0: + raise ValueError( + f"Log-scale uniform bounds must be strictly positive: {var}" + ) + ppfs[var] = loguniform(a=lo, b=hi) + else: + ppfs[var] = uniform( + loc=dist_info["min"], scale=dist_info["max"] - dist_info["min"] + ) else: raise ValueError(f"Unsupported distribution type: {dist_type}") diff --git a/tests/test_api_workflows.py b/tests/test_api_workflows.py index 75326fd..4e40f0b 100644 --- a/tests/test_api_workflows.py +++ b/tests/test_api_workflows.py @@ -469,3 +469,59 @@ def test_cv_metrics_reject_non_positive_log_response(self, samples): bad.loc[bad.index[0], RESPONSE] = -1.0 with pytest.raises(SumoInputError, match="log-scale but hold values"): evaluate_cv_metrics(bad, VARIABLES, RESPONSE, preprocessing=_LOG_SCALE) + + +class TestLogScaleSobol: + """Log must reach the Sobol sampler too (SPEC T27fr -- 'log everywhere'): a + log-scale variable is drawn log-uniform, so the variance decomposition is + taken over that distribution rather than a linear one.""" + + def test_log_input_changes_the_variance_decomposition(self, samples): + linear = evaluate_sobol( + samples, VARIABLES, RESPONSE, distributions=_UQ_DISTS, seed=7 + ) + logw = evaluate_sobol( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + preprocessing=_LOG_WIDTH, + seed=7, + ) + assert set(logw.indices) == set(VARIABLES) + # width drawn log-uniform is compressed toward its low end, so it explains + # slightly less response variance than a linear width -- and height, taking + # the residual share, rises. The direction proves the flag reached the + # sampler, not just the surrogate fit. + assert logw.indices["width"]["total"] < linear.indices["width"]["total"] + assert logw.indices["height"]["total"] > linear.indices["height"]["total"] + + def test_log_input_rejects_a_normal_distribution(self, samples): + normal_width = { + "width": DistributionSpec("normal", mean=3.0, std=0.5), + "height": _UQ_DISTS["height"], + } + with pytest.raises(SumoInputError, match="only a uniform supports log"): + evaluate_sobol( + samples, + VARIABLES, + RESPONSE, + distributions=normal_width, + preprocessing=_LOG_WIDTH, + seed=7, + ) + + def test_log_input_rejects_a_non_positive_lower_bound(self, samples): + zero_min = { + "width": DistributionSpec("uniform", minimum=0.0, maximum=5.0), + "height": _UQ_DISTS["height"], + } + with pytest.raises(SumoInputError, match="not strictly positive"): + evaluate_sobol( + samples, + VARIABLES, + RESPONSE, + distributions=zero_min, + preprocessing=_LOG_WIDTH, + seed=7, + ) From 914a1c24d24f2c0d11efa0c4a8ae9dc1f61975f8 Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 22:46:52 +0200 Subject: [PATCH 05/13] docs(spec): honour log-scale in every workflow (T27fr / V44ls / B17sc) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Record the "log everywhere" behaviour just shipped in code. - V44ls: a scale="log" override must be honoured by every scale-consuming api workflow (surrogate fit, UQ/Sobol sampling, MOGA search domain, CV metric), never silently ignored outside the surrogate fit. - B17sc (backprop): log reached the surrogate but was a near-no-op for UQ/MOGA/cv-metrics/Sobol after the initial port; the fix wires log-space into each path, guarded by V44ls. - T27fr note: backend log-scale now complete end-to-end; remaining scope narrowed to the domain⊥distribution split (V26dd) + mmux_vite frontend consumer migration. --- SPEC.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/SPEC.md b/SPEC.md index e392723..760edeb 100644 --- a/SPEC.md +++ b/SPEC.md @@ -109,6 +109,7 @@ V40gz: commit-time hooks (`prek` tool, run via `uvx prek run --all-files`; mirro V41hp: `ty check` ! zero diagnostics repo-wide (blocking in both the local `prek` hook and CI); version-gated stdlib/typing symbols (e.g. `datetime.UTC`, `typing.Self`, `tomllib` — all 3.11+-only, caught mid-stack once actually run, T42jn) ! resolved against `requires-python`'s floor, not the CI runner's own interpreter — this is strictly cheaper than the 3.10-3.13 test matrix (V18rs) for that bug class, though the matrix stays as the ground-truth runtime check V19nd: ∀ Dependabot ecosystem entry in `.github/dependabot.yml` → weekly Monday 03:00 Europe/Zurich schedule ∧ exactly one wildcard dependency group; ⊥ unspecified default schedule or one PR per dependency V20qx: ∀ ignored Dependabot dependency → ignore reason names active compatibility constraint ∧ references tracked resolution task; `itis-dakota` remains ignored until T16mo resolves the Dakota 6.23+ interface-cache regression +V44ls: ∀ scale-consuming `itis_sumo.api` workflow (surrogate fit ∧ UQ/Sobol input sampling ∧ MOGA search domain ∧ CV accuracy metric) → a `scale="log"` override ! be honoured in that workflow's own computation (log-space surrogate + log-space sampling/search where a distribution or domain is involved, positivity-guarded); ⊥ a `scale` override accepted but silently ignored outside the surrogate fit ## §R R1: `export_model`/`import_model` child keywords; formats `text_archive`(.sps)/`binary_archive`(.bsps)/`algebraic_file`(.alg); naming `{prefix}.{resp}.{ext}` | branch R2 @@ -149,7 +150,7 @@ T23bn|✓|`itis_sumo.api`: `cross_validate()` + `evaluate_along_axes()` — inte T24cm|✓|remove `preprocess/models.py::{FunctionJob,JobVariableSelection}`; re-express the Dakota sufficiency rule over tabular data; jobs→table adapter moves to flaskapi (this port DELETES itis-sumo code)|V20dm T25dp|.|`itis_sumo.api`: remaining 6 workflows (UQ-w-uncertainty incl. the ~120-line erfinv/histogram block, correlation, Sobol, grid, MOGA, cv-accuracy-metrics) + E1 `export_model`/`evaluate_stored_model` facade|V22rs,R1-R4 T26eq|.|POST-PORT: extract the fitted-model handle (`fit()` → methods → `save()`/`load()`); carries the fitted preprocessing config ⇒ closes the E1 gap (model store persists archive+metadata+training copy but ⊥ preprocessor config, so a reloaded model cannot inverse-transform to original units)|V27fq,V10jk -T27fr|~|POST-PORT: split `domain` vs `distribution` config + consumer migration; absorb the mmux_vite `jgo/fullstack-logscale` work|V26dd +T27fr|~|POST-PORT: split `domain` vs `distribution` config + consumer migration; absorb the mmux_vite `jgo/fullstack-logscale` work — backend log-scale now honoured end-to-end (surrogate + along-axes + grid + CV + UQ sampler + CV-metrics + MOGA + Sobol, V44ls); remaining = the `domain`⊥`distribution` config split (V26dd) ∧ the mmux_vite frontend consumer migration|V26dd,V44ls T28gs|.|POST-PORT `?`: decide whether `distribution` gets an auto-generated default — explicit discussion required, ⊥ silently defaulted|§C `?` T29hw|~|`publish.yml` tag trigger accepts `v`-prefixed PEP 440 prereleases ✓; tagging `v0.1.0a1` BLOCKED on the one-time PyPI Trusted Publisher config (user action), then clean-venv install + `itis-sumo validate` + headless smoke|T1pw T30qa|✓|release/CI workflow refresh: build→TestPyPI automatic on tag push (✓), PyPI+Release gated manual (✓, T35cc); auto-tag-on-branch redesigned to PR-time check + tag-only-at-merge (no bot commit) — see T31xx/T32yy; keep git-cliff release notes, dependency-review + concurrency (✓); skip weekly cron/healthchecks for now|§C,V17rt @@ -191,3 +192,4 @@ B14zt|2026-08-27|`itis-dakota==1.5.9` was deleted from PyPI upstream (release li B15tg|2026-08-27|the `v0.1.0a2` tag (created by auto-tag.yml on the develop merge) exists, but no TestPyPI publish followed — auto-tag's trailing `gh workflow run publish.yml` did not fire. Root cause: the `tag-merged-version` step invoked `gh` from an inline Python script that read `GITHUB_TOKEN` out of the step environment, but the step's `env:` block only exported `GH_REF_NAME`, so `gh` got no usable token and the dispatch silently did nothing|add `GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}` to the `tag-merged-version` step's `env:` block so the inline `gh` call inherits a valid token; one-off `gh workflow run publish.yml --ref develop --field tag=v0.1.0a2 --field registry=testpypi` published a2 manually to recover the missed alpha; V29yn B16mo|2026-08-28|T16mo rung 2 bumped the engine `itis-dakota==1.5.11` (Dakota 6.20) → `==6.24.7` (Dakota 6.24). 6.24.7 still carries the 6.23+ `Interface::interface_cache()` regression (R5): `dakota.environment.study()`'s ctor unconditionally calls `interface_cache(problem_db)` and throws `IndexError: map::at` ("called with nonexistent study!") for itis-sumo's pure data-fit surrogate confs, which construct no Interface (only `import_build_points_file`, no `interface` block) — the cache static map is never populated. All 45 real-Dakota surrogate tests failed on it (training, cross-validation, UQ propagation, MOGA, import/export, evaluation). Upstream fix (get-or-create `operator[]`) is still pending|shim in itis-sumo, not waiting on upstream: `config/funs_create_dakota_conf.py::add_r5_interface_cache_workaround()` (appended to every conf via `start_dakota_file()`) declares an unused `single` model `R5_WA_MODEL` wired to a no-op `fork` interface (`analysis_drivers='true'`) and points each surrogate `model` at it via `truth_model_pointer='R5_WA_MODEL'`, forcing the Interface to be constructed (populating the cache) during study setup; the dummy model is never referenced by any method/iterator so the interface is never executed. 318-test suite green on 6.24.7. Took 6.24.7 (not 6.24.9) because 6.24.8+ dropped cp311 wheels, which would have forced a Python 3.12 migration of both repos; R5 B8hm|2026-09-11|Dependabot entries used unspecified weekly timing, and `itis-dakota` was intentionally ignored without a durable policy invariant|V19nd,V20qx +B17sc|2026-09-28|`scale="log"` on `PreprocessingSpec.overrides` reached the surrogate fit but was a near-no-op for UQ/MOGA/cv-metrics/Sobol — the logscale port (T27fr) landed the preprocessor transform but not per-workflow log-space sampling/search, so `evaluate_uncertainty` with a log input returned linear-space statistics (probe: mean 12.057 linear vs 12.058 log). Symptom never seen in flaskapi, whose frontend drove log at the request-payload layer|wire log-space into every scale-consuming path: `create_manual_uq_samples` log_scale branch (log-uniform, uniform-only + positivity), `_uq_engine_distributions` flags ∧ validates log vars for UQ + Sobol, `evaluate_sobol_indices` samples `scipy.stats.loguniform`, `optimize_pareto_front` maps a log variable's search domain into ln space ∧ fits log objectives; `evaluate_cv_metrics` inherits log via `cross_validate`. Guarded by V44ls ∧ tests across the four workflows|V44ls,T27fr From e7cbcc0e7ed4468869fbbdc27af11c9eea2c3327 Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 22:51:43 +0200 Subject: [PATCH 06/13] chore(release): 0.1.0a5 -> 0.1.0a6 version-check (V28tz) requires a develop PR's version to be strictly newer than every existing tag; v0.1.0a5 is already tagged at develop's tip, so this PR must carry a6. pyproject + uv.lock self-version moved together (uv lock --locked clean). --- pyproject.toml | 2 +- uv.lock | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/pyproject.toml b/pyproject.toml index 64863b7..5fc7660 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "uv_build" [project] name = "itis-sumo" -version = "0.1.0a5" +version = "0.1.0a6" description = "Surrogate Modeling functionality for IT'IS Foundation / ZMT Modeling Intelligence suite: build, evaluate, cross-validate surrogates + UQ + sampling, headless or embedded" readme = "README.md" # T16mo rung 1 (1.5.9->1.5.11): 1.5.11 ships cp313 wheels, so the <3.13 diff --git a/uv.lock b/uv.lock index c84e42c..c39abdb 100644 --- a/uv.lock +++ b/uv.lock @@ -458,7 +458,7 @@ wheels = [ [[package]] name = "itis-sumo" -version = "0.1.0a5" +version = "0.1.0a6" source = { editable = "." } dependencies = [ { name = "itis-dakota" }, From 83c150d9a6dad32092b8a72ed4ce3c18628ad741 Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 23:20:35 +0200 Subject: [PATCH 07/13] docs(spec): V45ls scale-is-first-class + V21pf amendment, T46ls opened scale must reach every value-producing entry point (samplers + correlations included), enforced structurally (required-arg helpers) + behaviourally (linear<->log flip-test). Decisions locked in review discussion. --- SPEC.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/SPEC.md b/SPEC.md index 760edeb..aba480a 100644 --- a/SPEC.md +++ b/SPEC.md @@ -84,7 +84,7 @@ V17ab: ∀ `.py` file in repo ! `ruff check` + `ruff format --check` clean (CI-e V18rs: ∀ `.github/workflows/**` change ! at least one validation job executes the affected workflow path; shared CI workflow changes ⊥ leave the validation matrix entirely skipped V19cn: public API ! use VOCAB nouns (sample/variable=parameter/response=QoI); ⊥ "job" in any itis_sumo signature, docstring, or result field V20dm: consumer-facing entrypoints ! accept plain tabular data; ⊥ `FunctionJob`/`JobVariableSelection`/oSPARC status strings inside itis_sumo (test-enforced, sibling of V4ty) -V21pf: preprocessing ! auto-defaulted + overridable in domain vocabulary only; ⊥ transform vocabulary reachable from a public signature; effective config inspectable in the result +V21pf: preprocessing ! auto-defaulted + overridable in domain vocabulary only; ⊥ transform vocabulary reachable from a public signature; effective config inspectable in the result; ∀ value-producing entry point ! accept `scale` (default linear — ⊥ caller must pass) ∧ ! apply it (V45ls) V22rs: results ! typed dataclass, original units + original names, `dataclasses.asdict()`-serializable; ⊥ `_hat`/`_std_hat` suffix keys in the public shape V23er: ∀ error escaping `itis_sumo.api` ! be a `SumoError` subclass; ⊥ raw `KeyError`/`IndexError`/`ValueError` crossing the boundary V24af: run dir ! discarded on success, PRESERVED on failure w/ its path + stderr tail attached to the raised `SumoEngineError` @@ -110,6 +110,7 @@ V41hp: `ty check` ! zero diagnostics repo-wide (blocking in both the local `prek V19nd: ∀ Dependabot ecosystem entry in `.github/dependabot.yml` → weekly Monday 03:00 Europe/Zurich schedule ∧ exactly one wildcard dependency group; ⊥ unspecified default schedule or one PR per dependency V20qx: ∀ ignored Dependabot dependency → ignore reason names active compatibility constraint ∧ references tracked resolution task; `itis-dakota` remains ignored until T16mo resolves the Dakota 6.23+ interface-cache regression V44ls: ∀ scale-consuming `itis_sumo.api` workflow (surrogate fit ∧ UQ/Sobol input sampling ∧ MOGA search domain ∧ CV accuracy metric) → a `scale="log"` override ! be honoured in that workflow's own computation (log-space surrogate + log-space sampling/search where a distribution or domain is involved, positivity-guarded); ⊥ a `scale` override accepted but silently ignored outside the surrogate fit +V45ls: ∀ value-producing `itis_sumo.api` entry point → `scale` ! be threaded into the computation that yields those values ∧ observable in the result; ⊥ a value the api produced while ignoring the caller's scale; enforced: structural (internal value-producers `scale_distribution`/`resolve_log_scale` take scale/flag as a required arg — unwired code breaks at call w/ `TypeError`, never silent drift) ∧ behavioural (linear↔log flip-test over ∀ 10 entry points; rank-correlation exempt — monotone-invariant, asserted unchanged) ## §R R1: `export_model`/`import_model` child keywords; formats `text_archive`(.sps)/`binary_archive`(.bsps)/`algebraic_file`(.alg); naming `{prefix}.{resp}.{ext}` | branch R2 @@ -160,6 +161,7 @@ T33zz|✓|swap itis-sumo LICENSE + `pyproject.toml` `license`/classifiers to mat T34aa|✓|`make publish-testpypi-dev` ! source `TESTPYPI_TOKEN` from local `.env` (gitignored) instead of requiring a pre-exported shell var|§C T35cc|✓|gate `publish`/`release` jobs behind explicit `workflow_dispatch`; `build`+`verify` stay CI-only for release tags; feature `.devN` uploads move to local Make target|V31vp,V32bb T36dd|✓|add `.env` to `.gitignore`; make `publish-testpypi` source `.env` without printing token, then build/check/publish|V37bb +T46ls|.|scale first-class across the WHOLE api surface: `generate_lhs_samples` ∧ `generate_grid_samples` ∧ `compute_correlations` take `preprocessing` ∧ honour scale; one `scale_distribution` unit→value helper absorbs the log10-power hand-roll ∧ the duplicate log guards ∧ the `_is_log` closure; V45ls linear↔log flip-tests over ∀ 10 entry points + 5 gap tests (MOGA log·maximize, log-variable axis x-restoration, log-response UQ spread, Sobol mixed partition, repo-wide `ty`); rows land in the V&V report Category I|V45ls,V44ls,V21pf,V22rs T37ef|✓|validate manual publish tag + artifact version before PyPI upload|V35wk T38dd|✓|make `publish-testpypi-dev` auto-compute/write `.devN`, build/check/upload directly to TestPyPI, then restore project version; CI verifies alpha/beta/rc before real PyPI; no dev tag cascade|V38cc,V39qf T39pk|.|POST-PORT/deepen candidate: `itis_sumo.api.workflows` repeats the same build-session → translate-config → call → translate-result → wrap-errors skeleton per function; pull the shared shape down into `_session.py` so each `workflows.py` function shrinks to signature+one call — behavior-preserving refactor, tests green before/after; natural precursor to T26eq's handle extraction|V16qf,V27fq,T26eq From c4e3abd8b2ed7aa70032aebe1c61e24ac9dcc42c Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 23:27:39 +0200 Subject: [PATCH 08/13] fix(tests): narrow Optional match in dependabot ecosystem test Pre-existing ty diagnostic (unguarded .group on re.search) reding prek on every PR into develop, incl. #47. Walrus + assert keeps the test strict (silently skipping a block would hide config drift). --- tests/test_dependabot_config.py | 14 ++++++++------ 1 file changed, 8 insertions(+), 6 deletions(-) diff --git a/tests/test_dependabot_config.py b/tests/test_dependabot_config.py index 2c3281a..1fa1861 100644 --- a/tests/test_dependabot_config.py +++ b/tests/test_dependabot_config.py @@ -27,12 +27,14 @@ def test_v19nd_dependabot_updates_are_grouped_and_scheduled(): blocks = _update_blocks() assert len(blocks) == 2 - assert { - re.search(r"package-ecosystem: (\S+)", block).group(1) for block in blocks - } == { - "github-actions", - "uv", - } + ecosystems = set() + for block in blocks: + # _update_blocks only yields blocks anchored on this key, but ty needs + # the Optional narrowed -- assert rather than silently skip a block. + ecosystem = re.search(r"package-ecosystem: (\S+)", block) + assert ecosystem is not None + ecosystems.add(ecosystem.group(1)) + assert ecosystems == {"github-actions", "uv"} for block in blocks: assert "target-branch: develop" in block From e691a65ae4654e4c478c761640de1707fb303ae4 Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 23:27:39 +0200 Subject: [PATCH 09/13] feat(api): scale first-class on samplers + correlations (T46ls) V45ls: every value-producing api entry point takes scale and applies it. - data.funs_data_processing: one unit->value map scale_distribution(min,max,*,scale) (uniform vs scipy.loguniform) + resolve_log_scale + scale_values; scale is a REQUIRED arg -- unwired value producers break at call with TypeError, never silently default. Absorbs the log10-power hand-roll and the duplicated log guards in manual-UQ sampling and the Sobol ppf path. - api.generate_lhs_samples / generate_grid_samples: accept preprocessing; log domains fill log-uniform/geometric (linear default bit-identical); non-positive log domain -> SumoInputError. - api.compute_correlations: accepts preprocessing; compute_correlation_indices requires input_scales/output_scale (Pearson moves under log, Spearman is monotone-invariant, asserted). - api._session: _is_log closure + method getter collapse into module-level column_scale(spec, column). Tests: correlator unit tests adopt required scales + new scale-shift/invariance and TypeError-tripwire cases. 359-pass suite, ruff/ty clean. --- src/itis_sumo/api/_session.py | 28 +++-- src/itis_sumo/api/workflows.py | 90 +++++++++++++--- src/itis_sumo/data/funs_data_processing.py | 117 ++++++++++++++------- src/itis_sumo/evaluate/funs_evaluate.py | 34 +++--- tests/test_correlation_indices.py | 109 +++++++++++++++++-- 5 files changed, 289 insertions(+), 89 deletions(-) diff --git a/src/itis_sumo/api/_session.py b/src/itis_sumo/api/_session.py index 5530b33..a8dc8ff 100644 --- a/src/itis_sumo/api/_session.py +++ b/src/itis_sumo/api/_session.py @@ -81,6 +81,17 @@ def _stderr_tail(run_dir: Path | None) -> str: return "\n".join(lines[-_STDERR_TAIL_LINES:]) +def column_scale(spec: PreprocessingSpec | None, column: str) -> str: + """The scale asked for on ``column``; ``None`` spec ≡ all-linear (V21pf). + + One home for the lookup, shared by the session, the optimizer and the + standalone samplers (V45ls: every value producer reads scale through here). + """ + if spec is None: + return "linear" + return spec.overrides.get(column, VariableSpec()).scale + + def _validate_samples( samples: pd.DataFrame, variables: Sequence[str], @@ -526,7 +537,7 @@ def _uq_engine_distributions( def _scale_of(self, column: str) -> str: """The scale the caller asked for on ``column`` (default: linear).""" - return self._spec.overrides.get(column, VariableSpec()).scale + return column_scale(self._spec, column) def _to_original_std( self, @@ -687,13 +698,17 @@ def optimize_pareto_front( f"unknown={unknown_domains}" ) - def _is_log(column: str) -> bool: - return spec.overrides.get(column, VariableSpec()).scale == "log" + log_inputs = [ + variable for variable in variables if column_scale(spec, variable) == "log" + ] + log_objectives = [ + objective for objective in objectives if column_scale(spec, objective) == "log" + ] non_positive_domain = sorted( name for name, dom in domains.items() - if _is_log(name) and (dom.minimum <= 0 or dom.maximum <= 0) + if name in log_inputs and (dom.minimum <= 0 or dom.maximum <= 0) ) if non_positive_domain: raise SumoInputError( @@ -704,9 +719,6 @@ def _is_log(column: str) -> bool: # undefined for <= 0 outputs) and reuses the shared validation path. validated = _validate_samples(samples, variables, list(objectives), spec) - log_inputs = [variable for variable in variables if _is_log(variable)] - log_objectives = [objective for objective in objectives if _is_log(objective)] - run_dir = ( create_run_dir(Path(workspace), "sumo") if workspace is not None @@ -741,7 +753,7 @@ def _is_log(column: str) -> bool: ] mapped_domains: dict[str, dict[str, float | str]] = {} for name, dom in domains.items(): - if _is_log(name): + if name in log_inputs: dom = DomainSpec( minimum=float(np.log(dom.minimum)), maximum=float(np.log(dom.maximum)), diff --git a/src/itis_sumo/api/workflows.py b/src/itis_sumo/api/workflows.py index b57205e..62e8e4b 100644 --- a/src/itis_sumo/api/workflows.py +++ b/src/itis_sumo/api/workflows.py @@ -19,9 +19,14 @@ from collections.abc import Mapping, Sequence from pathlib import Path +import numpy as np import pandas as pd -from itis_sumo.api._session import SumoSession, optimize_pareto_front +from itis_sumo.api._session import ( + SumoSession, + column_scale, + optimize_pareto_front, +) from itis_sumo.api.errors import SumoInputError from itis_sumo.api.types import ( DEFAULT_SEED, @@ -42,6 +47,7 @@ compute_correlation_indices, create_grid_samples, load_data, + scale_distribution, ) from itis_sumo.evaluate.funs_evaluate import compute_cv_accuracy_metrics from itis_sumo.sampling.lhs import lhs as _lhs @@ -180,14 +186,26 @@ def compute_correlations( samples: pd.DataFrame, variables: Sequence[str], response: str, + *, + preprocessing: PreprocessingSpec | None = None, ) -> CorrelationResult: - """Compute response correlations from a caller-owned sample table.""" + """Compute response correlations from a caller-owned sample table. + + Each column is correlated on its own scale (V45ls): Pearson is not + invariant to log, so a log-scaled variable's coefficient shifts; Spearman + is monotone-invariant and therefore unchanged. + """ missing = sorted((set(variables) | {response}) - set(samples.columns)) if missing: raise SumoInputError(f"Samples do not contain columns: {missing}") + spec = preprocessing or PreprocessingSpec() try: coefficients = compute_correlation_indices( - samples, samples[response].tolist(), list(variables) + samples, + samples[response].tolist(), + list(variables), + input_scales={v: column_scale(spec, v) for v in variables}, + output_scale=column_scale(spec, response), ) except ValueError as exc: raise SumoInputError(str(exc)) from exc @@ -281,47 +299,66 @@ def generate_lhs_samples( domains: Mapping[str, DomainSpec], n_samples: int, *, + preprocessing: PreprocessingSpec | None = None, seed: int = DEFAULT_SEED, ) -> pd.DataFrame: """Draw a Latin-hypercube design over the given variable domains. + A ``scale="log"`` domain (via ``preprocessing``) is filled log-uniformly -- + the unit design maps through the same ``scale_distribution`` every value + producer uses (V45ls); linear domains are unchanged. + Args: domains: Variable name -> allowed range to draw from. n_samples: Number of sample rows to generate. + preprocessing: Per-column scale overrides (default: all linear). seed: Controls the draw. Raises: - SumoInputError: No domains given. + SumoInputError: No domains given, or a log-scale domain is not strictly + positive. """ if not domains: raise SumoInputError("At least one variable domain is required.") names = list(domains) design = _lhs(len(names), n_samples, seed=seed) - return pd.DataFrame( - { - name: design[:, i] * (domains[name].maximum - domains[name].minimum) - + domains[name].minimum - for i, name in enumerate(names) - } - ) + columns = {} + for i, name in enumerate(names): + dom = domains[name] + try: + dist = scale_distribution( + dom.minimum, dom.maximum, scale=column_scale(preprocessing, name) + ) + except ValueError as exc: + raise SumoInputError(f"'{name}': {exc}") from exc + columns[name] = dist.ppf(design[:, i]) + return pd.DataFrame(columns) def generate_grid_samples( domains: Mapping[str, DomainSpec], points_per_variable: Mapping[str, int], *, + preprocessing: PreprocessingSpec | None = None, workspace: Path | None = None, ) -> pd.DataFrame: """Generate a full-factorial grid of samples over the given variable domains. + A ``scale="log"`` domain (via ``preprocessing``) is gridded log-uniformly: + the axis is built in log space (``ln`` bounds to the linspace, then exp back), + so its points are geometrically spaced -- matching the LHS/UQ log spacing + (V45ls). Linear domains are unchanged. + Args: domains: Variable name -> allowed range to draw from. points_per_variable: Variable name -> number of grid points along that axis. + preprocessing: Per-column scale overrides (default: all linear). workspace: If given, working files are written here and kept. If omitted, they are discarded on success and kept on failure. Raises: - SumoInputError: No domains given, or a variable is missing its point count. + SumoInputError: No domains given, a variable is missing its point count, + or a log-scale domain is not strictly positive. """ if not domains: raise SumoInputError("At least one variable domain is required.") @@ -330,6 +367,24 @@ def generate_grid_samples( if missing: raise SumoInputError(f"Missing points_per_variable for: {', '.join(missing)}") + log_names = set() + build_bounds: dict[str, tuple[float, float]] = {} + for name in names: + dom = domains[name] + if column_scale(preprocessing, name) == "log": + if dom.minimum <= 0: + raise SumoInputError( + f"'{name}': log-scale domain must be strictly positive, got " + f"[{dom.minimum}, {dom.maximum}]" + ) + log_names.add(name) + build_bounds[name] = ( + float(np.log(dom.minimum)), + float(np.log(dom.maximum)), + ) + else: + build_bounds[name] = (dom.minimum, dom.maximum) + run_dir = ( create_run_dir(Path(workspace), "grid-sample") if workspace is not None @@ -340,14 +395,17 @@ def generate_grid_samples( run_dir=run_dir, grid_vars=names, input_vars=names, - mins=[domains[name].minimum for name in names], + mins=[build_bounds[name][0] for name in names], cut_values=[ - (domains[name].minimum + domains[name].maximum) / 2 for name in names + (build_bounds[name][0] + build_bounds[name][1]) / 2 for name in names ], - maxs=[domains[name].maximum for name in names], + maxs=[build_bounds[name][1] for name in names], n_points_per_dimension=[points_per_variable[name] for name in names], ) - return load_data(grid_file)[names].astype(float) + grid = load_data(grid_file)[names].astype(float) + for name in log_names: + grid[name] = np.exp(grid[name]) + return grid finally: if workspace is None: shutil.rmtree(run_dir, ignore_errors=True) diff --git a/src/itis_sumo/data/funs_data_processing.py b/src/itis_sumo/data/funs_data_processing.py index e71226d..a477960 100644 --- a/src/itis_sumo/data/funs_data_processing.py +++ b/src/itis_sumo/data/funs_data_processing.py @@ -479,6 +479,59 @@ def extract_predictions_gridpoints( return results +def resolve_log_scale(var: str, dist_info: dict[str, float | str]) -> bool: + """Validate a sampler distribution entry's optional ``log_scale`` flag. + + Log-space sampling is only defined for a strictly-positive uniform, and only + ``uniform`` entries carry usable bounds to log-transform -- V45ls guards, + enforced at the api boundary as ``SumoInputError`` and again here so direct + users of this layer cannot drift past them silently. + """ + log_scale = bool(dist_info.get("log_scale", False)) + if log_scale and dist_info["distribution"] != "uniform": + raise ValueError( + f"log_scale is only supported for uniform distributions: {var}" + ) + return log_scale + + +def scale_distribution(minimum: float, maximum: float, *, scale: str): + """The frozen distribution mapping unit probability -> value on ``scale``. + + One home for the unit->value map every value producer shares (LHS and grid + samplers, UQ draws, Sobol ppfs): linear fills the interval uniformly, log + fills it log-uniformly (``scipy.stats.loguniform``, base-agnostic since + ``ln(10^u)`` is affine in ``u``). ``scale`` is a REQUIRED argument -- + V45ls' structural tripwire: code producing values without deciding a scale + breaks at call with ``TypeError``, never silently defaults to linear. + """ + from scipy.stats import loguniform, uniform + + if scale == "log": + lo, hi = float(minimum), float(maximum) + if lo <= 0 or hi <= lo: + raise ValueError( + f"Log-scale bounds must be strictly positive and increasing, " + f"got [{lo}, {hi}]" + ) + return loguniform(a=lo, b=hi) + if scale == "linear": + return uniform(loc=minimum, scale=maximum - minimum) + raise ValueError(f"Unknown scale: {scale!r}") + + +def scale_values(values: Sequence[float] | np.ndarray, *, scale: str) -> np.ndarray: + """Map raw column values onto the scale the model treats them on.""" + array = np.asarray(values, dtype=float) + if scale == "log": + if (array <= 0).any(): + raise ValueError("log scale is undefined for values <= 0") + return np.log(array) + if scale == "linear": + return array + raise ValueError(f"Unknown scale: {scale!r}") + + def create_manual_uq_samples( input_vars: list[str], distributions: dict[str, dict[str, float | str]], @@ -494,18 +547,14 @@ def create_manual_uq_samples( matched against a preprocessor fit on original names. """ - from scipy.stats import norm, uniform + from scipy.stats import norm rng = np.random.default_rng(seed=seed) samples = {} for var in input_vars: dist_info = distributions[var] dist_type = dist_info["distribution"] - log_scale = bool(dist_info.get("log_scale", False)) - if log_scale and dist_type != "uniform": - raise ValueError( - f"log_scale is only supported for uniform distributions: {var}" - ) + log_scale = resolve_log_scale(var, dist_info) if dist_type == "normal": mean = float(dist_info["mean"]) std = float(dist_info["std"]) @@ -513,33 +562,18 @@ def create_manual_uq_samples( size=num_samples, loc=mean, scale=std, random_state=rng ).tolist() elif dist_type == "uniform": - min_val = float(dist_info["min"]) - max_val = float(dist_info["max"]) - if log_scale: - if min_val <= 0: - raise ValueError( - f"Log-scale uniform bounds must be strictly positive: {var}" - ) - # Draw uniformly in log10 space, map back to the caller's original - # units -- the surrogate's own preprocessor re-applies the log, so - # what reaches it is a uniform-in-log sample (SPEC T27fr). - log_min, log_max = np.log10(min_val), np.log10(max_val) - samples[var] = np.power( - 10, - uniform.rvs( - size=num_samples, - loc=log_min, - scale=log_max - log_min, - random_state=rng, - ), - ).tolist() - else: - samples[var] = uniform.rvs( - size=num_samples, - loc=min_val, - scale=max_val - min_val, - random_state=rng, - ).tolist() + # A log-scale uniform is drawn log-uniform in the caller's original + # units -- the surrogate's preprocessor re-applies the log, so what + # reaches it is uniform-in-log (V44ls). + samples[var] = ( + scale_distribution( + float(dist_info["min"]), + float(dist_info["max"]), + scale="log" if log_scale else "linear", + ) + .rvs(size=num_samples, random_state=rng) + .tolist() + ) elif dist_type == "constant": value = dist_info["value"] samples[var] = [float(value)] * num_samples @@ -718,6 +752,9 @@ def compute_correlation_indices( input_samples: pd.DataFrame | dict[str, list[float]], output_samples: list[float] | np.ndarray, input_vars: list[str], + *, + input_scales: dict[str, str], + output_scale: str, ) -> dict[str, dict[str, float]]: """ Compute per-input <-> output Pearson and Spearman correlation coefficients. @@ -734,13 +771,19 @@ def compute_correlation_indices( output_samples: Sample values of the output QoI, paired index-for-index with each input variable's samples (i.e. same Monte Carlo run). input_vars: Input variable names to compute correlations for. + input_scales: Per-variable ``"linear"``/``"log"`` scale; correlations run + on values mapped onto that scale (V45ls -- Pearson is not invariant to + log, so honouring scale changes it; Spearman is monotone-invariant and + therefore provably unaffected). REQUIRED, never defaulted. + output_scale: Scale of ``output_samples``. REQUIRED, never defaulted. Returns: Dict mapping each input variable name to `{"pearson": float, "spearman": float}`. Raises: ValueError: If `input_vars` is empty, a variable is missing from - `input_samples`, or sample lengths are mismatched. + `input_samples`, sample lengths are mismatched, or a log-scaled + column holds a value <= 0. """ from scipy.stats import pearsonr, spearmanr @@ -752,13 +795,15 @@ def compute_correlation_indices( col: input_samples[col].tolist() for col in input_samples.columns } - output_array = np.asarray(output_samples, dtype=float) + output_array = scale_values(output_samples, scale=output_scale) correlations: dict[str, dict[str, float]] = {} for var in input_vars: if var not in input_samples: raise ValueError(f"Input variable '{var}' not found in input samples.") - input_array = np.asarray(input_samples[var], dtype=float) + input_array = scale_values( + input_samples[var], scale=input_scales.get(var, "linear") + ) if len(input_array) != len(output_array): raise ValueError( f"Sample length mismatch for variable '{var}': " diff --git a/src/itis_sumo/evaluate/funs_evaluate.py b/src/itis_sumo/evaluate/funs_evaluate.py index 8880dd4..92206da 100644 --- a/src/itis_sumo/evaluate/funs_evaluate.py +++ b/src/itis_sumo/evaluate/funs_evaluate.py @@ -32,7 +32,9 @@ get_results, load_data, process_input_file, + resolve_log_scale, sanitize_varnames, + scale_distribution, ) _logger = logging.getLogger(__name__) @@ -1035,7 +1037,7 @@ def evaluate_sobol_indices( import math import pandas as pd - from scipy.stats import loguniform, norm, sobol_indices, uniform + from scipy.stats import norm, sobol_indices from scipy.stats.qmc import Sobol # NOTE: input_vars/distributions must stay in the caller's original @@ -1059,33 +1061,23 @@ def evaluate_sobol_indices( d_varying = len(varying_vars) - # Build frozen scipy distributions with .ppf for each varying variable. A - # log-scale variable is drawn uniformly in log space (log-uniform in the - # caller's original units), mirroring create_manual_uq_samples -- the - # surrogate's own preprocessor re-applies the log downstream. + # Build frozen scipy distributions with .ppf for each varying variable. The + # scale map is the shared scale_distribution: a log-scale variable is drawn + # log-uniform in the caller's original units (V44ls), the surrogate's + # preprocessor re-applies the log downstream. ppfs = {} for var in varying_vars: dist_info = distributions[var] dist_type = dist_info["distribution"] - log_scale = bool(dist_info.get("log_scale", False)) - if log_scale and dist_type != "uniform": - raise ValueError( - f"log_scale is only supported for uniform distributions: {var}" - ) + log_scale = resolve_log_scale(var, dist_info) if dist_type == "normal": ppfs[var] = norm(loc=dist_info["mean"], scale=dist_info["std"]) elif dist_type == "uniform": - if log_scale: - lo, hi = float(dist_info["min"]), float(dist_info["max"]) - if lo <= 0: - raise ValueError( - f"Log-scale uniform bounds must be strictly positive: {var}" - ) - ppfs[var] = loguniform(a=lo, b=hi) - else: - ppfs[var] = uniform( - loc=dist_info["min"], scale=dist_info["max"] - dist_info["min"] - ) + ppfs[var] = scale_distribution( + float(dist_info["min"]), + float(dist_info["max"]), + scale="log" if log_scale else "linear", + ) else: raise ValueError(f"Unsupported distribution type: {dist_type}") diff --git a/tests/test_correlation_indices.py b/tests/test_correlation_indices.py index 91e3c91..cfb995b 100644 --- a/tests/test_correlation_indices.py +++ b/tests/test_correlation_indices.py @@ -12,6 +12,13 @@ from itis_sumo.data.funs_data_processing import compute_correlation_indices +_LINEAR = "linear" + + +def _lin(*names: str) -> dict[str, str]: + """All-linear scale map (the correlator requires scales -- V45ls).""" + return dict.fromkeys(names, _LINEAR) + class TestComputeCorrelationIndices: """Unit tests for the pure correlation-computation function.""" @@ -23,7 +30,11 @@ def test_perfectly_correlated_variable(self): output = 3.0 * x1 + 2.0 # perfectly linear, positive correlation correlations = compute_correlation_indices( - {"x1": x1.tolist()}, output.tolist(), ["x1"] + {"x1": x1.tolist()}, + output.tolist(), + ["x1"], + input_scales=_lin("x1"), + output_scale=_LINEAR, ) assert correlations["x1"]["pearson"] == pytest.approx(1.0, abs=1e-6) @@ -36,7 +47,11 @@ def test_perfectly_anticorrelated_variable(self): output = -5.0 * x1 + 1.0 correlations = compute_correlation_indices( - {"x1": x1.tolist()}, output.tolist(), ["x1"] + {"x1": x1.tolist()}, + output.tolist(), + ["x1"], + input_scales=_lin("x1"), + output_scale=_LINEAR, ) assert correlations["x1"]["pearson"] == pytest.approx(-1.0, abs=1e-6) @@ -50,7 +65,11 @@ def test_uncorrelated_variable_near_zero(self): output = rng.uniform(-1, 1, size=n) # independent of x1 correlations = compute_correlation_indices( - {"x1": x1.tolist()}, output.tolist(), ["x1"] + {"x1": x1.tolist()}, + output.tolist(), + ["x1"], + input_scales=_lin("x1"), + output_scale=_LINEAR, ) assert abs(correlations["x1"]["pearson"]) < 0.05 @@ -68,6 +87,8 @@ def test_multiple_input_vars_one_response_per_var(self): {"x_sensitive": x_sensitive.tolist(), "x_noise": x_noise.tolist()}, output.tolist(), ["x_sensitive", "x_noise"], + input_scales=_lin("x_sensitive", "x_noise"), + output_scale=_LINEAR, ) assert set(correlations.keys()) == {"x_sensitive", "x_noise"} @@ -81,9 +102,15 @@ def test_accepts_dataframe_input(self): output = 2.0 * x1 df = pd.DataFrame({"x1": x1}) - correlations_df = compute_correlation_indices(df, output.tolist(), ["x1"]) + correlations_df = compute_correlation_indices( + df, output.tolist(), ["x1"], input_scales=_lin("x1"), output_scale=_LINEAR + ) correlations_dict = compute_correlation_indices( - {"x1": x1.tolist()}, output.tolist(), ["x1"] + {"x1": x1.tolist()}, + output.tolist(), + ["x1"], + input_scales=_lin("x1"), + output_scale=_LINEAR, ) assert correlations_df == correlations_dict @@ -91,16 +118,82 @@ def test_accepts_dataframe_input(self): def test_empty_input_vars_raises(self): """Empty input_vars list is rejected.""" with pytest.raises(ValueError, match="input_vars cannot be empty"): - compute_correlation_indices({"x1": [1.0, 2.0]}, [1.0, 2.0], []) + compute_correlation_indices( + {"x1": [1.0, 2.0]}, + [1.0, 2.0], + [], + input_scales={}, + output_scale=_LINEAR, + ) def test_missing_variable_raises(self): """Requesting a variable absent from input_samples raises ValueError.""" with pytest.raises(ValueError, match="not found in input samples"): compute_correlation_indices( - {"x1": [1.0, 2.0, 3.0]}, [1.0, 2.0, 3.0], ["x2"] + {"x1": [1.0, 2.0, 3.0]}, + [1.0, 2.0, 3.0], + ["x2"], + input_scales=_lin("x2"), + output_scale=_LINEAR, ) def test_mismatched_lengths_raises(self): """Input/output sample length mismatch raises ValueError.""" with pytest.raises(ValueError, match="Sample length mismatch"): - compute_correlation_indices({"x1": [1.0, 2.0, 3.0]}, [1.0, 2.0], ["x1"]) + compute_correlation_indices( + {"x1": [1.0, 2.0, 3.0]}, + [1.0, 2.0], + ["x1"], + input_scales=_lin("x1"), + output_scale=_LINEAR, + ) + + +class TestComputeCorrelationIndicesScale: + """Scale is a required axis of the correlator (V45ls).""" + + def test_scales_are_required_arguments(self): + """Omitting the scales breaks at call -- V45ls structural tripwire.""" + with pytest.raises(TypeError): + compute_correlation_indices( # ty: ignore[missing-argument] + {"x1": [1.0, 2.0]}, [1.0, 2.0], ["x1"] + ) + + def test_log_input_shifts_pearson_but_not_spearman(self): + """A monotone log reparametrization moves Pearson, cannot move rank.""" + rng = np.random.default_rng(11) + log_x = rng.uniform(0.0, 3.0, size=800) # x = exp(log_x), log-uniform x + x = np.exp(log_x) + # Output linear in x (not in log_x): Pearson(x, y) beats Pearson(ln x, y) + # while both share the identical rank sequence. + output = 3.0 * x + + linear = compute_correlation_indices( + {"x": x.tolist()}, + output.tolist(), + ["x"], + input_scales=_lin("x"), + output_scale=_LINEAR, + ) + log = compute_correlation_indices( + {"x": x.tolist()}, + output.tolist(), + ["x"], + input_scales={"x": "log"}, + output_scale=_LINEAR, + ) + + assert linear["x"]["pearson"] == pytest.approx(1.0, abs=1e-6) + assert log["x"]["pearson"] < 0.99 + assert log["x"]["spearman"] == pytest.approx(linear["x"]["spearman"], abs=1e-12) + + def test_log_scale_rejects_non_positive_values(self): + """A log-scaled column holding <= 0 is refused with a clear message.""" + with pytest.raises(ValueError, match="log scale is undefined"): + compute_correlation_indices( + {"x": [1.0, 0.0, 3.0]}, + [1.0, 2.0, 3.0], + ["x"], + input_scales={"x": "log"}, + output_scale=_LINEAR, + ) From 02e0e5b93f952d6b3a4e184fcfc1615285a4021a Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 23:27:44 +0200 Subject: [PATCH 10/13] test(api): V45ls flip matrix + log-scale gap tests (T46ls) - TestScaleFlipMatrix: all 10 public value-producing entry points' outputs must move when width turns log -- the machine guard against silent scale-ignore, shipped or future. - TestScaleAwareSamplers: log LHS/grid fill log-uniform, linear default unchanged, non-positive log domains refused, correlations honour log (Pearson moves, Spearman bit-identical, untouched column identical). - TestScaleGapCoverage: MOGA log objective + maximize (sign-after-log inverse order), log-variable sweep geometric in original units, log-response uncertainty multiplicative spread, Sobol mixed log+constant partition. - Existing grid test: cast result.data for repo-wide ty (fixes the prek red on #47). --- tests/test_api_workflows.py | 305 +++++++++++++++++++++++++++++++++++- 1 file changed, 304 insertions(+), 1 deletion(-) diff --git a/tests/test_api_workflows.py b/tests/test_api_workflows.py index 4e40f0b..b7479ec 100644 --- a/tests/test_api_workflows.py +++ b/tests/test_api_workflows.py @@ -11,6 +11,7 @@ import dataclasses import json import math +from collections.abc import Callable from typing import cast import numpy as np @@ -30,6 +31,8 @@ evaluate_grid, evaluate_sobol, evaluate_uncertainty, + generate_grid_samples, + generate_lhs_samples, optimize, ) @@ -353,7 +356,11 @@ def test_grid_log_response_returns_original_units(self, samples): preprocessing=_LOG_SCALE, points_per_variable=5, ) - flat = [v for row in result.data["stress"] for v in row] + flat = [ + v + for row in cast("dict[str, list[list[float]]]", result.data)["stress"] + for v in row + ] assert min(flat) > 0.0 assert max(flat) < 10 * samples[RESPONSE].max() @@ -525,3 +532,299 @@ def test_log_input_rejects_a_non_positive_lower_bound(self, samples): preprocessing=_LOG_WIDTH, seed=7, ) + + +_MOGA_DOMAINS = { + "width": DomainSpec(minimum=WIDTH_RANGE[0], maximum=WIDTH_RANGE[1]), + "height": DomainSpec(minimum=HEIGHT_RANGE[0], maximum=HEIGHT_RANGE[1]), +} +_SAMPLER_DOMAINS = { + "width": DomainSpec(minimum=1.0, maximum=100.0), + "height": DomainSpec(minimum=10.0, maximum=20.0), +} + + +class TestScaleGapCoverage: + """The edges the first log-scale pass left untested (T46ls).""" + + def test_moga_log_objective_maximize_front_in_original_units(self, samples): + result = optimize( + samples, + VARIABLES, + {RESPONSE: "maximize"}, + domains=_MOGA_DOMAINS, + max_evaluations=200, + preprocessing=_LOG_SCALE, + ) + front = result.data[RESPONSE] + assert front + # Maximizing a log objective inverts as -(ln y) -> ln y -> exp: the front + # must come back as HIGH stresses in original units. A broken sign/log + # order would return ln-space values (O(3)) or a minimized front. + assert min(front) > 0.0 + assert max(front) < 10 * samples[RESPONSE].max() + assert max(front) > float(samples[RESPONSE].mean()) + + def test_log_variable_sweep_returns_geometric_original_units(self, samples): + result = evaluate_along_axes( + samples, + VARIABLES, + RESPONSE, + preprocessing=_LOG_WIDTH, + points_per_variable=7, + ) + xs = np.asarray(result.sweeps["width"].x, dtype=float) + assert xs.min() > 0.5 + assert xs.max() <= samples["width"].max() * 1.01 + # Swept in log space, restored to original units => ln-equispaced. + diffs = np.diff(np.log(xs)) + assert np.allclose(diffs, diffs.mean(), rtol=1e-6) + + def test_uncertainty_log_response_spread_is_multiplicative(self, samples): + result = evaluate_uncertainty( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + preprocessing=_LOG_SCALE, + num_samples=300, + n_histograms=5, + seed=7, + ) + assert result.q1 > 0.0 + assert result.q1 <= result.median <= result.q3 + # Noise is injected in log space then exp-restored -> bounded relative + # (multiplicative) spread, not an additive tail around the mean. + assert 0.5 < result.q1 / result.median + assert result.q3 / result.median < 2.0 + + def test_sobol_mixed_log_and_constant(self, samples): + dists = { + "width": DistributionSpec("uniform", minimum=1.0, maximum=5.0), + "height": DistributionSpec("constant", value=300.0), + } + result = evaluate_sobol( + samples, + VARIABLES, + RESPONSE, + distributions=dists, + preprocessing=_LOG_WIDTH, + seed=7, + ) + assert set(result.indices) == set(VARIABLES) + assert result.indices["width"]["total"] > 0.8 + assert result.indices["height"]["total"] < 0.3 + + +class TestScaleAwareSamplers: + """Standalone samplers and correlations honour scale like every other + value producer (V45ls; defaults to linear, so old calls are unchanged).""" + + def test_lhs_honours_log_scale(self): + linear = generate_lhs_samples(_SAMPLER_DOMAINS, 500, seed=11) + logw = generate_lhs_samples( + _SAMPLER_DOMAINS, 500, preprocessing=_LOG_WIDTH, seed=11 + ) + widths = np.asarray(logw["width"], dtype=float) + assert widths.min() >= 1.0 and widths.max() <= 100.0 + # log-uniform in [1,100]: geometric mean at the log-midpoint. + assert float(np.exp(np.log(widths).mean())) == pytest.approx(10.0, rel=0.15) + assert float(widths.mean()) < 0.5 * float(np.mean(linear["width"])) + + def test_lhs_linear_default_unchanged(self): + first = generate_lhs_samples(_SAMPLER_DOMAINS, 200, seed=5) + second = generate_lhs_samples(_SAMPLER_DOMAINS, 200, seed=5) + pd.testing.assert_frame_equal(first, second) + widths = np.asarray(first["width"], dtype=float) + assert widths.min() >= 1.0 and widths.max() <= 100.0 + + def test_lhs_rejects_non_positive_log_domain(self): + domains = {"width": DomainSpec(minimum=0.0, maximum=100.0)} + log_width = PreprocessingSpec(overrides={"width": VariableSpec(scale="log")}) + with pytest.raises(SumoInputError, match="strictly positive"): + generate_lhs_samples(domains, 50, preprocessing=log_width, seed=1) + + def test_grid_log_axis_is_geometric(self): + from scipy.stats import loguniform + + domains = {"width": DomainSpec(minimum=1.0, maximum=100.0)} + linear = generate_grid_samples(domains, {"width": 5}) + logw = generate_grid_samples(domains, {"width": 5}, preprocessing=_LOG_WIDTH) + assert np.allclose( + np.sort(np.unique(logw["width"])), + loguniform(a=1.0, b=100.0).ppf(np.linspace(0.0, 1.0, 5)), + rtol=1e-6, + ) + assert np.allclose( + np.sort(np.unique(linear["width"])), + np.linspace(1.0, 100.0, 5), + rtol=1e-6, + ) + + def test_grid_rejects_non_positive_log_domain(self): + domains = {"width": DomainSpec(minimum=0.0, maximum=100.0)} + log_width = PreprocessingSpec(overrides={"width": VariableSpec(scale="log")}) + with pytest.raises(SumoInputError, match="strictly positive"): + generate_grid_samples(domains, {"width": 4}, preprocessing=log_width) + + def test_correlations_honour_log_scale(self, samples): + linear = compute_correlations(samples, VARIABLES, RESPONSE) + logw = compute_correlations( + samples, VARIABLES, RESPONSE, preprocessing=_LOG_WIDTH + ) + # Pearson on the log-scale column must move; the rank statistic is + # monotone-invariant; the untouched variable is bit-identical. + assert logw.coefficients["width"]["pearson"] != pytest.approx( + linear.coefficients["width"]["pearson"], rel=1e-3 + ) + assert logw.coefficients["width"]["spearman"] == pytest.approx( + linear.coefficients["width"]["spearman"], abs=1e-12 + ) + assert logw.coefficients["height"] == pytest.approx( + linear.coefficients["height"] + ) + + +class TestScaleFlipMatrix: + """V45ls behavioural enforcement in one place: EVERY public value-producing + entry point's output must move when a column turns log. A silent scale-ignore + in any path -- shipped or future -- fails here.""" + + @staticmethod + def _fp(values: object) -> list[float]: + return [float(v) for v in np.asarray(values, dtype=float).ravel()] + + def test_every_entry_point_moves_under_log(self, samples): + cases: list[tuple[str, Callable[[PreprocessingSpec | None], list[float]]]] = [ + ( + "cross_validate", + lambda p: self._fp( + cross_validate( + samples, VARIABLES, RESPONSE, preprocessing=p, seed=7 + ).predicted + ), + ), + ( + "evaluate_cv_metrics", + lambda p: self._fp( + evaluate_cv_metrics( + samples, VARIABLES, RESPONSE, preprocessing=p, seed=7 + ).root_mean_squared + ), + ), + ( + "evaluate_along_axes", + lambda p: self._fp( + [ + v + for sweep in evaluate_along_axes( + samples, + VARIABLES, + RESPONSE, + preprocessing=p, + points_per_variable=5, + ).sweeps.values() + for v in sweep.predicted + ] + ), + ), + ( + "evaluate_grid", + lambda p: self._fp( + cast( + "dict[str, list[list[float]]]", + evaluate_grid( + samples, + VARIABLES, + RESPONSE, + grid_variables=VARIABLES, + preprocessing=p, + points_per_variable=5, + ).data, + )[RESPONSE] + ), + ), + ( + "evaluate_uncertainty", + lambda p: self._fp( + evaluate_uncertainty( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + preprocessing=p, + num_samples=150, + n_histograms=4, + seed=7, + ).mean + ), + ), + ( + "evaluate_sobol", + lambda p: self._fp( + [ + entry["total"] + for entry in evaluate_sobol( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + preprocessing=p, + seed=7, + ).indices.values() + ] + ), + ), + ( + "optimize", + lambda p: sorted( + self._fp( + optimize( + samples, + VARIABLES, + {RESPONSE: "minimize"}, + domains=_MOGA_DOMAINS, + max_evaluations=150, + preprocessing=p, + ).data[RESPONSE] + ) + ), + ), + ( + "compute_correlations", + lambda p: self._fp( + [ + entry["pearson"] + for entry in compute_correlations( + samples, VARIABLES, RESPONSE, preprocessing=p + ).coefficients.values() + ] + ), + ), + ( + "generate_lhs_samples", + lambda p: self._fp( + generate_lhs_samples( + _SAMPLER_DOMAINS, 200, preprocessing=p, seed=7 + )["width"] + ), + ), + ( + "generate_grid_samples", + lambda p: self._fp( + generate_grid_samples( + _SAMPLER_DOMAINS, + {"width": 5, "height": 4}, + preprocessing=p, + )["width"] + ), + ), + ] + assert len(cases) == 10 + for name, fingerprint in cases: + linear = fingerprint(None) + log = fingerprint(_LOG_WIDTH) + assert log != pytest.approx(linear, rel=1e-9, nan_ok=True), ( + f"{name}: output is identical under log scale -- the scale " + "override was silently ignored (V45ls)" + ) From 1dfdc6fcd2565de203e3e511dfe4f582d83b46ca Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Mon, 28 Sep 2026 23:27:44 +0200 Subject: [PATCH 11/13] docs: V&V Category I scale semantics + tier-1 scale contracts (T46ls) verification-validation.md: new Category I (I1-I8, incl. H6 preprocessor log round-trip), status counts 245->359 / 26->30, engine refs 1.5.9/6.20 -> 6.24.7, summary A-H -> A-I. TIER plan: section 7 documents the scale helper contracts (required-scale tripwires, log guards, delta-method std inverse). SPEC: T46ls x. --- SPEC.md | 2 +- docs/TIER1_TIER2_UNIT_TESTS_PLAN.md | 26 +++++++++++++++++++++ docs/verification-validation.md | 36 +++++++++++++++++++++++++---- 3 files changed, 58 insertions(+), 6 deletions(-) diff --git a/SPEC.md b/SPEC.md index aba480a..f8166ee 100644 --- a/SPEC.md +++ b/SPEC.md @@ -161,7 +161,7 @@ T33zz|✓|swap itis-sumo LICENSE + `pyproject.toml` `license`/classifiers to mat T34aa|✓|`make publish-testpypi-dev` ! source `TESTPYPI_TOKEN` from local `.env` (gitignored) instead of requiring a pre-exported shell var|§C T35cc|✓|gate `publish`/`release` jobs behind explicit `workflow_dispatch`; `build`+`verify` stay CI-only for release tags; feature `.devN` uploads move to local Make target|V31vp,V32bb T36dd|✓|add `.env` to `.gitignore`; make `publish-testpypi` source `.env` without printing token, then build/check/publish|V37bb -T46ls|.|scale first-class across the WHOLE api surface: `generate_lhs_samples` ∧ `generate_grid_samples` ∧ `compute_correlations` take `preprocessing` ∧ honour scale; one `scale_distribution` unit→value helper absorbs the log10-power hand-roll ∧ the duplicate log guards ∧ the `_is_log` closure; V45ls linear↔log flip-tests over ∀ 10 entry points + 5 gap tests (MOGA log·maximize, log-variable axis x-restoration, log-response UQ spread, Sobol mixed partition, repo-wide `ty`); rows land in the V&V report Category I|V45ls,V44ls,V21pf,V22rs +T46ls|x|scale first-class across the WHOLE api surface: `generate_lhs_samples` ∧ `generate_grid_samples` ∧ `compute_correlations` take `preprocessing` ∧ honour scale; one `scale_distribution` unit→value helper absorbs the log10-power hand-roll ∧ the duplicate log guards ∧ the `_is_log` closure; V45ls linear↔log flip-tests over ∀ 10 entry points + 5 gap tests (MOGA log·maximize, log-variable axis x-restoration, log-response UQ spread, Sobol mixed partition, repo-wide `ty`); rows land in the V&V report Category I|V45ls,V44ls,V21pf,V22rs T37ef|✓|validate manual publish tag + artifact version before PyPI upload|V35wk T38dd|✓|make `publish-testpypi-dev` auto-compute/write `.devN`, build/check/upload directly to TestPyPI, then restore project version; CI verifies alpha/beta/rc before real PyPI; no dev tag cascade|V38cc,V39qf T39pk|.|POST-PORT/deepen candidate: `itis_sumo.api.workflows` repeats the same build-session → translate-config → call → translate-result → wrap-errors skeleton per function; pull the shared shape down into `_session.py` so each `workflows.py` function shrinks to signature+one call — behavior-preserving refactor, tests green before/after; natural precursor to T26eq's handle extraction|V16qf,V27fq,T26eq diff --git a/docs/TIER1_TIER2_UNIT_TESTS_PLAN.md b/docs/TIER1_TIER2_UNIT_TESTS_PLAN.md index 4807ac7..c602ecb 100644 --- a/docs/TIER1_TIER2_UNIT_TESTS_PLAN.md +++ b/docs/TIER1_TIER2_UNIT_TESTS_PLAN.md @@ -109,6 +109,32 @@ Supporting utility used across the workflow above, not a workflow stage on its o **`DataPreprocessor`** — roundtrip fidelity (see Tier 2, P3/P6) +### 7. Scale semantics helpers (`funs_data_processing.py`, `api/_session.py`) + +The `unit→value` map shared by every value producer (SPEC V44ls/V45ls). + +**`resolve_log_scale(var, dist_info) -> bool`** +- `log_scale=True` on a non-`uniform` entry → `ValueError` naming the variable +- absent flag → `False` (linear default) + +**`scale_distribution(minimum, maximum, *, scale)`** +- `scale` is a REQUIRED keyword — omitting it raises `TypeError` (structural tripwire) +- `"linear"` → `uniform(loc=min, scale=max-min)` (bit-identical to the pre-scale sampler math) +- `"log"` → `loguniform(a=min, b=hi)`: geometric ppf/rvs, geometric mean at the log-midpoint +- `min <= 0` or `max <= min` under log → `ValueError` ("strictly positive and increasing") +- unknown scale string → `ValueError` + +**`scale_values(values, *, scale) -> np.ndarray`** +- `"log"` of any `<= 0` → `ValueError("log scale is undefined...")`; otherwise `np.log` +- `"linear"` → unchanged array + +**`compute_correlation_indices(..., *, input_scales, output_scale)`** +- scales required (tripwire test); log-reparametrized input shifts Pearson, leaves Spearman bit-identical + +**`create_manual_uq_samples`** — `log_scale` uniform drawn log-uniform in original units; log+normal and log+min≤0 refused + +**`DataPreprocessor` log transform** — `setup_log_transform`/fit/transform/inverse round-trip; delta-method `inverse_transform_output_std` (`tests/test_data_preprocessor.py::TestLogTransform`) + --- ## Tier 2: Property-Based / Invariant Tests diff --git a/docs/verification-validation.md b/docs/verification-validation.md index 4320900..720e189 100644 --- a/docs/verification-validation.md +++ b/docs/verification-validation.md @@ -17,11 +17,11 @@ Live results as of the last run of the ported V&V suite: | Suite | Tests | Result | |---|---|---| -| Full standalone suite (`uv run pytest`) | 245 | **all passing** | -| Analytical/integration tier (`-m analytical`, real Dakota subprocess, no mocking) | 26 | **all passing** | +| Full standalone suite (`uv run pytest`) | 359 | **all passing** | +| Analytical/integration tier (`-m analytical`, real Dakota subprocess, no mocking) | 30 | **all passing** | | Sobol' / Ishigami acceptance gate (`test_sobol_indices.py`) | 4 | **all passing** | -The analytical tier spawns a real `itis-dakota` (1.5.9 / Dakota 6.20) +The analytical tier spawns a real `itis-dakota` (6.24.7) process per test — nothing here is mocked. Re-run locally with: ```sh @@ -131,6 +131,32 @@ See [Data preprocessing](reference/preprocess.md). | H1-H3 | Z-score / min-max / sign-switch round-trip to 1e-10 | ✅ | | H4 | Round-trip on 1000×20 dataset | ✅ | | H5 | Normalization improves accuracy on badly-scaled `f(x) = 1000x+1` | ✅ | +| H6 | `log_transform` round-trip (`np.log`/`np.exp`), non-positive values refused, delta-method std inverse `std_orig ≈ |y_hat|·std_log` | ✅ | + +## Category I: Scale semantics (log) + +`scale` is a first-class, un-forgettable axis (SPEC V44ls/V45ls): every public +value-producing entry point accepts it (defaulting to linear, so existing calls +are unchanged) and every value the api yields is computed under it. The one +`unit→value` map behind all of it is `scale_distribution` (`uniform` vs +`scipy.stats.loguniform`), and it requires the scale argument — unwired code +breaks with `TypeError`, it can never silently default. + +| Test | What | Status | +|---|---|---| +| I1 | LHS + grid samplers: log domain ⇒ log-uniform/geometric fill; linear default bit-identical; non-positive log domain ⇒ `SumoInputError` | ✅ | +| I2 | Correlations: Pearson moves under log, Spearman provably unchanged (monotone-invariant), untouched columns bit-identical; correlator's scale args are required (`TypeError` tripwire) | ✅ | +| I3 | Surrogate / CV / along-axes / grid eval: log response exp-restored to original units; non-positive training outputs rejected pre-Dakota | ✅ | +| I4 | UQ propagation: log inputs drawn log-uniform (response skews low vs linear, directionally asserted); log+normal / log+min≤0 rejected; log response ⇒ multiplicative (not additive) spread | ✅ | +| I5 | CV accuracy metrics: inherit log through `cross_validate` (metrics differ from linear); reject non-positive log responses | ✅ | +| I6 | MOGA: log variable explored in ln-space (domain mapped, positivity-guarded); log objective exp-restored for **both** minimize and maximize (sign-after-log inverse order verified) | ✅ | +| I7 | Sobol: log input shifts the variance decomposition in the expected direction (compressed variable explains less); mixed log+constant partition | ✅ | +| I8 | Flip matrix: all 10 public value-producing entry points' outputs move when a column turns log — the V45ls machine guard against any silent scale-ignore, shipped or future | ✅ | + +Tests: `tests/test_api_workflows.py` (`TestLogScale*`, `TestScaleGapCoverage`, +`TestScaleAwareSamplers`, `TestScaleFlipMatrix`), +`tests/test_correlation_indices.py`, `tests/test_dakota_funs_data_processing.py`, +`tests/test_data_preprocessor.py`. ## Ishigami analytical acceptance gate @@ -161,7 +187,7 @@ first-order-only check and fail this one. ## Known limitations Carried over from the original test plan, still true of the pinned engine -(`itis-dakota==1.5.9`, Dakota 6.20): +(`itis-dakota==6.24.7`): - **Built-in Dakota CV parsing is unreliable** — `log_output` comes back hardcoded empty on some study configurations, which is why @@ -185,7 +211,7 @@ Carried over from the original test plan, still true of the pinned engine ## Summary -Every category in the original V&V plan (A through H, plus the Ishigami +Every category in the original V&V plan (A through I, plus the Ishigami acceptance gate) currently passes against its analytical reference, with real (unmocked) Dakota execution for every category that requires the engine. The known limitations above are pre-existing engine/pipeline From 5d5292fd1a6f19e27b2b40188693c5b8b271ad51 Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Tue, 29 Sep 2026 08:54:49 +0200 Subject: [PATCH 12/13] feat(api): evaluate_correlations -- MC-through-surrogate correlation (T47mt) The 11th value-producing entry point. The web UI's correlation workflow (/compute_correlation_indices, #470) draws Monte Carlo samples from caller distributions, predicts them through the fitted surrogate, and correlates each input against the prediction over the SHARED sample set. Until now the api only offered table-mode compute_correlations, so a consumer migration would either keep the computation in flaskapi or silently degrade to correlating the training table (what superseded branch #537 did unnoticed). - evaluate/funs_evaluate: correlate_manual_uq_samples -- the endpoint's sampling chain, with input_scales/output_scale REQUIRED (V45ls tripwire; no erfinv noise -- Pearson/Spearman run on the surrogate's own prediction). - SumoSession.correlations: exact-cover guard + log flags via the existing _uq_engine_distributions (log requires uniform with min > 0). - api.evaluate_correlations: one-shot workflow mirroring evaluate_uncertainty. - CorrelationResult gains seed (default None; table mode unchanged). Tests: dominant-variable recovery, seed reproducibility, log-mode movement, guard rejections, TypeError tripwire, __all__ surface, flip matrix now over all 11 entry points. Suite 365 passed; ruff/ty clean. --- src/itis_sumo/api/__init__.py | 2 + src/itis_sumo/api/_session.py | 43 +++++++ src/itis_sumo/api/types.py | 1 + src/itis_sumo/api/workflows.py | 48 +++++++ src/itis_sumo/data/funs_data_processing.py | 4 +- src/itis_sumo/evaluate/funs_evaluate.py | 72 ++++++++++- tests/test_api_contract.py | 1 + tests/test_api_workflows.py | 138 ++++++++++++++++++++- 8 files changed, 305 insertions(+), 4 deletions(-) diff --git a/src/itis_sumo/api/__init__.py b/src/itis_sumo/api/__init__.py index f2b5d8a..40fc761 100644 --- a/src/itis_sumo/api/__init__.py +++ b/src/itis_sumo/api/__init__.py @@ -42,6 +42,7 @@ compute_correlations, cross_validate, evaluate_along_axes, + evaluate_correlations, evaluate_cv_metrics, evaluate_grid, evaluate_sobol, @@ -75,6 +76,7 @@ "compute_correlations", "cross_validate", "evaluate_along_axes", + "evaluate_correlations", "evaluate_cv_metrics", "evaluate_grid", "evaluate_sobol", diff --git a/src/itis_sumo/api/_session.py b/src/itis_sumo/api/_session.py index a8dc8ff..02a44f4 100644 --- a/src/itis_sumo/api/_session.py +++ b/src/itis_sumo/api/_session.py @@ -36,6 +36,7 @@ from itis_sumo.api.types import ( AlongAxesResult, AxisSweep, + CorrelationResult, CrossValidationResult, Direction, DistributionSpec, @@ -48,6 +49,7 @@ VariableSpec, ) from itis_sumo.evaluate.funs_evaluate import ( + correlate_manual_uq_samples, evaluate_sobol_indices, evaluate_sumo_along_axes, evaluate_sumo_manual_crossvalidation, @@ -507,6 +509,47 @@ def uncertainty( maximum=summary["max"], ) + def correlations( + self, + *, + distributions: Mapping[str, DistributionSpec], + num_samples: int, + seed: int, + ) -> CorrelationResult: + """Correlate each variable with the surrogate-predicted response over a + shared Monte Carlo sample set drawn from ``distributions``. + + Scale comes from the session's own spec (V45ls): each column's samples + and the prediction are correlated on their declared scale. + """ + missing = sorted(set(self._variables) - set(distributions)) + unknown = sorted(set(distributions) - set(self._variables)) + if missing or unknown: + raise SumoInputError( + f"Distributions must cover variables exactly; missing={missing}, " + f"unknown={unknown}" + ) + engine_distributions = self._uq_engine_distributions(distributions) + coefficients = self._run_engine( + "correlating through the surrogate", + correlate_manual_uq_samples, + self._run_dir, + self._training_file, + self._variables, + self._response, + engine_distributions, + self._preprocessor, + num_samples, + input_scales={ + variable: self._scale_of(variable) for variable in self._variables + }, + output_scale=self._scale_of(self._response), + seed=seed, + ) + return CorrelationResult( + response=self._response, seed=seed, coefficients=coefficients + ) + def _uq_engine_distributions( self, distributions: Mapping[str, DistributionSpec] ) -> dict[str, dict[str, float | str]]: diff --git a/src/itis_sumo/api/types.py b/src/itis_sumo/api/types.py index 2afe541..c9b4dac 100644 --- a/src/itis_sumo/api/types.py +++ b/src/itis_sumo/api/types.py @@ -149,6 +149,7 @@ class CorrelationResult: response: str coefficients: dict[str, dict[str, float]] + seed: int | None = None @dataclass(frozen=True) diff --git a/src/itis_sumo/api/workflows.py b/src/itis_sumo/api/workflows.py index 62e8e4b..47657b5 100644 --- a/src/itis_sumo/api/workflows.py +++ b/src/itis_sumo/api/workflows.py @@ -212,6 +212,54 @@ def compute_correlations( return CorrelationResult(response=response, coefficients=coefficients) +def evaluate_correlations( + samples: pd.DataFrame, + variables: Sequence[str], + response: str, + *, + distributions: Mapping[str, DistributionSpec], + num_samples: int = 1000, + seed: int = DEFAULT_SEED, + preprocessing: PreprocessingSpec | None = None, + workspace: Path | None = None, +) -> CorrelationResult: + """Correlate a response against every variable over the SAME Monte Carlo + sample set the surrogate predicts. + + Draws ``num_samples`` from ``distributions``, evaluates the fitted surrogate + once, and reports Pearson + Spearman of each variable's samples against the + prediction (original units, original names). This is the sampling-through- + the-surrogate correlation the web UI needs -- distinct from + ``compute_correlations``, which correlates columns of a table the caller + already owns. + + Args: + samples: Training samples (the surrogate is fitted on these). + variables: Input variable columns. + response: Response column, named as in the table. + distributions: Per-variable Monte Carlo distributions; must cover the + variables exactly. + num_samples: Monte Carlo sample size. + seed: Controls the draw and the surrogate. + preprocessing: Per-column scale overrides (V45ls): each column's samples + and the prediction are correlated on their declared scale, so a + log-scale variable's Pearson moves while its Spearman is provably + unchanged. + workspace: If given, working files are written here and kept. If + omitted, they are discarded on success and kept on failure. + + Raises: + SumoInputError: Distributions do not cover the variables exactly, or a + log-scale variable's distribution cannot support log sampling. + """ + with SumoSession( + samples, variables, response, preprocessing=preprocessing, workspace=workspace + ) as session: + return session.fit().correlations( + distributions=distributions, num_samples=num_samples, seed=seed + ) + + def evaluate_cv_metrics( samples: pd.DataFrame, variables: Sequence[str], diff --git a/src/itis_sumo/data/funs_data_processing.py b/src/itis_sumo/data/funs_data_processing.py index a477960..2dbe0e0 100644 --- a/src/itis_sumo/data/funs_data_processing.py +++ b/src/itis_sumo/data/funs_data_processing.py @@ -3,7 +3,7 @@ import logging import os import re -from collections.abc import Callable, Sequence +from collections.abc import Callable, Mapping, Sequence from pathlib import Path from typing import Literal, TypeVar, overload @@ -753,7 +753,7 @@ def compute_correlation_indices( output_samples: list[float] | np.ndarray, input_vars: list[str], *, - input_scales: dict[str, str], + input_scales: Mapping[str, str], output_scale: str, ) -> dict[str, dict[str, float]]: """ diff --git a/src/itis_sumo/evaluate/funs_evaluate.py b/src/itis_sumo/evaluate/funs_evaluate.py index 92206da..b1c0712 100644 --- a/src/itis_sumo/evaluate/funs_evaluate.py +++ b/src/itis_sumo/evaluate/funs_evaluate.py @@ -2,7 +2,7 @@ import os import re import shutil -from collections.abc import Callable +from collections.abc import Callable, Mapping from pathlib import Path from typing import Literal @@ -23,6 +23,7 @@ from itis_sumo.core.dakota_object import DakotaObject from itis_sumo.core.sumo_model_store import stage_model_for_import, store_exported_model from itis_sumo.data.funs_data_processing import ( + compute_correlation_indices, create_grid_samples, create_manual_uq_samples, create_samples_along_axes, @@ -252,6 +253,75 @@ def propagate_manual_uq_with_uncertainty( ) +def correlate_manual_uq_samples( + run_dir: Path, + PROCESSED_TRAINING_FILE: Path, + input_vars: list[str], + output_response: str, + distributions: dict[str, dict[str, float | str]], + preprocessor, + num_samples: int, + *, + input_scales: Mapping[str, str], + output_scale: str, + seed: int = 42, +) -> dict[str, dict[str, float]]: + """Correlate each input variable with the surrogate-predicted response over + a shared Monte Carlo sample set. + + Draws ``num_samples`` per-variable samples from ``distributions`` (in the + caller's original units), evaluates the surrogate once, then correlates + every input's samples against the inverse-transformed prediction -- the + workflow behind the historical ``/flask/dakota/compute_correlation_indices`` + route (#470), owned here so no consumer re-implements the sampling/ + surrogate/correlation chain. + + Correlations run on values mapped onto each column's own scale (V45ls): + ``input_scales``/``output_scale`` are REQUIRED keyword arguments -- a + producer of values cannot forget to decide a scale (``TypeError``, never a + silent linear default). Pearson moves under a log reparametrization; + Spearman is monotone-invariant by construction. + """ + samples = create_manual_uq_samples(input_vars, distributions, num_samples, seed) + df_samples = pd.DataFrame(samples) + SAMPLES_FILE = run_dir / "correlation_samples.csv" + df_samples.to_csv(SAMPLES_FILE, index=False) + + df_samples_transformed = preprocessor.transform(df_samples) + PROCESSED_SAMPLES_FILE = run_dir / "correlation_samples_processed.csv" + df_samples_transformed.to_csv(PROCESSED_SAMPLES_FILE, sep=" ", index=False) + + mapped_input_vars = [ + preprocessor.input_variables[var].mapped_name for var in input_vars + ] + mapped_response = preprocessor.output_variables[output_response].mapped_name + results = evaluate_sumo( + run_dir, + PROCESSED_TRAINING_FILE, + PROCESSED_SAMPLES_FILE, + mapped_input_vars, + mapped_response, + ) + + prediction_key = mapped_response + "_hat" + if prediction_key not in results: + raise ValueError( + f"Cannot compute correlation indices without '{prediction_key}' " + f"predictions. Available keys: {list(results.keys())}." + ) + + predictions_original = preprocessor.inverse_transform( + {mapped_response: results[prediction_key]} + ) + return compute_correlation_indices( + df_samples, + predictions_original[output_response], + input_vars, + input_scales=input_scales, + output_scale=output_scale, + ) + + def summarize_uncertainty_samples( values: np.ndarray, num_bins: int | None = None ) -> dict: diff --git a/tests/test_api_contract.py b/tests/test_api_contract.py index 2605937..0f1fa01 100644 --- a/tests/test_api_contract.py +++ b/tests/test_api_contract.py @@ -76,6 +76,7 @@ def test_exports_only_the_documented_names(self): "compute_correlations", "cross_validate", "evaluate_along_axes", + "evaluate_correlations", "evaluate_grid", "evaluate_sobol", "evaluate_cv_metrics", diff --git a/tests/test_api_workflows.py b/tests/test_api_workflows.py index b7479ec..bb41a8d 100644 --- a/tests/test_api_workflows.py +++ b/tests/test_api_workflows.py @@ -12,6 +12,7 @@ import json import math from collections.abc import Callable +from pathlib import Path from typing import cast import numpy as np @@ -27,6 +28,7 @@ compute_correlations, cross_validate, evaluate_along_axes, + evaluate_correlations, evaluate_cv_metrics, evaluate_grid, evaluate_sobol, @@ -685,6 +687,123 @@ def test_correlations_honour_log_scale(self, samples): ) +class TestEvaluateCorrelations: + """`evaluate_correlations` owns the MC-through-surrogate correlation + workflow (#470) so no consumer re-implements the sampling chain (V45ls).""" + + def test_recovers_dominant_variable_over_shared_sample_set(self, samples): + result = evaluate_correlations( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + num_samples=300, + seed=7, + ) + assert result.response == RESPONSE + assert result.seed == 7 + assert set(result.coefficients) == set(VARIABLES) + # stress = 3*width + 0.01*height: width dominates on the shared set. + assert abs(result.coefficients["width"]["pearson"]) > 0.9 + assert abs(result.coefficients["width"]["pearson"]) > abs( + result.coefficients["height"]["pearson"] + ) + + def test_is_seed_reproducible(self, samples): + first = evaluate_correlations( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + num_samples=150, + seed=11, + ) + second = evaluate_correlations( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + num_samples=150, + seed=11, + ) + assert first.coefficients == second.coefficients + + def test_log_scale_moves_the_coefficients(self, samples): + # NOTE: unlike table-mode correlation, log mode ALSO changes the draw + # (log-uniform vs uniform), so only "the numbers move" is guaranteed -- + # no rank-invariance claim here. + linear = evaluate_correlations( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + num_samples=300, + seed=7, + ) + logw = evaluate_correlations( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + num_samples=300, + preprocessing=_LOG_WIDTH, + seed=7, + ) + assert logw.coefficients["width"]["pearson"] != pytest.approx( + linear.coefficients["width"]["pearson"], rel=1e-3 + ) + + def test_log_variable_requires_uniform_positive_support(self, samples): + non_uniform = { + "width": DistributionSpec("normal", mean=3.0, std=0.5), + "height": _UQ_DISTS["height"], + } + with pytest.raises(SumoInputError, match="only a uniform supports log"): + evaluate_correlations( + samples, + VARIABLES, + RESPONSE, + distributions=non_uniform, + num_samples=50, + preprocessing=_LOG_WIDTH, + seed=7, + ) + non_positive = { + "width": DistributionSpec("uniform", minimum=0.0, maximum=5.0), + "height": _UQ_DISTS["height"], + } + with pytest.raises(SumoInputError, match="strictly positive"): + evaluate_correlations( + samples, + VARIABLES, + RESPONSE, + distributions=non_positive, + num_samples=50, + preprocessing=_LOG_WIDTH, + seed=7, + ) + + def test_distributions_must_cover_variables_exactly(self, samples): + with pytest.raises(SumoInputError, match="cover variables exactly"): + evaluate_correlations( + samples, + VARIABLES, + RESPONSE, + distributions={"width": _UQ_DISTS["width"]}, + num_samples=50, + seed=7, + ) + + def test_producer_requires_scales(self): + """V45ls structural tripwire on the new value-producing entry point.""" + from itis_sumo.evaluate.funs_evaluate import correlate_manual_uq_samples + + with pytest.raises(TypeError): + correlate_manual_uq_samples( # ty: ignore[missing-argument] + Path("."), Path("."), ["x"], "y", {}, None, 10, seed=1 + ) + + class TestScaleFlipMatrix: """V45ls behavioural enforcement in one place: EVERY public value-producing entry point's output must move when a column turns log. A silent scale-ignore @@ -790,6 +909,23 @@ def test_every_entry_point_moves_under_log(self, samples): ) ), ), + ( + "evaluate_correlations", + lambda p: self._fp( + [ + entry["pearson"] + for entry in evaluate_correlations( + samples, + VARIABLES, + RESPONSE, + distributions=_UQ_DISTS, + num_samples=150, + preprocessing=p, + seed=7, + ).coefficients.values() + ] + ), + ), ( "compute_correlations", lambda p: self._fp( @@ -820,7 +956,7 @@ def test_every_entry_point_moves_under_log(self, samples): ), ), ] - assert len(cases) == 10 + assert len(cases) == 11 for name, fingerprint in cases: linear = fingerprint(None) log = fingerprint(_LOG_WIDTH) From 92ee066b08ce32c832884e59377793e342a97c7d Mon Sep 17 00:00:00 2001 From: Javier Garcia Ordonez Date: Tue, 29 Sep 2026 08:54:49 +0200 Subject: [PATCH 13/13] docs(spec): T47mt + B18mt, V45ls entry count 11, V&V Category I9 B18mt records why the surface had a hole (T25dp landed only table-mode correlation) so no future consumer migration re-degrades the endpoint. V&V report: flip-matrix row says 11, new I9 for evaluate_correlations, counts 359->365. TIER plan documents the new producer's contracts. --- SPEC.md | 4 +++- docs/TIER1_TIER2_UNIT_TESTS_PLAN.md | 3 +++ docs/verification-validation.md | 5 +++-- 3 files changed, 9 insertions(+), 3 deletions(-) diff --git a/SPEC.md b/SPEC.md index f8166ee..e268522 100644 --- a/SPEC.md +++ b/SPEC.md @@ -110,7 +110,7 @@ V41hp: `ty check` ! zero diagnostics repo-wide (blocking in both the local `prek V19nd: ∀ Dependabot ecosystem entry in `.github/dependabot.yml` → weekly Monday 03:00 Europe/Zurich schedule ∧ exactly one wildcard dependency group; ⊥ unspecified default schedule or one PR per dependency V20qx: ∀ ignored Dependabot dependency → ignore reason names active compatibility constraint ∧ references tracked resolution task; `itis-dakota` remains ignored until T16mo resolves the Dakota 6.23+ interface-cache regression V44ls: ∀ scale-consuming `itis_sumo.api` workflow (surrogate fit ∧ UQ/Sobol input sampling ∧ MOGA search domain ∧ CV accuracy metric) → a `scale="log"` override ! be honoured in that workflow's own computation (log-space surrogate + log-space sampling/search where a distribution or domain is involved, positivity-guarded); ⊥ a `scale` override accepted but silently ignored outside the surrogate fit -V45ls: ∀ value-producing `itis_sumo.api` entry point → `scale` ! be threaded into the computation that yields those values ∧ observable in the result; ⊥ a value the api produced while ignoring the caller's scale; enforced: structural (internal value-producers `scale_distribution`/`resolve_log_scale` take scale/flag as a required arg — unwired code breaks at call w/ `TypeError`, never silent drift) ∧ behavioural (linear↔log flip-test over ∀ 10 entry points; rank-correlation exempt — monotone-invariant, asserted unchanged) +V45ls: ∀ value-producing `itis_sumo.api` entry point → `scale` ! be threaded into the computation that yields those values ∧ observable in the result; ⊥ a value the api produced while ignoring the caller's scale; enforced: structural (internal value-producers `scale_distribution`/`resolve_log_scale` take scale/flag as a required arg — unwired code breaks at call w/ `TypeError`, never silent drift) ∧ behavioural (linear↔log flip-test over ∀ entry points (10 table-mode workflows + `evaluate_correlations`); rank-correlation exempt — monotone-invariant, asserted unchanged) ## §R R1: `export_model`/`import_model` child keywords; formats `text_archive`(.sps)/`binary_archive`(.bsps)/`algebraic_file`(.alg); naming `{prefix}.{resp}.{ext}` | branch R2 @@ -162,6 +162,7 @@ T34aa|✓|`make publish-testpypi-dev` ! source `TESTPYPI_TOKEN` from local `.env T35cc|✓|gate `publish`/`release` jobs behind explicit `workflow_dispatch`; `build`+`verify` stay CI-only for release tags; feature `.devN` uploads move to local Make target|V31vp,V32bb T36dd|✓|add `.env` to `.gitignore`; make `publish-testpypi` source `.env` without printing token, then build/check/publish|V37bb T46ls|x|scale first-class across the WHOLE api surface: `generate_lhs_samples` ∧ `generate_grid_samples` ∧ `compute_correlations` take `preprocessing` ∧ honour scale; one `scale_distribution` unit→value helper absorbs the log10-power hand-roll ∧ the duplicate log guards ∧ the `_is_log` closure; V45ls linear↔log flip-tests over ∀ 10 entry points + 5 gap tests (MOGA log·maximize, log-variable axis x-restoration, log-response UQ spread, Sobol mixed partition, repo-wide `ty`); rows land in the V&V report Category I|V45ls,V44ls,V21pf,V22rs +T47mt|✓|`api.evaluate_correlations`: move the #470 MC-through-surrogate correlation workflow (draw → surrogate predict → Pearson+Spearman over the SHARED sample set) out of the flaskapi endpoint into the api — scale-native from birth (engine producer `correlate_manual_uq_samples` takes `input_scales`/`output_scale` as required kwargs); `CorrelationResult.seed` records the draw; flip matrix extended over ∀ 11 entry points|V45ls,V22rs T37ef|✓|validate manual publish tag + artifact version before PyPI upload|V35wk T38dd|✓|make `publish-testpypi-dev` auto-compute/write `.devN`, build/check/upload directly to TestPyPI, then restore project version; CI verifies alpha/beta/rc before real PyPI; no dev tag cascade|V38cc,V39qf T39pk|.|POST-PORT/deepen candidate: `itis_sumo.api.workflows` repeats the same build-session → translate-config → call → translate-result → wrap-errors skeleton per function; pull the shared shape down into `_session.py` so each `workflows.py` function shrinks to signature+one call — behavior-preserving refactor, tests green before/after; natural precursor to T26eq's handle extraction|V16qf,V27fq,T26eq @@ -195,3 +196,4 @@ B15tg|2026-08-27|the `v0.1.0a2` tag (created by auto-tag.yml on the develop merg B16mo|2026-08-28|T16mo rung 2 bumped the engine `itis-dakota==1.5.11` (Dakota 6.20) → `==6.24.7` (Dakota 6.24). 6.24.7 still carries the 6.23+ `Interface::interface_cache()` regression (R5): `dakota.environment.study()`'s ctor unconditionally calls `interface_cache(problem_db)` and throws `IndexError: map::at` ("called with nonexistent study!") for itis-sumo's pure data-fit surrogate confs, which construct no Interface (only `import_build_points_file`, no `interface` block) — the cache static map is never populated. All 45 real-Dakota surrogate tests failed on it (training, cross-validation, UQ propagation, MOGA, import/export, evaluation). Upstream fix (get-or-create `operator[]`) is still pending|shim in itis-sumo, not waiting on upstream: `config/funs_create_dakota_conf.py::add_r5_interface_cache_workaround()` (appended to every conf via `start_dakota_file()`) declares an unused `single` model `R5_WA_MODEL` wired to a no-op `fork` interface (`analysis_drivers='true'`) and points each surrogate `model` at it via `truth_model_pointer='R5_WA_MODEL'`, forcing the Interface to be constructed (populating the cache) during study setup; the dummy model is never referenced by any method/iterator so the interface is never executed. 318-test suite green on 6.24.7. Took 6.24.7 (not 6.24.9) because 6.24.8+ dropped cp311 wheels, which would have forced a Python 3.12 migration of both repos; R5 B8hm|2026-09-11|Dependabot entries used unspecified weekly timing, and `itis-dakota` was intentionally ignored without a durable policy invariant|V19nd,V20qx B17sc|2026-09-28|`scale="log"` on `PreprocessingSpec.overrides` reached the surrogate fit but was a near-no-op for UQ/MOGA/cv-metrics/Sobol — the logscale port (T27fr) landed the preprocessor transform but not per-workflow log-space sampling/search, so `evaluate_uncertainty` with a log input returned linear-space statistics (probe: mean 12.057 linear vs 12.058 log). Symptom never seen in flaskapi, whose frontend drove log at the request-payload layer|wire log-space into every scale-consuming path: `create_manual_uq_samples` log_scale branch (log-uniform, uniform-only + positivity), `_uq_engine_distributions` flags ∧ validates log vars for UQ + Sobol, `evaluate_sobol_indices` samples `scipy.stats.loguniform`, `optimize_pareto_front` maps a log variable's search domain into ln space ∧ fits log objectives; `evaluate_cv_metrics` inherits log via `cross_validate`. Guarded by V44ls ∧ tests across the four workflows|V44ls,T27fr +B18mt|2026-09-29|T25dp's "correlation" port landed only table-mode `api.compute_correlations`, leaving the #470 MC→surrogate→correlate endpoint workflow (sample `distributions` ∧ predict ∧ correlate over the SHARED set) inside the flaskapi blueprint — surfaced while re-landing the mmux_vite consumption migration: either compute stays in flaskapi (⊥ §G) or the endpoint silently degrades to table correlation (the superseded branch #537 chose the latter, ignoring `distributions`/`num_samples`/`seed` unnoticed)|`api.evaluate_correlations` + `SumoSession.correlations` (log guards via existing `_uq_engine_distributions`) + engine `correlate_manual_uq_samples` w/ REQUIRED `input_scales`/`output_scale`; flip matrix extended over ∀ 11 entry points|V45ls,V22rs,T47mt diff --git a/docs/TIER1_TIER2_UNIT_TESTS_PLAN.md b/docs/TIER1_TIER2_UNIT_TESTS_PLAN.md index c602ecb..62a89e2 100644 --- a/docs/TIER1_TIER2_UNIT_TESTS_PLAN.md +++ b/docs/TIER1_TIER2_UNIT_TESTS_PLAN.md @@ -131,6 +131,9 @@ The `unit→value` map shared by every value producer (SPEC V44ls/V45ls). **`compute_correlation_indices(..., *, input_scales, output_scale)`** - scales required (tripwire test); log-reparametrized input shifts Pearson, leaves Spearman bit-identical +**`correlate_manual_uq_samples(...)` → `api.evaluate_correlations`** +- MC→surrogate→correlate workflow (#470) with REQUIRED `input_scales`/`output_scale` (tripwire test); distributions must cover variables exactly; log variable requires uniform ∧ strictly-positive lower bound + **`create_manual_uq_samples`** — `log_scale` uniform drawn log-uniform in original units; log+normal and log+min≤0 refused **`DataPreprocessor` log transform** — `setup_log_transform`/fit/transform/inverse round-trip; delta-method `inverse_transform_output_std` (`tests/test_data_preprocessor.py::TestLogTransform`) diff --git a/docs/verification-validation.md b/docs/verification-validation.md index 720e189..3dc46a6 100644 --- a/docs/verification-validation.md +++ b/docs/verification-validation.md @@ -17,7 +17,7 @@ Live results as of the last run of the ported V&V suite: | Suite | Tests | Result | |---|---|---| -| Full standalone suite (`uv run pytest`) | 359 | **all passing** | +| Full standalone suite (`uv run pytest`) | 365 | **all passing** | | Analytical/integration tier (`-m analytical`, real Dakota subprocess, no mocking) | 30 | **all passing** | | Sobol' / Ishigami acceptance gate (`test_sobol_indices.py`) | 4 | **all passing** | @@ -151,7 +151,8 @@ breaks with `TypeError`, it can never silently default. | I5 | CV accuracy metrics: inherit log through `cross_validate` (metrics differ from linear); reject non-positive log responses | ✅ | | I6 | MOGA: log variable explored in ln-space (domain mapped, positivity-guarded); log objective exp-restored for **both** minimize and maximize (sign-after-log inverse order verified) | ✅ | | I7 | Sobol: log input shifts the variance decomposition in the expected direction (compressed variable explains less); mixed log+constant partition | ✅ | -| I8 | Flip matrix: all 10 public value-producing entry points' outputs move when a column turns log — the V45ls machine guard against any silent scale-ignore, shipped or future | ✅ | +| I8 | Flip matrix: all 11 public value-producing entry points' outputs move when a column turns log — the V45ls machine guard against any silent scale-ignore, shipped or future | ✅ | +| I9 | MC-through-surrogate correlation (`evaluate_correlations`, #470 workflow): dominant variable recovered over the shared sample set, seed-reproducible, log-scale coefficients move, log+non-uniform / log+non-positive / non-covering distributions rejected, engine producer requires its scales (`TypeError` tripwire) | ✅ | Tests: `tests/test_api_workflows.py` (`TestLogScale*`, `TestScaleGapCoverage`, `TestScaleAwareSamplers`, `TestScaleFlipMatrix`),