Skip to content

Differentiable SCETlib prediction as a rabbit parameter model - #715

Draft
lucalavezzo wants to merge 12 commits into
WMass:mainfrom
lucalavezzo:scetlib-ad-param-model
Draft

Differentiable SCETlib prediction as a rabbit parameter model#715
lucalavezzo wants to merge 12 commits into
WMass:mainfrom
lucalavezzo:scetlib-ad-param-model

Conversation

@lucalavezzo

@lucalavezzo lucalavezzo commented Aug 19, 2026

Copy link
Copy Markdown
Collaborator

What

A new package wremnants/postprocessing/scetlib_ad/ providing a rabbit ParamModel
in which every theory parameter SCETlib exposes is a continuous fit parameter with
exact derivatives
, instead of a discrete template morph whose joint response with
the others is an outer product.

parameter group count (Z) cache/runcard flag it needs
alphaS 1 has_as for the PDF-consistent version (see below)
nonperturbative λ (Collins–Soper + TMD) 8
theory nuisance parameters 10 a non-off [TNPs] block
resummation kappa_R 1 diff_scales
matching transition points x1..x3 3 diff_scalessign-inverted, see Validation
kappa_F 1 diff_scales and has_muf (inert without it)
PDF eigenvector coefficients n_eig n_eig > 0

Which of these exist is a property of the cache, not of this code: the model
reads gradient_param_names() and registers what it finds.

The prediction comes from the SCETlib autodiff-sigmaul branch, whose
ScetlibCachedXsecTF replays a prepared cache (compressed bin rules for the
resummed piece + a frozen fixed-order grid for the nonsingular) and returns exact
first and second derivatives from clad.

How it works

1  SCETlib cached rule replay  ->  sigma(p; g)       boson level, gen grid
2  fold through the response R ->  sigma_reco(p; b)  gen -> reco
3  ratio to the reference      ->  rnorm(b, proc)    handed to rabbit

There is no surrogate: autodiff differentiates the real prediction.
ScetlibCachedXsecTF is an ordinary TF-differentiable function whose backward pass
is itself a custom_gradient contracting Hessian-vector products, so nested
GradientTapes work and TF drives every C++ call. The model just calls it inside
the graph, exactly as examples/matched_ad/tf_gradients.py does. An earlier
revision injected a local quadratic behind stop_gradient instead; that path was
removed (62fb5881, "one differentiation path, no fallback") — there is now one
path, so there is nothing to keep in sync.

That imposes one requirement which is easy to break by accident: map rabbit's fit
vector into SCETlib's layout with a constant 0/1 matrix multiply, never
tensor_scatter_nd_update.
rabbit's vector holds only the fitted parameters, POIs
first, while SCETlib's holds every registered parameter in registry order, so some
mapping is unavoidable. A scatter's backward pass contains a gather, whose gradient
TF represents as tf.IndexedSlices, and the bridge's second-order py_function
payloads call .numpy() on the incoming cotangent and fail on it — so anything past
first order breaks. The matmul is bit-identical (entries are exactly 0 and 1) and
free at these sizes (at most ~25 × 25).

Other notes:

  • One rabbit job. Second derivatives come from the bridge's own HVP
    contraction, so the composite Hessian is exact and fit + postfit covariance run
    together. Confirmed on the Asimov fits below: one job, edmval ~ 1e-27.
  • --jitCompile off is required (XLA cannot compile a PyFunc) and is
    enforced at construction with an actionable message.
  • GenFold sums cache bins onto the card's gen grid — handling a different
    nesting order, a signed-Y cache folded onto |Y|, and a cache finer than the fit
    binning — and verifies every gen bin is exactly tiled rather than assuming it.
    Note it does not normalise the Y convention out of the values it returns: a
    positive-side-only cache yields half the |Y| cross section. That cancels in the
    ratio the fit uses, and does not cancel if you compare absolutely.

Profile scales as unit nuisances

set_diff_scales(1) makes kappa_R and the transition points differentiable. But
their templates encode multiplicative (κ_R: 0.5 / 1 / 2) or asymmetric
(x2: 0.35 / 0.6 / 0.75) steps, which a symmetric Gaussian prior on the physical
value cannot express. So params.REPARAM maps them to unit nuisances —
exp(θ·ln2) for the scales, a quadratic for the transition point — and TF
differentiates the map, so the chain rule keeps gradient and Hessian exact.

The map itself is verified bit-exact (0.000e+00): θ = −1/0/+1 land on the
physical values, and the analytic Jacobian column times dκ/dθ equals TF's AD
gradient per bin, with dκ/dθ = 0.6931471806 = ln2 and 0.2 respectively. (The map
being right is not the same as the underlying scale response being right — see the
transition points below.)

Methodology note for anyone re-testing this: do not validate these
derivatives with a central difference across the anchor. The value surrogate has a
knot there (c_val forces exactness at the anchor), so a symmetric difference
averages two different slopes and the error is flat in h over four decades —
which looks like a real failure and is not. Compare against the analytic Jacobian.

Validation

Against a validated SCETlib production run and the histmaker, on the analysis
binning (Z, ptll × yll, 210 gen bins, cache_aspair):

test result
cached rules vs a live SCETlib evaluation (small cache) 6e-15 value, 1.5e-15 gradient, 3.4e-14 Hessian
σ_gen vs the production driver, full space 7.4e-06
σ_reco vs the histmaker's corrected reco, shape 0.128% yield-weighted
σ_reco vs indata.norm, absolute (no rescaling) 0.149%, total 0.998847
injection closure (small cache) exact, every other parameter unmoved, EDM 6e-22

Per-variation response against the Corr templates, all 37 labels the reference
carries (validate_variations.py), as max|dev| / mean|dev|:

block agreement verdict
20 TNP variations 2.2e-16 … 7.4e-04 / 2e-18 … 3.3e-05 excellent
8 NP λ variations 4.6e-04 … 4.9e-03 / 9e-06 … 2.1e-04 good
mufdown / mufup 2.9e-03 / 1.4e-02 usable, worth understanding
kappaFO2.-kappaf0.5 4.5e-03 good
kappaFO0.5-kappaf2. 4.0e-02 the κ_R down direction is 10× worse than up
3 transition_points* 1.1e-01 … 2.0e-01 WRONG SIGN — see below

The transition-point directions disagree in sign with their templates. All three
move the prediction the opposite way from the reference:

label                          model rng        ref rng
transition_points0.2_0.35_1.0  [1.0000,1.1593]  [0.9602,1.0000]   model UP,   ref DOWN
transition_points0.2_0.75_1.0  [0.8957,1.0000]  [1.0000,1.0207]   model DOWN, ref UP
transition_points0.3_0.6_0.9   [0.8971,1.0103]  [0.9990,1.0135]   model DOWN, ref UP

The third moves x1/x3 rather than x2, so this is not a mapping slip in the
validation table, and the reparametrisation map is separately verified bit-exact.
The central values match the templates' (x = 0.2, 0.6, 1.0). Leading suspect is a
convention difference between set_diff_scales and the production
transition_points setting. Until it is resolved, resumTransition* must not be
floated for a physics result
— and note that the σ(α_s) in fit B below did float
resumTransition2, so its number and its ρ are provisional.

Also worth flagging: the κ_R down direction agreeing only to 4% matters more than it
looks, because κ_R is the direction that dominates σ(α_s).

Asimov reco fits

-t -1, real card, login-node CPU, one job each:

floating edmval saturated 2ΔNLL / ndof wall σ(α_s)
6 (α_s + 5 λ) 9.7e-28 0.0 / 774 46 min 6.16e-04
8 (+ κ_R, x2 — x2 provisional) 9.7e-28 0.0 / 772 58 min 1.81e-03

Every parameter returns exactly at truth, and the Hessian is finite and sensible.

ρ(α_s, resumScaleMuR) = +0.927. κ_R and α_s both set the strength of the
resummed logs, so they trade off almost freely in the qT shape — floating κ_R as a
continuous nuisance costs a factor 2.9 on σ(α_s). The κ_R treatment is therefore
the dominant choice for the α_s uncertainty, more than any NP λ. (Also
ρ(λ2, λ2_ν) = −0.97, the two low-qT damping knobs being near-degenerate as
expected.) Neither σ(α_s) is quotable: these caches carry no PDF eigenvectors.

Two bugs this validation caught

The AD path initially disagreed with a validated production run by 1–3%. Every
earlier test compared the AD path against itself (exact to 1e-15), so none of them
could see it. Running the same configuration through SCETlib's own production driver
split it into two independent bugs whose product reproduced the discrepancy to 1e-5
in every bin:

  1. Ours — the cache builder configured the calculation without
    configure_ew_parameters / configure_fiducial_volumes /
    configure_calculation, so it silently used SCETlib's default EW inputs rather
    than the runcard's. Flat +1.61%. Fixed here.
  2. UpstreamDrellYan::operator()'s ad::Node_shared was hoisted out of the
    node loop, so every node reused node 0's shared sub-expressions. qT-dependent,
    ±1.4%. Reported and cherry-picked upstream as b919b61 (which also turned up
    a second occurrence).

After both: A/driver = 0.000e+00, B_cacheON/driver = 3.028e-06.

Cache contents are load-bearing — read this before quoting anything

The model registers what the cache has, and a cache built with --no-pdf is missing
directions silently unless the card still carries the templates:

  • has_as = 0α_s is a derivative at fixed PDF. The α_s pair rides on the
    existing alphas slot so one parameter moves the calculation and the PDF;
    without it, do not quote α_s. --no-pdf now says so out loud.
  • has_muf = 0resumScaleMuF has an identically zero derivative. It is
    refused at fit time rather than silently doing nothing.
  • n_eig = 0 → no PDF uncertainty from the model, so the card's pdf* templates
    must be kept.

Measured build cost at 210 bins: rules 9.0 min, FO warm 20.6 min, and 19.4 min per
PDF member
for the fixed-order variations. The α_s + μF cache (4 members) is
~1.9 h; the full 29 eigenvector pairs (62 members) is ~20 h. The member loop is
the only serial axis left — the bin loop inside it is already _parallel_run — so
that is what sharding would have to split.

Not yet done

  • The transition-point sign inversion above. Highest priority; it gates
    floating resumTransition* at all.
  • No PDF eigenvector cache yet (the ~20 h build, plus beamfunc grids for 58
    members). Until then the card must keep pdf*.
  • TNPs are not floated by defaultfit_params excludes resumTNP_*. They
    work when asked for; the default is deliberate, not an oversight.
  • Below qT ≈ 2 GeV the model's nonsingular cutoff (0.1 GeV) differs from the
    reference's (--qtCutoff 1.0). It largely cancels in the ratio the fit uses, and
    is being handled separately.
  • Real-data fits, impacts, and a comparison of the continuous κ_R treatment against
    the discrete resumFOScale templates it would replace.

Notes for review

  • The package is standalone: no code dependency on the scetlib_np package of
    SCETlib NP continuous parameter model #701. The two scetlib_np strings in response.py are datacard conventions
    (the auxiliary group name and the metadata key existing cards carry) and are
    overridable via response_group=. One optional script,
    compare_to_np_model.py, imports scetlib_np lazily — that script exists to
    cross-check the two models, so it is the one place the dependency is intended.
  • Guards for the failure modes that are otherwise silent: cache/card binning
    mismatch, cache anchor vs the card's recorded nonperturbative values, template
    nuisances that would double-count a fitted parameter (regex-based, so ^pdf\d+
    catches the eigenvectors without catching pdfAlphaS), TNPs fitted without
    priors, and any parameter whose Jacobian column is identically zero.
  • The diff is exactly the 16 files of this package: no unrelated changes ride along.
  • pylint is not installed in the local container, so the pre-commit hook's final
    step could not be run; isort, black and flake8 are clean.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c

lucalavezzo and others added 12 commits August 19, 2026 10:52
Adds `wremnants/postprocessing/scetlib_ad/`, a rabbit ParamModel in which every
theory parameter SCETlib exposes is a continuous fit parameter with exact
derivatives, rather than a discrete template morph whose joint response with the
others is an outer product: alpha_s, the 8 nonperturbative lambdas, the 10 theory
nuisance parameters, and PDF eigenvector coefficients. Which of these exist is a
property of the cache, not of the code -- the model reads
`gradient_param_names()` and registers what it finds. Only the profile-scale
parameters (kappaFO, kappaf, muf, transition points) are outside SCETlib's
autodiff and still need template nuisances.

The prediction comes from the SCETlib `autodiff-sigmaul` branch, whose
`ScetlibCachedXsecTF` replays a prepared cache -- compressed bin rules for the
resummed piece plus a frozen fixed-order grid for the nonsingular -- and returns
exact first and second derivatives from clad.

Design notes:

* Derivatives are injected, not traced. SCETlib returns value, Jacobian and
  exact Hessian from C++, so instead of differentiating through the py_function
  boundary the model hands autodiff an exact local quadratic
  `stop_gradient(val) + J.d + 0.5 d^T K d` with `d = p - stop_gradient(p)`.
  Value, first and second derivative are then exact at the evaluation point while
  everything downstream stays pure TF. This also keeps the PyFunc behind
  stop_gradient, so `GradientTape.jacobian` never re-enters C++ once per fit
  parameter. `differentiate=through` selects the alternative so the claim is
  checkable; the two agree on the gradient to 1.3e-16.
* One rabbit job. The exact second-derivative term is always included, so the
  composite Hessian is exact and the fit and postfit covariance run together.
* `--jitCompile off` is required (XLA cannot compile a PyFunc) and enforced at
  construction with an actionable message.
* `GenFold` sums cache bins onto the card's gen grid, handling a different
  nesting order, a signed-Y cache folded onto |Y|, and a cache finer than the
  fit's binning; it verifies every gen bin is exactly tiled rather than assuming.
* Guards for the failure modes that are otherwise silent: cache/card binning
  mismatch, cache anchor vs the card's recorded nonperturbative values, template
  nuisances that would double-count a fitted parameter, TNPs fitted without
  priors, and any parameter whose Jacobian column is identically zero.

Validated on a small cache: injection closure is exact in alpha_s and the
lambdas, the cached rules reproduce a live SCETlib evaluation to 1e-15 (values),
1.5e-15 (gradient) and 3.4e-14 (Hessian), and TNPs fit with priors.

Scripts under `scripts/rabbit/scetlib_ad/` build a cache for a card's binning,
check a cache standalone, build a self-contained closure card, and validate the
resummed piece against a native SCETlib production run. `conf/` carries runcards
reproducing the analysis configuration -- note that the analysis order is defined
by the `[TNPs]` block, so an analysis-faithful cache has 19 parameters, not 9.

Not yet exercised: the reco fold path, and any production-binning cache.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
The production runs carry an off-peak mass bin (10-60) alongside the 60-120 one,
so summing the reference's Q axis would have compared our mass window against a
wider one. Select the bin whose edges match the cache's window and fail loudly if
there is none.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
rabbit's fit vector is not SCETlib's: it holds only the fitted parameters, POIs
first, while SCETlib's holds every registered parameter in registry order. The
model has to map between them inside the differentiated graph, and it was doing so
with tensor_scatter_nd_update.

That op's backward pass contains a gather, whose gradient TF represents as
tf.IndexedSlices, and the SCETlib bridge's second-order py_function payloads call
.numpy() on the incoming cotangent -- so anything past first order failed with
"'IndexedSlices' object has no attribute 'numpy'". Isolated to the scatter itself:
a nested-tape HVP works on a bare Variable, and fails with a scatter in front even
when the scatter covers the whole vector.

Replacing it with a multiplication by a constant 0/1 selection matrix is
bit-identical (the entries are exactly 0 and 1) and negligible at these sizes
(at most ~25 x 25), and TF's matmul gradient rule always yields a dense cotangent.

This makes differentiate=through usable, which turns the straight-through default
from an assertion into a measured claim: the two now agree to 1.3e-16 on the
gradient, 4e-15 on HVPs and 2.4e-17 on the Hessian, and a full fit driven either
way returns the same alphaS and the same uncertainty on every parameter. Second
order had no cross-check at all before this.

Straight-through stays the default on cost rather than correctness, since rabbit
builds the postfit Hessian as t2.jacobian(grad, self.x) over the whole fit vector
and pfor cannot vectorise a PyFunc; the reasoning is now written down in the
module docstring and the README.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
Now that the fit vector is mapped with a dense matmul, letting TF differentiate
the SCETlib call is not only possible but the better default: it is the ordinary
TF idiom, it is what examples/matched_ad/tf_gradients.py does, and on a
6-parameter gen-level card it is faster -- 15.8 s of fit time against 40.9 s --
with identical postfit values and uncertainties and a slightly better EDM.

straightthrough is kept, because the two scale oppositely in the number of FIT
parameters (counting every datacard nuisance, not just ours). rabbit builds the
postfit Hessian as t2.jacobian(grad, self.x) and pfor cannot vectorise a PyFunc,
so `through` costs one C++ HVP sweep per fit parameter, while `straightthrough`
pays one value+Jacobian and one full Hessian per distinct parameter point whatever
the count. With an HVP at ~2.5x a gradient and a materialised Hessian at ~40x, the
crossover is a few tens of parameters: `through` for a gen-level fit,
`straightthrough` if a reco card with hundreds of nuisances makes the postfit
Hessian the bottleneck.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
Differentiate through the SCETlib bridge and nothing else. The straight-through
surrogate, the `differentiate` option and the mode checker are removed: with the
dense-matmul mapping in place, letting TF differentiate the real prediction works
at every order, is the ordinary TF idiom, and is faster at these parameter counts,
so carrying a second path that agrees with the first to 1e-16 was surface without
a reason.

The one requirement it imposes is now documented where it can be seen rather than
guarded by a flag: map rabbit's fit vector into SCETlib's layout with a constant
0/1 matrix multiply, never tensor_scatter_nd_update, or everything past first order
breaks while first order keeps working.

Fit unchanged: alphaS 0.1195 +/- 0.00045, same uncertainty on every parameter,
EDM 4.7e-21, and 12.8 s of fit time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
configure() called configure_calculation() and set_vary() but not
configure_ew_parameters() or configure_fiducial_volumes(). Neither is part
of configure_calculation; prod/scetlib_run/scetlib-run-qT.py -- the path
every production correction was made with -- calls both.

So the runcard's [Electroweak] block was silently ignored and SCETlib's
defaults used instead. On the analysis card, which sets mZ = 91.1535,
GammaZ = 2.4932 and custom alphaem / sin2_thw / CKM, that is a flat 1.61%
normalization error, essentially independent of qT.

Measured against the production driver on the reference runcard read
verbatim (calculation_piece = sing), Q [60,120], Y [1.0,1.5]:

  before:  operator() / driver = 1.01637 .. 1.01664 across qT 2 -> 10
  after:   operator() / driver = 1.00000000 (bit-identical, max dev 0.0)

examples/matched_ad/prepare_cache.py upstream has the same omission; the
docstring now says why this must not be "simplified" back to match it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
Everything SCETlib exposes is now a continuous fit parameter, so the
corresponding card templates can be dropped instead of double-counted.

Backend (xsec_backend.py):
- configure() gains diff_scales (registers scale_kappa_R, scale_x1..x3 and an
  inert scale_kappa_F slot via set_diff_scales(1)) and fo_resolve_muR (resolves
  the fixed-order muR dependence into the frozen grid, so the FO piece follows
  kappa_R in closed form). fo_resolve_muR must be set BEFORE the grid is built
  and costs ~3x on the warm.
- diff_scales REFUSES muf_follows_muB = yes: with muf tied to muB a live
  kappa_R would move muF while the beam convolutions stay frozen at their own.
- cache_param_names() peeks the npz 'names' array without loading the cache, so
  the calculation is configured the way the cache expects rather than the way we
  would prefer. A mismatch is otherwise a hard load failure -- the fingerprint
  hashes the names in order.

Cache builder (prepare_cache_for_card.py):
- --pdf-eig / --as-pair / --no-muf / --no-pdf / --grid-jobs, and
  build_variations() reusing the upstream helpers.
- The alphaS pair rides on the EXISTING alphas slot, so one parameter moves the
  calculation and the PDF together. Without it alphas is a derivative at FIXED
  PDF, which --no-pdf now says out loud: do not quote alphaS from such a cache.

Parameters (params.py) and model (param_model.py):
- Name maps for the scales and the PDF eigenvector coefficients, plus
  resumScale / resumTransition impact groups and pdf_group().
- REPARAM: the profile scales become UNIT nuisances, since their templates
  encode multiplicative (kappa_R: 0.5/1/2) or asymmetric (x2: 0.35/0.6/0.75)
  steps that a symmetric Gaussian on the physical value cannot express. The map
  is exp(theta*ln2) or a quadratic; TF differentiates it so the chain rule keeps
  the gradient and Hessian exact. Verified bit-exact against the analytic
  Jacobian at theta = -1, 0, +1.
- Construction asserts theta = 0 reproduces the anchor to 1e-12, so a mistyped
  map cannot silently shift the start point.
- _check_double_counting refuses to float a direction whose card templates are
  still present (regex, so ^pdf\d+ catches the eigenvectors without catching
  pdfAlphaS), and _check_no_inert_params refuses a parameter with an identically
  zero derivative -- which is what scale_kappa_F is unless the cache was built
  with the muF member pair, and which would otherwise make the covariance
  singular.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
validate_reco.py -- the model's sigma_reco against the histmaker's corrected
reco hist, folded through the response exactly as the fit does. Documents the
four traps this comparison has: R sums helicitySig while N_gen takes UL, the
ptVGen [44,100] bin is an OVERFLOW (it holds qT > 100 while sigma_SC stops
there), the two sides differ by pb-vs-fb so the plots density-normalise, and
hist.project() on a cropped hist silently re-adds the flow. Closes at 0.128%
yield-weighted. --reference card compares against indata.norm's signal column
instead (sliced start:stop, not [:nbins]), --no-match-norm drops the global
scale so the ABSOLUTE normalisation is tested (0.149%), and --y-fold defaults
off the GenFold's own y_convention -- a positive-side-only cache holds HALF the
|Y| cross section, which cancels in the ratio construction and does not cancel
absolutely.

validate_variations.py -- every variation the model produces against the
corresponding template from the Corr file, for all 38 labels. 28 of 38 agree at
1e-6..1e-8 above qT ~ 4; the residual is concentrated at low qT, where the
nonsingular cutoff differs from the reference (ours 0.1 GeV vs --qtCutoff 1.0).

compare_to_np_model.py -- AD model vs the scetlib_np model. Their centrals are
NOT required to agree (different nonsingulars, and the fit only ever uses the
ratio to each model's own central); what must agree is the response to a shared
parameter.

compare_cards.py -- row-sum audit and event-vs-matrix granularity between two
cards, needing neither cache nor SCETlib. Written for the deferred
uncorrected-histmaker route, but the row-sum audit stands alone as a check that
marginalisation, cropping and axis order are right.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
…tion status

The docstring still claimed the profile-scale parameters "are outside SCETlib's
autodiff and still need template nuisances". That has been false since
set_diff_scales(1) was wired in: kappa_R and x1..x3 are differentiable, and
kappa_F has a slot that is inert unless the cache carries the muF member pair.

Replaced with the actual status, which is NOT uniform across the scale
directions and should not be summarised as "validated":
- TNPs reproduce their templates to 1e-4..1e-16, the NP lambdas to ~1e-3;
- kappa_R reproduces kappaFO2.-kappaf0.5 to 4.5e-03 but kappaFO0.5-kappaf2.
  only to 4.0e-02, i.e. the down direction is 10x worse than the up direction,
  and that is the direction driving sigma(alpha_s) (rho = +0.93);
- all three transition_points variations move the prediction the OPPOSITE way
  from their templates (model [1.0000,1.1593] vs reference [0.9602,1.0000], and
  likewise for the other two). The one that moves x1/x3 rather than x2 inverts
  too, so it is not a mapping slip in the validation table, and the
  reparametrisation map is separately verified bit-exact. Suspect a convention
  difference between set_diff_scales and the production transition_points
  setting. Documented as: do not float resumTransition* for a physics result
  until resolved.

Also drops a comment referring to "the differentiate=through path", an option
removed in 62fb588.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
max|dev| alone cannot tell a low-qT cutoff artefact from a broken response, and
the two need completely different follow-up. Adds a "worst qT" column (always)
and --profile (the qT profile of max|dev| over |Y|, with a bar chart).

It immediately separates three different failure modes that all looked like one
number before:

- lambdas and TNPs: the residual is ONLY the low-qT feature. lambda21.0 goes
  4.9e-03 at [0,1] -> 2.2e-03 -> 5.3e-04 -> 1.8e-04 and is at 1e-05 by 5 GeV,
  1e-06 by 20. s1. has the same shape at 7.1e-04. Nothing to fix in the
  response; this is the known nonsingular cutoff mismatch.
- kappa_R: the low-qT feature PLUS a broad shoulder. kappaFO0.5-kappaf2. is
  4.0e-02 at [0,1] but stays at 5e-03..9e-03 through 3-7 GeV and ~1e-03 out to
  14 GeV -- 10-100x the lambda residual at the same qT, so something beyond the
  cutoff is wrong.
- muF: the low-qT feature PLUS a FLAT ~2e-04 pedestal at every qT out to 100,
  which is an offset in the response rather than a low-qT artefact.
- transition points: EXACTLY 0.00e+00 below 12 GeV, then monotonically rising to
  1.99e-01 at [33,44]. Not a low-qT problem at all -- it fails precisely where
  the resummed -> fixed-order matching lives, which is consistent with these
  being the matching transition points, and with the sign inversion already
  documented in param_model's docstring.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
The main CorrZ carries no alphaS and no PDF variations at all -- they live in
two sidecars produced by the same job, `*_pdfas_CorrZ` and `*_pdfvars_CorrZ`.
Validating those two directions therefore needs more than one reference file,
so --corr now takes a list and the per-file work moved into _one_file().

Their labels do not follow the main file's convention either, so they are
resolved by pattern rather than enumerated:

  * `pdfCT18ZNNLO_as_0116` / `ALPHAS_116` (HERAPDF spells it differently)
    -> alphas = 0.116. NB `_as_0118` is that file's CENTRAL, not a variation,
    which central_label() has to know or every ratio comes out against the
    wrong denominator.
  * `pdf0` is the central of the other file and `pdf(2i+1)`/`pdf(2i+2)` are
    eigenvector i up/down, i.e. c_e = +-1 by construction of
    build_pdf_variations. Reported as skipped, not silently passed, when the
    cache was built with n_eig = 0.

Measured with these: alphaS reproduces its template to 2.0e-03 (up) and
2.4e-03 (down), worst in the lowest qT bin, which is the same low-qT feature
the scale directions show and not an alphaS-specific problem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
The transition-point response we hand the fitter has the wrong sign, and the
cause is upstream rather than in this model. Making the identical physical
change through the runcard with set_diff_scales off reproduces the production
template to 2e-6; making it by moving the registered scale_x2 parameter with
set_diff_scales(1) gives 0.966985 where the template says 1.159163 at
qT [33,44] -- opposite sign, roughly -7x in slope, and linear from zero.

Mechanism: the transition points move muF by ~20% (muF has its own profile over
the same points) while the per-node beam convolutions stay frozen at the
config's muF -- conv_probe shows they shift 7-16% over that range. kappa_R
escapes it because set_muR_factor holds muF fixed by construction. It is not
fixable from Python: SCETlib's muF machinery interpolates a GLOBAL member while
the induced shift is per node, and DrellYan.hpp:586 shows a per-node dconv was
already considered and rejected. Eliminated by measurement along the way: the
REPARAM map, the label mapping in the validation table, the shared
formulas::f_run, the ported node scalars (node_scalars_probe agrees to 0.00e+00),
the separately inlined node_value, the compressed bin rules, calculation_piece,
the frozen nonsingular, and make_theory_corr.

So resumTransition2 joins 1 and 3 in DEFAULT_FROZEN. That is a KNOWN GAP, not a
fix: it drops the transition-point uncertainty from the fit. It is still
preferable to profiling a nuisance whose response points the wrong way, which
biases the POI rather than merely mis-sizing an error. Floating one anyway stays
possible -- that is how the upstream fix will be tested -- but now prints a
warning, driven by a KNOWN_BAD_RESPONSE table so the reason travels with the
name instead of living only in a comment.

Also drops the blank line isort wants gone in param_model's import block, which
is what the linting CI job has been failing on. NB run isort from a checkout
with submodules populated: in a linked worktree `rabbit/` is empty, isort then
classifies it third-party, and "fixing" the file there produces exactly the
grouping CI rejects.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnJ9YKK8c1q1sCDouM6y1c
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant