Skip to content

Use case-weighted survival in Fleming-Harrington tests - #1698

Open
dnncha wants to merge 1 commit into
CamDavidsonPilon:masterfrom
dnncha:fix/fleming-harrington-case-weights
Open

dnncha wants to merge 1 commit into
CamDavidsonPilon:masterfrom
dnncha:fix/fleming-harrington-case-weights

Conversation

@dnncha

@dnncha dnncha commented Sep 8, 2026 •

Copy link
Copy Markdown

Fleming–Harrington weighting currently fits its pooled Kaplan–Meier curve without the supplied case weights, although the test's risk sets, event counts and covariance calculation use those weights. This makes the result depend on whether duplicate observations are expanded or represented as frequency counts.

Compute left-continuous pooled survival directly from the same weighted risk sets already used by the test: cumulatively multiply 1 - d_i / n_i, then shift by one with an initial value of 1. This also avoids the assumption that the KM timeline has an extra origin row before every observed time, resolving the zero-time length mismatch in #1681.

On load_waltons(), 163 rows can be compressed losslessly into 39 rows with integer frequency weights. For FH(p=3,q=3), the original expanded representation gives p=0.0607523, but the compressed representation gives p=4.46375e-9. After this patch, both give p=0.0607523. A full grid is recorded rather than only this crossing: (0,0), (1,0), (0,1), (1,1), (2,2), (3,3), (4,4). The (0,0) log-rank control agrees before and after; every nontrivial weighting in this grid disagrees before correction. These are the same observations, not different weighting estimands or new data.

Validation on base 7a8fc34a013ecd79fa405017b89e1697c1cc6e17: seven cases compare expanded/compressed results against a separate scalar risk-set implementation of the two-group weighted log-rank statistic; two cases test invariance to a shift from a zero time origin. Before: eight fail, one passes. The complete existing statistics test file plus these cases gives 51 passed. Black 22.8.0 formatting passes. The full package suite was not run. Released 0.30.3 also reproduces the case-weight discrepancy.

Fixes #1681. Prepared with AI assistance. This checks a bundled Drosophila dataset and does not establish an effect on a published conclusion or clinical decision. The public audit includes the reproduction scripts, separate patches, environment records and before/after execution logs in a downloadable evidence archive.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

weightings=“fleming-harrington” cant handle zero values

1 participant