fix: label by= strata from the fit, so subset= can drop a group - #310
Conversation
kaplan() and nelson() counted the by= groups from the data and handed out labels by position, so a survfit() option passed through ... that removed a whole group (subset, start.time) stopped with "the 'by' column has 2 groups but the fit has 1 strata". CRAN 3.5.3 relabelled the remaining groups instead. Fit on a factor bound to `grp`, so survfit() names each stratum "grp=<level>", and read the labels back from those names. A fit left with one stratum is unnamed; its group is found from the rows the fit used. Raised in review on #308. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #310 +/- ##
==========================================
- Coverage 91.15% 91.14% -0.02%
==========================================
Files 60 60
Lines 5949 5964 +15
==========================================
+ Hits 5423 5436 +13
- Misses 526 528 +2
🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Negative row-index subsets are mishandled, and the documented start.time behavior lacks regression coverage.
Review effort: Balanced
Findings: 1
Open (2)
What changed in this PR
Updates stratified Kaplan–Meier and Nelson–Aalen labeling when survfit() options remove groups.
Changes:
- Derives labels from fitted strata.
- Handles single-stratum fits using retained rows.
- Adds subset-based regression coverage.
| File | Description |
|---|---|
R/utils.R |
Adds stratification and fit-row helpers. |
R/kaplan.R |
Uses fit-derived group labels. |
R/nelson.R |
Applies matching labels to weighted and unweighted fits. |
tests/testthat/test_kaplan_nelson.R |
Tests groups removed by subset. |
NEWS.md |
Documents corrected labeling. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| } else if (!is.null(subset)) { | ||
| kept <- kept & seq_along(kept) %in% subset |
There was a problem hiding this comment.
Confirmed: with subset = -arm_2 the fit kept only trt == 1 and label recovery rejected it. Fixed in 3810943: the subscript is applied to the row numbers (seq_along(kept)[subset]), so logical, positive and negative subscripts mean what they mean to survfit(). The test adds the negative case for both estimators.
| vet <- survival::veteran | ||
| for (est in list(kaplan, nelson)) { | ||
| one <- est(interval = "time", censor = "status", data = vet, by = "trt", | ||
| subset = vet$trt == 2) | ||
| expect_equal(unique(one$groups), 2) | ||
| expect_equal(one$surv, | ||
| est(interval = "time", censor = "status", | ||
| data = vet[vet$trt == 2, ])$surv) |
There was a problem hiding this comment.
Right, the test did not cover start.time, and checking it showed my PR description was also wrong. A start.time that leaves a single stratum does not lose its name: survfit() still returns grp=squamous, so it is labelled from the name and does not error. Only a subset down to one group returns no $strata. Added in 3810943: start.time = 200 drops adeno (last time 186) and labels the other three; start.time = 600 leaves only squamous, labelled correctly. Description corrected.
Review on #310. .fit_rows() tested row numbers with %in% subset, so an exclusion such as subset = -(1:2) selected no rows and a fit left with one group failed label recovery. Also pins start.time: dropping an early group, and leaving one stratum, which survfit() still names. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…int-v3 Backport #310 to maint/v3: label by= strata from the fit


Follow-up to #305, raised by Copilot in review on #308.
What was wrong
kaplan()andnelson()pass...tosurvfit(). An option there that removes a whole group (subset,start.time) leaves the fit with fewer strata than thebycolumn has values.main(since fix: restart kaplan()/nelson() interval lags in every by= stratum #305) this stops:kaplan("time", "status", data = veteran, by = "trt", subset = veteran$trt == 1)gives "the 'by' column has 2 groups but the fit has 1 strata".trt == 2rows as group1.The change
byis turned into a factor bound to the namegrp(.strata_factor()), and both fits usesurvfit(srv ~ grp, ...).survfit()then names each stratumgrp=<level>, and.label_strata()reads the labels back from those names, mapping them to the original values sogroupskeeps the column's type. The fit is the same as withstrata(data[[by]]); I checkedsurvand the strata counts onveteran.subsetleaves a single group,survfit()returns no$strata, so there is nothing to read. That case takes the group from the rows the fit used: complete cases inside thesubsetsubscript, applied to the row numbers so logical, positive and negative subscripts all work (.fit_rows()). A single stratum left bystart.timekeeps its name and is read like any other. If the kept rows do not come to one group, it is an error that says to subsetdatainstead.nelson()'s weighted fit uses the samegrp.Verification
subsetleaving one group (labelled2and equal to that arm fitted alone); a negativesubsetleaving the other;start.time = 200droppingadeno, andstart.time = 600leaving onlysquamous; andsubsetdropping one of fourcelltypegroups (labels andsurvmatch each group fitted alone).devtools::document(): no changes.lintr::lint_package(): 0.NOT_CRAN=true VDIFFR_RUN_TESTS=true devtools::test(): 0 failures, 2317 passed, 6 skipped (existing skips). No snapshot files touched.R CMD check --as-cranwith the manual, from agit archiveexport: 1 NOTE (CRAN incoming feasibility).A backport to
maint/v3is ready locally and will follow once this merges.🤖 Generated with Claude Code