Backport #310 to maint/v3: label by= strata from the fit - #311
Conversation
kaplan() and nelson() counted the by= groups from the data and handed out labels by position, so a survfit() option passed through ... that removed a whole group (subset, start.time) stopped with "the 'by' column has 2 groups but the fit has 1 strata". CRAN 3.5.3 relabelled the remaining groups instead. Fit on a factor bound to `grp`, so survfit() names each stratum "grp=<level>", and read the labels back from those names. A fit left with one stratum is unnamed; its group is found from the rows the fit used. Raised in review on #308. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit c43f63f)
Review on #310. .fit_rows() tested row numbers with %in% subset, so an exclusion such as subset = -(1:2) selected no rows and a fit left with one group failed label recovery. Also pins start.time: dropping an early group, and leaving one stratum, which survfit() still names. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 3810943)
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## maint/v3 #311 +/- ##
============================================
- Coverage 89.13% 89.12% -0.01%
============================================
Files 50 50
Lines 4667 4682 +15
============================================
+ Hits 4160 4173 +13
- Misses 507 509 +2
🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
subset expressions can be evaluated twice, potentially producing labels inconsistent with the fitted rows.
Review effort: Balanced
Findings: 2
Open (2)
What changed in this PR
Backports robust survival-strata labeling to the 3.5.x maintenance line.
Changes:
- Derives group labels from fitted
survfit()strata. - Handles groups removed by
subsetorstart.time. - Adds regression coverage and release notes.
| File | Description |
|---|---|
R/kaplan.R |
Uses fitted strata for Kaplan labels. |
R/nelson.R |
Applies equivalent Nelson labeling. |
R/utils.R |
Adds strata and fit-row helpers. |
tests/testthat/test_kaplan_nelson.R |
Covers subset and start-time cases. |
NEWS.md |
Documents the corrected behavior. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| if (!is.null(by)) { | ||
| tbl <- .label_strata(tbl, srv_tab, data[[by]][!is.na(srv)]) | ||
| tbl <- .label_strata(tbl, srv_tab, strat, | ||
| .fit_rows(srv, list(...)$subset)) |
There was a problem hiding this comment.
Checked, and it does not happen. A ... argument is a promise: survfit()'s model-frame call forces it, R caches the value, and list(...)$subset afterwards reads that cached value without evaluating the expression again. I tested it with a subset that counts its own evaluations:
n_eval <- 0
kaplan("time", "status", data = veteran, by = "trt",
subset = { n_eval <- n_eval + 1; veteran$trt == n_eval })n_eval is 1 afterwards for both kaplan() and nelson() (including nelson()'s second, weighted fit), on this branch and on main, and the result is labelled group 1, matching the fit. A random single-group subset also came back with the label of the group the fit kept. No change made.
| if (!is.null(by)) { | ||
| tbl <- .label_strata(tbl, srv_tab, data[[by]][!is.na(srv)]) | ||
| tbl <- .label_strata(tbl, srv_tab, strat, | ||
| .fit_rows(srv, list(...)$subset)) |
There was a problem hiding this comment.
Checked, and it does not happen. A ... argument is a promise: survfit()'s model-frame call forces it, R caches the value, and list(...)$subset afterwards reads that cached value without evaluating the expression again. I tested it with a subset that counts its own evaluations:
n_eval <- 0
kaplan("time", "status", data = veteran, by = "trt",
subset = { n_eval <- n_eval + 1; veteran$trt == n_eval })n_eval is 1 afterwards for both kaplan() and nelson() (including nelson()'s second, weighted fit), on this branch and on main, and the result is labelled group 1, matching the fit. A random single-group subset also came back with the label of the group the fit kept. No change made.

Backports #310, which Copilot raised in review on #308. #308 shipped the strata fix from #305 to
maint/v3without this follow-up, so 3.5.4 as it stands errors when asubsetpassed through...drops a wholebygroup.What it carries
Both #310 commits, cherry-picked with
-x:grp, sosurvfit()names each stratumgrp=<level>, and read the labels back from those names. Asubsetorstart.timethat drops a group now labels the groups that are left; CRAN 3.5.3 relabelled them as the first groups indata.subsetleaves a single group (the one casesurvfit()does not name), the group comes from the rows the fit used, with the subscript applied to the row numbers so negative indices work.Differences from
mainNEWS.md: the note goes into the existing 3.5.4 strata bullet, worded against 3.5.3 rather than the v4 development line. It also rewraps an overlong line Backport #303, #304, #299 fixes to maint/v3 (3.5.4) #308 introduced.Verification
lintr::lint_package(): 0.NOT_CRAN=true VDIFFR_RUN_TESTS=true devtools::test(): 0 failures, 1578 passed, 5 skipped (existing skips). No snapshot files touched.R CMD check --as-cranwith the manual, from agit archiveexport: 1 NOTE (CRAN incoming feasibility).🤖 Generated with Claude Code