Skip to content

fix(sc): set step_finished=True in async GRPO logging - #3816

Merged
yuki-97 merged 5 commits into
mainfrom
ruit/v2_2766_log_switch
Aug 27, 2026
Merged

yuki-97 merged 5 commits into
mainfrom
ruit/v2_2766_log_switch

Conversation

@RayenTian

@RayenTian RayenTian commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor

Streams #2766 to the v2 single-controller entrypoint.

#2766 added step_finished=True to the final timing/train log_metrics call in the v1 async GRPO loop (grpo.py:async_grpo_train), so W&B commits the step (commit=True) instead of buffering the step's last metrics until the next step.

The v2 entrypoint examples/run_grpo_single_controller.py drives SingleControllerActor (nemo_rl/algorithms/single_controller.py) instead of grpo.py, so #2766 didn't cover it. Its async training loop _train_pump has the same structure, where timing/train is the final per-step log. This applies the identical fix there.

Follow-up: the stall watchdog

Thanks to @yuki-97 for catching this — the port is not equivalent as-is, because v2 has a concurrent writer to the same step that v1's async_grpo_train does not.

_stall_watchdog_pump publishes rollout counters every async_rl.stall_watchdog.interval_s (default 30s) at step=self._train_steps. _train_steps is incremented at single_controller.py:1258, before the logging block this PR touches, and is not incremented again until the next step finishes. So committing at the end of step N leaves every watchdog tick for the whole duration of step N+1 naming step N — a step W&B has already closed.

W&B 0.28.1 drops a write below the current step (handler.go:972-993) and prints ...this data will be ignored on every tick. Measured on grpo-qwen2.5-math-1.5b-instruct-1n8g-megatron-single-controller-sync (20 steps, ~10.9s/step, 7 ticks): with the timing/train commit alone, 1 tick landed and 6 of 7 were dropped, one warning each. Without a second change this PR would have silently killed the rollout/* metrics.

Fix

Route the watchdog log through step_metric, which :1453 already populates:

self._logger.log_metrics(
    metrics, step=self._train_steps, step_metric="rollout/train_steps"
)

logger.py:384-386 takes that branch and calls run.log(metrics, commit=False) with no step=. W&B's monotonicity check is guarded on the step field being present, so it never runs, and the tick accumulates into the currently open step instead of being discarded. TensorBoard (logger.py:164-165) and MLflow ignore step_metric, so nothing changes there.

Note that step_finished=False would not work here: it is already the default, and the watchdog's write is already commit=False. What gets it dropped is sending a step number at all, not committing.

Trade-offs, both unchanged in kind from today's behaviour:

  • Watchdog values land in the next step's history row rather than the just-committed one. rollout/train_steps is logged as a value, so the tick's true step number stays recoverable from the data.
  • Multiple ticks within one step still overwrite each other.

Test

test_watchdog_pump.py::TestMetrics::test_ticks_never_name_the_committed_step records step_metric in the fake logger and asserts every tick passes "rollout/train_steps" and carries that key in its payload. logger.py:384 needs both — dropping the key alone would silently fall back to sending a step and restore the bug with nothing failing. Verified red without the fix, green with it.

W&B panels to expect

The rollout/* group (committed_total, inflight, idle_s, redispatch_total, …) goes from "dropped after the first tick" back to fully populated.

The visible addition is a rollout/train_steps panel — a monotonically rising step counter. It was already in the metrics dict at :1453, but promoting it to the step_metric argument is what makes it meaningful to read on its own. This is the direct analog of the existing ray/ray_step panel that RayGpuMonitorLogger produces through the same mechanism (logger.py:1019-1029).

One difference from the GPU monitor: that path pairs its step_metric with a define_metric("ray/*", step_metric="ray/ray_step") call, so ray/* plots against ray/ray_step. This call site does not, so rollout/* plots against W&B's internal step and rollout/train_steps is read as an ordinary series. Worth adding a define_metric later if the one-step offset becomes confusing in practice.

🤖 Generated with Claude Code

@copy-pr-bot

copy-pr-bot Bot commented Aug 25, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@RayenTian RayenTian added the CI:Lfast Runs a fast test suite and re-use nightly `main` container (but sync dependencies to PRs version) label Aug 25, 2026
@RayenTian

Copy link
Copy Markdown
Contributor Author

/ok to test 92d0ab0

@RayenTian
RayenTian marked this pull request as ready for review August 25, 2026 06:13
@RayenTian
RayenTian requested a review from a team as a code owner August 25, 2026 06:13
@RayenTian
RayenTian requested review from terrykong, yfw and yuki-97 August 25, 2026 06:13
@RayenTian
RayenTian force-pushed the ruit/v2_2766_log_switch branch from 92d0ab0 to f5d0611 Compare August 25, 2026 08:49
@RayenTian
RayenTian requested a review from a team as a code owner August 25, 2026 08:49
@RayenTian

Copy link
Copy Markdown
Contributor Author

/ok to test f5d0611

@RayenTian

Copy link
Copy Markdown
Contributor Author

/ok to test fb4529d

@yuki-97 yuki-97 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One note on the SC port: the concurrent stall watchdog logs to the same step, which v1's async loop does not have.

Comment thread nemo_rl/algorithms/single_controller.py
RayenTian and others added 5 commits August 26, 2026 22:45
Mirror #2766 for the v2 single-controller entrypoint: mark the final
timing/train log of each step as step_finished so W&B commits the step.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: ruit <ruit@nvidia.com>
The train_pump now passes step_finished=True to the final timing/train
log; align the test's fake logger signature with the real Logger so it
does not raise TypeError.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: ruit <ruit@nvidia.com>
Record step_finished in the fake logger and assert the final timing/train
log carries step_finished=True while the train log does not, so the fix is
guarded against regression instead of merely tolerated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: ruit <ruit@nvidia.com>
_train_pump now commits its step, and _train_steps is not incremented
again until the next step finishes, so every stall-watchdog tick in
between named a step wandb had already closed and was dropped with a
warning. Route the tick through step_metric, which takes the logger's
commit=False branch and sends no step, so it accumulates into the open
step instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: ruit <ruit@nvidia.com>
Record step_metric in the fake logger and assert every watchdog tick passes
"rollout/train_steps" and carries that key in its payload. WandbLogger takes
the no-step branch only when both hold, so dropping the key alone would
silently restore the dropped-tick bug.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: ruit <ruit@nvidia.com>
@RayenTian
RayenTian force-pushed the ruit/v2_2766_log_switch branch from fb4529d to 5657f01 Compare August 27, 2026 06:03
@RayenTian

Copy link
Copy Markdown
Contributor Author

/ok to test 5657f01

@terrykong

Copy link
Copy Markdown
Collaborator

/ok to test 5657f01

@yuki-97
yuki-97 merged commit e8f4457 into main Aug 27, 2026
108 of 121 checks passed
@yuki-97
yuki-97 deleted the ruit/v2_2766_log_switch branch August 27, 2026 15:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI:Lfast Runs a fast test suite and re-use nightly `main` container (but sync dependencies to PRs version)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants