Skip to content

Make Nemotron 3.5 Lightning the default answer and judge model - #2591

Merged
jioffe502 merged 9 commits into
NVIDIA:mainfrom
jioffe502:jioffe502/replace-retired-answer-defaults
Sep 29, 2026
Merged

jioffe502 merged 9 commits into
NVIDIA:mainfrom
jioffe502:jioffe502/replace-retired-answer-defaults

Conversation

@jioffe502

@jioffe502 jioffe502 commented Aug 26, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Make Nemotron 3.5 Lightning the default answer generator and LLM judge across runtime configuration, evaluation tools, and examples. Answer generation remains opt-in.

  • Update Helm and Compose to NIM image 2.0.9-variant, one GPU, automatic profile selection, and the nemotron_v3 reasoning parser.
  • Raise the service answer budget from 512 to 4096 across Python, bundled YAML, Helm, and Compose, matching LiteLLMClient while reasoning remains enabled.
  • Keep answer-model tool calling opt-in; update the existing judge tool parser to qwen3_coder.
  • Update matching tests and documentation. Retain Nano as an optional example and preserve local Super-49B compatibility and historical benchmark attribution.

Production Python changes are model identifiers and the coupled service token-budget default. No new runtime control flow or abstractions.

Validation

  • 261 tests and 38 subtests passed; the example follow-up passed 17 targeted tests. Pre-commit hooks, Compose rendering, and documented Helm tool-call overrides passed.
  • All CI checks pass on 1eb31481, which applies the black formatting that main needed in _agentic/nemo_agent/agent.py after the latest merge from main.
  • Self-hosted smoke passed on one RTX PRO 6000 Blackwell (96 GB), including NIM startup, generation, judging, reasoning controls, and opt-in tool calling/agentic retrieval.
  • Hosted validation passed with a working NVIDIA API key: generation with reasoning enabled/disabled; correct-answer score 1.0 and incorrect-answer score 0.0; both judge prompts individually; the skill-eval config's 4096-token judge; a three-case retrieval/generation/judge pipeline; tool calling; and agentic selection of the correct document.
  • The hosted pipeline completed all three cases with scores 1.0, 0.5, and 1.0 and no reported generation/judge errors. These synthetic checks establish functionality, not benchmark quality. Long, variable request durations and an observed NFS package-import stall make this run unsuitable for performance conclusions.
  • Strict MkDocs has the same pre-existing embedded CLI link failure as main. Kubernetes reconciliation and other GPU families were not exercised.

Review follow-up

Addressed the service token-budget review by aligning all four service defaults at 4096. The follow-up passed 135 targeted tests and applicable pre-commit hooks; Compose and Helm explicit budget overrides still work. A live /v1/answer request with synthetic retrieval context and hosted Lightning inference sent max_tokens=4096 with reasoning enabled, returned HTTP 200 with the correct 25% answer, and finished with stop.

Review note

The GPU-variable finding compares an earlier PR revision rather than main. The existing Compose configuration and README use NIM_ANSWER_GPU_ID_0; this PR preserves it. Rendering with NIM_ANSWER_GPU_ID_0=1 correctly selects GPU 1.

@jioffe502
jioffe502 requested review from a team as code owners August 26, 2026 17:13
@jioffe502
jioffe502 requested a review from ChrisJar August 26, 2026 17:13
@greptile-apps

greptile-apps Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Changes default LLM model for answer generation and evaluation.

The PR appears safe to merge; no new actionable issue or outstanding previous finding was identified.

Findings

  1. P1 GPU override silently stops working ▶
Fix with agent prompt
### Issue 1
nemo_retriever/dev/compose/service-mode.compose.yaml:undefined-232
The single-GPU answer service now reads `NIM_ANSWER_GPU_ID_0` instead of the previously documented `NIM_ANSWER_GPU_ID`. Existing environments that set the old variable will be silently ignored, so the answer NIM falls back to GPU 7 and may collide with another service. Preserve the old variable as a fallback or retain the unsuffixed single-GPU name.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR makes Nemotron 3.5 Lightning the default answer and judge model while keeping answer generation opt-in.

  • Updates the Helm and Compose NIM defaults to a one-GPU configuration with reasoning parsing.
  • Aligns the service answer-token budget at 4096 and updates evaluation examples, tests, and documentation.

Reviews (10) · Last reviewed commit: "Apply black formatting to agentic citati..."

@jioffe502
jioffe502 force-pushed the jioffe502/replace-retired-answer-defaults branch from d56872a to fd675ed Compare August 26, 2026 17:44
@jioffe502 jioffe502 changed the title Replace retired answer model defaults with Nemotron 3 Nano Replace retired answer model defaults with Nemotron 3.5 Lightning Aug 26, 2026
@jioffe502
jioffe502 marked this pull request as draft August 26, 2026 17:48
Signed-off-by: jioffe502 <jioffe@nvidia.com>
@jioffe502 jioffe502 changed the title Replace retired answer model defaults with Nemotron 3.5 Lightning Make Nemotron 3.5 Lightning the default answer and judge model Sep 18, 2026
@jioffe502
jioffe502 marked this pull request as ready for review September 21, 2026 17:14
Comment thread nemo_retriever/src/nemo_retriever/tools/skill_eval/configs/skill_eval.yaml Outdated
Comment thread nemo_retriever/examples/eval_sweep.yaml Outdated
devices:
- <<: *nim-gpu
device_ids: ["${NIM_ANSWER_GPU_ID_0:-7}", "${NIM_ANSWER_GPU_ID_1:-8}"]
device_ids: ["${NIM_ANSWER_GPU_ID_0:-7}"]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 GPU override silently stops working

The single-GPU answer service now reads NIM_ANSWER_GPU_ID_0 instead of the previously documented NIM_ANSWER_GPU_ID. Existing environments that set the old variable will be silently ignored, so the answer NIM falls back to GPU 7 and may collide with another service. Preserve the old variable as a fallback or retain the unsuffixed single-GPU name.

Prompt To Fix With AI
This is a comment left during a code review.
Path: nemo_retriever/dev/compose/service-mode.compose.yaml
Line: 232

Comment:
**GPU override silently stops working**

The single-GPU answer service now reads `NIM_ANSWER_GPU_ID_0` instead of the previously documented `NIM_ANSWER_GPU_ID`. Existing environments that set the old variable will be silently ignored, so the answer NIM falls back to GPU 7 and may collide with another service. Preserve the old variable as a fallback or retain the unsuffixed single-GPU name.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Comment thread nemo_retriever/src/nemo_retriever/service/config.py
@jioffe502
jioffe502 requested a review from jperez999 September 22, 2026 15:17
jioffe502 and others added 2 commits September 28, 2026 18:13
Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@jioffe502
jioffe502 merged commit de42d5d into NVIDIA:main Sep 29, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants