Make Nemotron 3.5 Lightning the default answer and judge model - #2591
Conversation
|
d56872a to
fd675ed
Compare
Signed-off-by: jioffe502 <jioffe@nvidia.com>
| devices: | ||
| - <<: *nim-gpu | ||
| device_ids: ["${NIM_ANSWER_GPU_ID_0:-7}", "${NIM_ANSWER_GPU_ID_1:-8}"] | ||
| device_ids: ["${NIM_ANSWER_GPU_ID_0:-7}"] |
There was a problem hiding this comment.
GPU override silently stops working
The single-GPU answer service now reads NIM_ANSWER_GPU_ID_0 instead of the previously documented NIM_ANSWER_GPU_ID. Existing environments that set the old variable will be silently ignored, so the answer NIM falls back to GPU 7 and may collide with another service. Preserve the old variable as a fallback or retain the unsuffixed single-GPU name.
Prompt To Fix With AI
This is a comment left during a code review.
Path: nemo_retriever/dev/compose/service-mode.compose.yaml
Line: 232
Comment:
**GPU override silently stops working**
The single-GPU answer service now reads `NIM_ANSWER_GPU_ID_0` instead of the previously documented `NIM_ANSWER_GPU_ID`. Existing environments that set the old variable will be silently ignored, so the answer NIM falls back to GPU 7 and may collide with another service. Preserve the old variable as a fallback or retain the unsuffixed single-GPU name.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Description
Make Nemotron 3.5 Lightning the default answer generator and LLM judge across runtime configuration, evaluation tools, and examples. Answer generation remains opt-in.
2.0.9-variant, one GPU, automatic profile selection, and thenemotron_v3reasoning parser.LiteLLMClientwhile reasoning remains enabled.qwen3_coder.Production Python changes are model identifiers and the coupled service token-budget default. No new runtime control flow or abstractions.
Validation
1eb31481, which applies theblackformatting thatmainneeded in_agentic/nemo_agent/agent.pyafter the latest merge frommain.main. Kubernetes reconciliation and other GPU families were not exercised.Review follow-up
Addressed the service token-budget review by aligning all four service defaults at 4096. The follow-up passed 135 targeted tests and applicable pre-commit hooks; Compose and Helm explicit budget overrides still work. A live
/v1/answerrequest with synthetic retrieval context and hosted Lightning inference sentmax_tokens=4096with reasoning enabled, returned HTTP 200 with the correct 25% answer, and finished withstop.Review note
The GPU-variable finding compares an earlier PR revision rather than
main. The existing Compose configuration and README useNIM_ANSWER_GPU_ID_0; this PR preserves it. Rendering withNIM_ANSWER_GPU_ID_0=1correctly selects GPU1.