Skip to content

docs(extraction): report Recall@5 and NDCG@10 on Evaluate on your data - #2628

Closed
kheiss-uwzoo wants to merge 1 commit into
NVIDIA:mainfrom
kheiss-uwzoo:docs/eval-accuracy-metrics
Closed

kheiss-uwzoo wants to merge 1 commit into
NVIDIA:mainfrom
kheiss-uwzoo:docs/eval-accuracy-metrics

Conversation

@kheiss-uwzoo

Copy link
Copy Markdown
Collaborator

Summary

  • Adds an Accuracy metrics section to Evaluate on your data so readers know NRL end-to-end retrieval benchmarks report both Recall@5 and NDCG@10.
  • Keeps Recall@5 as the primary accuracy metric and explains why NDCG@10 is reported as well (public-benchmark comparison; Recall@5 varies more across queries).
  • Scoped to this PIC 2026-09-01 decision only. No CLI, eval-harness, or default-pipeline changes.

Why this PR exists

PIC 2026-09-01: Research and NRL E2E will report both metrics going forward. This is not a switch away from Recall. Nave and Even keep Recall as the accuracy north star. Janisha/Sean/Bo asked to include NDCG@10 because popular public benchmarks use it and Recall@5 is noisy.

The BEIR helpers already compute recall and NDCG at cutoffs 1, 3, 5, and 10 (DEFAULT_BEIR_KS). Tests already assert ndcg@10 and recall@5. This PR documents that reporting policy. It does not add a customer eval CLI.

Out of scope (same PIC; separate follow-ups)

  • Build endpoint deprecation coordination
  • GTC Berlin / AIQ-Retriever GTM after Nave's scope
  • LLM NIM SKU split-hosting (Nemotron Ultra / RTX Pro 6000) after Janisha's follow-up
  • Nemotron 3 Embed 1B migration / reindex guide
  • nemo-memory-datasets partner share
  • NIM patch timing for Berlin

Test plan

  • python -m mkdocs build --strict --config-file mkdocs.yml from docs/ (local untracked pages moved aside so they did not fail strict)
  • git diff --name-only upstream/main...HEAD is only docs/docs/extraction/evaluate-on-your-data.md
  • SME glance: Nave/Even okay with "primary accuracy metric" plus dual report in customer docs

Feedback requested

Please focus on the dual-report wording. Confirm we should not name experimental retriever eval / benchmark / skill-eval commands on this page.

PIC 2026-09-01: NRL E2E benchmarks report both metrics. Recall@5 remains the primary accuracy metric.

Signed-off-by: Kurt Heiss <kheiss@nvidia.com>
@kheiss-uwzoo kheiss-uwzoo self-assigned this Sep 2, 2026
@kheiss-uwzoo kheiss-uwzoo added 26.10 doc Improvements or additions to documentation labels Sep 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

26.10 doc Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant