Repository navigation
docs(extraction): report Recall@5 and NDCG@10 on Evaluate on your data - #2628
Closed
kheiss-uwzoo wants to merge 1 commit into
Closed
kheiss-uwzoo wants to merge 1 commit into
kheiss-uwzoo wants to merge 1 commit into
Conversation
PIC 2026-09-01: NRL E2E benchmarks report both metrics. Recall@5 remains the primary accuracy metric. Signed-off-by: Kurt Heiss <kheiss@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Recall@5andNDCG@10.Recall@5as the primary accuracy metric and explains whyNDCG@10is reported as well (public-benchmark comparison;Recall@5varies more across queries).Why this PR exists
PIC 2026-09-01: Research and NRL E2E will report both metrics going forward. This is not a switch away from Recall. Nave and Even keep Recall as the accuracy north star. Janisha/Sean/Bo asked to include NDCG@10 because popular public benchmarks use it and Recall@5 is noisy.
The BEIR helpers already compute recall and NDCG at cutoffs 1, 3, 5, and 10 (
DEFAULT_BEIR_KS). Tests already assertndcg@10andrecall@5. This PR documents that reporting policy. It does not add a customer eval CLI.Out of scope (same PIC; separate follow-ups)
nemo-memory-datasetspartner shareTest plan
python -m mkdocs build --strict --config-file mkdocs.ymlfromdocs/(local untracked pages moved aside so they did not fail strict)git diff --name-only upstream/main...HEADis onlydocs/docs/extraction/evaluate-on-your-data.mdFeedback requested
Please focus on the dual-report wording. Confirm we should not name experimental
retriever eval/benchmark/skill-evalcommands on this page.