Skip to content

fix(evaluation): score CAR scenarios on the mode key only - #627

Merged
ShuxinLin merged 1 commit into
aafeedback_changesfrom
fix/car-mode-key-only
Oct 8, 2026
Merged

ShuxinLin merged 1 commit into
aafeedback_changesfrom
fix/car-mode-key-only

Conversation

@ShuxinLin

Copy link
Copy Markdown
Collaborator

Description

Score clarification-abstain-response (CAR) scenarios only on the top-level mode key of the answer (response, clarification or abstain). The text under the key no longer affects pass or score.

Fix Details

  • src/evaluation/scorers/static_json.py, _evaluate_mode_json: an answer passes when it is a JSON object with exactly one top-level key and that key matches the gold mode (case-insensitive). score, f1 and strict_exact_match_accuracy are 1.0 on a pass and 0.0 otherwise.
  • Required-term coverage (mode_required_terms, mode_matched_terms, mode_term_coverage) is still reported as a diagnostic, but no longer feeds pass, score or the precision/recall/F1 key counts. The per-key details now hold only the mode-key comparison.
  • Harbor runs score with the same scorer through tests/test.sh, so Harbor rewards change too.
  • Tests: six CAR tests were already failing on aafeedback_changes. They targeted the groundtruth_eval.json term scoring and car_score that were removed in 8f1e538. I replaced them with key-only tests: a correct key with no matching terms passes, a wrong key with matching terms fails, the key is case-insensitive, an answer that isn't an object fails, and the StaticJsonScorer wrapper gives 1.0/0.0.
  • docs/static-json-evaluation.md: the CAR pass criteria and fields now match the code, and the docs say static_json does not read groundtruth_eval.json.

Impact on Benchmarking

  • No change to baselines: This fix only improves stability/performance.
  • Baseline change: This fix corrects a scoring error. (Please provide "Before vs. After" results).

CAR pass rates will go up for answers that pick the right mode but word it differently from the gold answer. Trade-off: an agent that always picks the same mode gets credit on every scenario expecting that mode. This PR has no before/after agent-run comparison.

Conflicts with origin/eval-scoring-relaxation (d1a668c), which rewrites the same function so that every required term is needed to pass.

Related Issues

  • Fixes: #

Verification Steps

  1. uv run pytest src/evaluation -q -k "not integration": 110 passed (6 failed before this change).
  2. uv run pytest src/ -q -k "not integration": 752 passed, 2 failed in src/observability/tests/test_file_exporter.py. Those 2 also fail on the base branch without this change.
  3. Manual: each gold answer of the 22 CAR tasks bundled under benchmarks/harbor/datasets/ passes when scored against itself.

Checklist

  • I have added tests that prove my fix is effective.
  • My code follows the project's Ruff formatting and linting rules. (No new findings; the 3 ruff reports in these files are already on the base branch.)
  • I have signed off my commits (DCO).

A CAR answer now passes when it is a single-key object whose key matches
the gold mode (response / clarification / abstain). The value under the
key no longer affects pass or score; required-term coverage stays in the
mode_* fields as a diagnostic.

Replaces the CAR tests that targeted the removed groundtruth_eval.json
term scoring and updates docs/static-json-evaluation.md to match.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
@ShuxinLin
ShuxinLin requested a review from DhavalRepo18 October 8, 2026 13:25
@ShuxinLin
ShuxinLin merged commit 9584e28 into aafeedback_changes Oct 8, 2026
5 checks passed
@ShuxinLin
ShuxinLin deleted the fix/car-mode-key-only branch October 8, 2026 14:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants