Skip to content

feat: add metric observations for vLLM and SGLang - #1068

Draft
podkidyshev wants to merge 3 commits into
mainfrom
ipod/llm-metrics
Draft

podkidyshev wants to merge 3 commits into
mainfrom
ipod/llm-metrics

Conversation

@podkidyshev

@podkidyshev podkidyshev commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add metric_observations to vLLM and SGLang for request throughput (requests/s), output-token throughput (tokens/s), TTFT/TPOT (ms), and optional semantic accuracy.
  • Dimensions distinguish produced values within a concrete test-case iteration or DSE step: TTFT/TPOT use statistic (mean, median, p99); throughput and accuracy have empty dimensions. Workload, model, and configured concurrency remain in the test/run context.
  • Support SOL defaults for scalar metrics and statistic-specific latency targets. Reuse the existing parsers, omit unavailable/non-finite measurements, and preserve existing scalar metric behavior.
  • Document metrics, dimensions, and a SOL example through a concise shared include in the vLLM/SGLang Reporting sections; test unique metric/dimension identities and independence from run-configuration metadata.

Test Plan

Environment: macOS, Python 3.14.3, locked development dependencies.

uv run --locked --extra dev pytest tests/workloads/vllm tests/workloads/sglang tests/workloads/common/test_llm_serving.py tests/test_metrics.py
170 passed in 0.42s

Ran uv run --locked --extra dev pre-commit run --files for changed files, including pyright, Ruff, vulture, and import-linter. All applicable checks passed after reviewing formatter edits. git diff --check passed.

Replayed five locally fetched benchmark executions without changing their result files. TestRun objects were reconstructed from saved metadata; historical SGLang semantic-evaluation settings were enabled in memory using the current schema.

Check Actual Expected
Two vLLM iterations 8 observations each Two throughput values and six latency statistics
Three SGLang executions (aggregated on one/two nodes; disaggregated on two nodes) 9 observations each Same measurements plus semantic accuracy
Original measurements and dimensions All 43 values matched source JSON/evaluation logs; only latency observations carried dimensions Exact values, scalar dimensions empty, latency dimensions limited to statistic
Observation identity Every metric/dimension combination unique within each execution No ambiguous values
SOL assessments Defaults and statistic-specific overrides matched expected targets and formulas Measured/target for throughput and accuracy; target/measured for latency

Built the documentation using uv run --locked --extra dev --extra docs sphinx-build -W -b html doc results/llm-metrics-docs-move with no warnings, and inspected both workload pages: five metric rows, the TOML example, and the reporting link render correctly. doc/reporting.rst has no changes relative to the PR base.

Additional Notes

This PR adds observation support. Built-in vLLM/SGLang comparison reporters retain their existing scalar path; NCCL/NIXL observations are outside this change's scope.

The generic metric comparison reporter's legacy Bokeh renderer has a known failure on categorical dimensions such as mean/median/p99 (TypeError while calculating numeric ticks). Reporter integration and that rendering fix are separate work.

Local replay scripts, fetched results, and generated artifacts are not committed. SOL values used for validation are illustrative, not hardware performance claims. No remote benchmark jobs were run.

Signed-off-by: Ivan Podkidyshev <ipodkidyshev@nvidia.com>
@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

Signed-off-by: Ivan Podkidyshev <ipodkidyshev@nvidia.com>
Signed-off-by: Ivan Podkidyshev <ipodkidyshev@nvidia.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant