Skip to content

refactor(sdk): type LLM streaming, message, and tokenizer boundaries - #5504

Merged
neubig merged 8 commits into
OpenHands:mainfrom
vedjoshi1:refactor/4976-typed-llm-boundaries
Oct 6, 2026
Merged

neubig merged 8 commits into
OpenHands:mainfrom
vedjoshi1:refactor/4976-typed-llm-boundaries

Conversation

@vedjoshi1

@vedjoshi1 vedjoshi1 commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

HUMAN:

trying to contribute towards issue #4976, this change makes the data/format of data that the llm actually expects a little clearer. hopefully this will make maintenance and debugging easier.


AGENT:

Why

LLM integration currently probes provider objects dynamically throughout core
streaming, message conversion, and tokenizer code. This change moves those
assumptions into explicit typed contracts and focused normalization helpers.
It also retains a completion observed in a stream when the wrapper's final
completed_response is stale None, following alanhuangyoo's work in #4772.

Summary

  • Introduce private stream/event and tokenizer capability helpers, preserving
    sync/async iterator support, optional Transformers loading, and fallback paths.
  • Normalize consumed provider message fields before conversion; preserve generic
    LiteLLM output-item events, opaque metadata, tool-call replay IDs, and ignored
    text parts. Access the declared authentication field directly.
  • Add focused regressions and remove 14 obsolete dynamic-access baseline entries.

Issue Number

Closes #4976.

How to Test

make build
uv run pytest tests/sdk/llm/ -q -o addopts='--tb=short'
uv run pytest tests/sdk -q -o addopts='--tb=short'
uv run pre-commit run --all-files
uv run python scripts/check_forbidden_dynamic_attributes.py --baseline-ref upstream/main
make build-server

Recorded local validation on 2026-10-05 after refreshing onto upstream
cb7eabaf7, on macOS arm64 / Python 3.13.3 / LiteLLM 1.93.0 / SDK 1.53.0:

  • Full SDK suite plus local development matrix: 8,175 passed, 1 failed,
    9 skipped, 12 xfailed
    in 323.84s. The sole failure was
    TestACPSessionIdPersistence.test_mask_callback_does_not_retain_agent,
    which timed out waiting for garbage collection. It passed in isolation,
    and all 39 tests in its surrounding group passed on rerun. The test
    and ACP implementation are unchanged from upstream. This suggests a
    cleanup/timing flake; the broad run is not represented as all-green.
    No LLM or matrix case failed.
  • Matrix rerun: 1,285 passed, including 240 generic-event reconstruction
    cases. The development harness is local; permanent regressions are under
    tests/sdk/llm/.
  • 12/12 real local HTTP/SSE checks passed.
  • Token counting passed with and without Transformers 5.17.0 (local count: 2).
  • The rebuilt packaged server returned HTTP 200 from /health; its owned
    process was stopped.
  • All-file pre-commit, staged hooks, dynamic-access gate, and whitespace checks
    passed.
  • Package-version guard against refreshed upstream/main passed: no package
    version changes detected
    . The branch inherits upstream 1.53.0.
  • No live provider calls were run because provider credentials were not configured.

See validation summary. Detailed development harnesses and
reports are preserved locally rather than included in the final PR diff.

Regression reproduction: construct GenericEvent or
BaseLiteLLMOpenAIResponseObject with type='response.output_item.done', followed
by a completion with empty output. Before the boundary correction, all six
generic-event tests failed; now reconstruction works across sync, async, and
async callers receiving sync streams, preserving output-item identity.

Video/Screenshots

Not applicable: internal SDK refactor; transport and packaged-runtime evidence
is recorded above.

Design Doc

No separate design artifact. Provider variability is isolated in three private
LLM helpers; core code consumes declared fields. Public exports, persisted event
fields, REST contracts, dependencies, and package versions are unchanged.

Type

  • Bug fix
  • Feature
  • Refactor
  • Breaking change
  • Docs / chore

Notes

  • Parent tracking issue: Remove dynamic attribute access from LLM and telemetry code #4904. Completion-selection semantics follow
    alanhuangyoo's earlier PR fix(sdk): keep the completion a Responses stream yielded #4772, credited here and in the implementation.

  • Refreshed onto upstream cb7eabaf7 without additional conflicts. Conflicts in LLM/test imports and async
    Responses iteration are resolved. Upstream hard/idle timeouts and transient
    error classification are preserved. Idle timing wraps the raw provider stream
    before event filtering; a new regression verifies ignored events keep the
    stream alive without producing callbacks. Ready for maintainer review; fork CI runs require maintainer approval.

  • Known-kind projections reject some malformed provider values previously
    tolerated (e.g. reasoning parts that are primitives or falsy non-list
    collections). No real provider emitting those shapes was demonstrated; this
    compatibility limitation remains explicit for review.

  • Only .pr/validation.md remains in the final diff. Earlier commits still
    contain development artifacts; rebasing changed commit IDs but did not strip
    those artifacts from historical commits.

  • Companion documentation: docs(sdk): describe typed provider boundaries and Responses streaming docs#891 (ready for review).

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

🧹 PR Artifact Cleanup Queued

The .pr/ directory reached main after merge. A cleanup PR has been opened or updated: #5482

vedjoshi1 and others added 8 commits October 4, 2026 20:26
Preserves the completion precedence and regression identified by alanhuangyoo in OpenHands#4772.

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: openhands <openhands@all-hands.dev>
Preserve absent LiteLLM fields, response replay metadata, and subscription config behavior. Cover stream output reconstruction without callbacks.

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: openhands <openhands@all-hands.dev>
Normalize generic LiteLLM output-item events at the stream boundary and preserve opaque metadata, ignored text parts, and Responses item IDs. Add focused regressions and retain validation evidence in temporary .pr artifacts.

Co-authored-by: openhands <openhands@all-hands.dev>
Keep only a concise validation summary in .pr; preserve detailed development artifacts in an ignored local archive.

Co-authored-by: openhands <openhands@all-hands.dev>
Verify that provider events filtered by the typed adapter still reset the upstream stream idle timer and do not emit callbacks.

Co-authored-by: openhands <openhands@all-hands.dev>
Record the green full SDK suite, preserved upstream stream timeouts, transport and tokenizer checks, and rebuilt packaged-server validation.

Co-authored-by: openhands <openhands@all-hands.dev>
@vedjoshi1
vedjoshi1 force-pushed the refactor/4976-typed-llm-boundaries branch from 5b07d0e to 120e17a Compare October 5, 2026 03:36
@vedjoshi1

Copy link
Copy Markdown
Contributor Author

This is ready for review from my side. Could a maintainer approve the pending CI workflows? Local validation passed: 6,905 SDK tests, the validation matrix, pre-commit checks, and transport/packaged-server checks.

@vedjoshi1
vedjoshi1 marked this pull request as ready for review October 5, 2026 03:50
@neubig neubig self-assigned this Oct 5, 2026
@neubig neubig added the integration-test Runs the integration tests and comments the results label Oct 6, 2026
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Hi! I started running the integration tests on your PR. You will receive a comment with the results shortly.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Test Suites Triggered

Results will be posted here when complete.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

🧪 Integration Tests Results

Overall Success Rate: 98.0%
Total Cost: $1.99
Models Tested: 5
Timestamp: 2026-10-06 00:19:19 UTC

📁 Detailed Logs & Artifacts

Click the links below to access detailed agent/LLM logs showing the complete reasoning process for each model. On the GitHub Actions page, scroll down to the 'Artifacts' section to download the logs.

📊 Summary

Model Overall Tests Passed Skipped Total Cost Tokens
litellm_proxy_openai_gpt_5.5 100.0% 10/10 0 10 $0.89 315,072
litellm_proxy_gemini_3.1_pro_preview 100.0% 10/10 0 10 $0.49 340,445
litellm_proxy_anthropic_claude_sonnet_4_6 90.0% 9/10 0 10 $0.59 400,991
litellm_proxy_deepseek_deepseek_v4_flash 100.0% 10/10 0 10 $0.02 418,192
litellm_proxy_minimax_MiniMax_M2.7 100.0% 9/9 1 10 $0.00 369,215

📋 Detailed Results

litellm_proxy_openai_gpt_5.5

  • Success Rate: 100.0% (10/10)
  • Total Cost: $0.89
  • Token Usage: prompt: 310,133, completion: 4,939, cache_read: 179,712, reasoning: 1,628
  • Run Suffix: litellm_proxy_openai_gpt_5.5_cb7eaba_gpt_5_5_run_N10_20261006_001721

litellm_proxy_gemini_3.1_pro_preview

  • Success Rate: 100.0% (10/10)
  • Total Cost: $0.49
  • Token Usage: prompt: 335,560, completion: 4,885, cache_read: 133,992, reasoning: 2,816
  • Run Suffix: litellm_proxy_gemini_3.1_pro_preview_cb7eaba_gemini_3_1_pro_run_N10_20261006_001721

litellm_proxy_anthropic_claude_sonnet_4_6

  • Success Rate: 90.0% (9/10)
  • Total Cost: $0.59
  • Token Usage: prompt: 395,390, completion: 5,601, cache_read: 282,356, cache_write: 112,975, reasoning: 550
  • Run Suffix: litellm_proxy_anthropic_claude_sonnet_4_6_cb7eaba_claude_sonnet_4_6_run_N10_20261006_001714

Failed Tests:

  • t02_add_bash_hello: Shell script is not executable (Cost: $0.06)

litellm_proxy_deepseek_deepseek_v4_flash

  • Success Rate: 100.0% (10/10)
  • Total Cost: $0.02
  • Token Usage: prompt: 412,619, completion: 5,573, cache_read: 313,856, reasoning: 1,777
  • Run Suffix: litellm_proxy_deepseek_deepseek_v4_flash_cb7eaba_deepseek_v4_flash_run_N10_20261006_001707

litellm_proxy_minimax_MiniMax_M2.7

  • Success Rate: 100.0% (9/9)
  • Total Cost: $0.00
  • Token Usage: prompt: 364,298, completion: 4,917, cache_read: 279,825
  • Run Suffix: litellm_proxy_minimax_MiniMax_M2.7_cb7eaba_minimax_m2_7_run_N10_20261006_001740
  • Skipped Tests: 1

Skipped Tests:

  • t08_image_file_viewing: This test requires a vision-capable LLM model. Please use a model that supports image input.

@neubig neubig removed the integration-test Runs the integration tests and comments the results label Oct 6, 2026

@neubig neubig left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving on behalf of @neubig.

Verified independently:

  • tests/sdk/llm/ -> 1218 passed locally on this branch.
  • Real agent sessions through Conversation + TerminalTool on this branch, instrumenting litellm_responses/litellm_completion to confirm the endpoint hit: all three sessions routed through the Responses path (including stream=True and a reasoning-heavy task), all reached FINISHED with correct tool-use outcomes.
  • Protocol/fallback semantics confirmed: OutputItemEvent correctly rejects pydantic-extra objects, _GenericOutputItemEvent catches them, and completed_response precedence behaves as documented.
  • All required status checks green (pre-commit, sdk-tests, tools-tests, cross-tests, agent-server-tests, build-binary-and-test (ubuntu-latest), Check OpenAPI Schema).

Rejection of malformed provider output is accepted as intended behavior.

Note for follow-up (non-blocking): the branch is behind main (main is at 1.53.0, branch at 1.51.0), which is why the non-required 'Check package versions' job reports version changes. The PR diff does not touch version files, so a squash merge will not revert main's versions.

This review was produced by an AI agent (OpenHands) on behalf of @neubig.

@neubig
neubig merged commit 69eb315 into OpenHands:main Oct 6, 2026
51 of 59 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Type LLM streaming, message, and tokenizer boundaries

2 participants