Skip to content

fix(condenser): split summarization prompt into system + user messages - #5143

Merged
juanmichelini merged 5 commits into
mainfrom
condenser/system-user-prompt-split
Sep 26, 2026
Merged

juanmichelini merged 5 commits into
mainfrom
condenser/system-user-prompt-split

Conversation

@juanmichelini

@juanmichelini juanmichelini commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

HUMAN:
Splitting the condenser's summarization prompt into a canonical system + user pair so the Responses API path stops dropping the event payload into instructions with empty input. I want the steering text and the <EVENT> data to land in the correct slots on every transport.

AGENT:

Why

The LLMSummarizingCondenser previously sent a single user message that bundled the "You are maintaining a context-aware state summary..." steering instructions with the <EVENT> list (prompts/summarizing_prompt.j2). This conflates "how to summarize" (steering) with "what to summarize" (data). On the OpenAI Responses API path, a lone system-style payload serialized to the instructions field with empty input (Message.to_responses_value returns a string for system; message_to_responses_dict returns [] for system), so the event data never landed as a real input item.

The system + user split is the canonical shape for both Chat Completions and Responses: the steering header goes to instructions/system, the event data goes to input/user. It also plays well with the subscription/Codex transport (transform_for_subscription), which prepends system chunks onto the first user message.

This change is safe because the condenser makes an independent, single-turn, side-channel call to its own LLM (self.llm.generate(messages=..., store=False)) — it builds its message list from scratch (no agent system prompt, no tools, no prior turns). It only affects that isolated call and does not touch the agent's main conversation history. The summary still re-enters the agent's view as a CondensationSummaryEvent (user message) at the forget boundary, exactly as before. Prompt caching is also fine: _apply_prompt_caching now marks the system block plus the trailing user block (two breakpoints), within Anthropic's limit.

Summary

  • New prompts/summarizing_system.j2 — the steering instructions, sent as a system message.
  • New prompts/summarizing_events.j2 — just the <EVENT> loop + "Now summarize the events using the rules above.", sent as a user message.
  • Removed prompts/summarizing_prompt.j2 (the old single combined template).
  • Added a shared LLMSummarizingCondenser._build_summary_messages() helper; both the sync _generate_condensation and async _agenerate_condensation now use it (dedupes the two paths).
  • Updated test_get_condensation_with_previous_summary to assert messages[0] is system and messages[1] (the events payload) is user, since the previous summary now lives in the user message.

Issue Number

Closes #5142.

How to Test

Run the condenser unit and broader context tests on this head:

pytest tests/sdk/context/condenser/test_llm_summarizing_condenser.py
pytest tests/sdk/context/ tests/sdk/conversation/test_condense.py tests/sdk/agent/test_agent_context_window_condensation.py
ruff format --check && ruff check

To verify the message roles end-to-end across transports, build the summary messages from a condenser instance and assert the serialization:

  • Chat Completions: roles are ['system', 'user'].
  • Responses API: instructions is the steering text and input is one user item carrying the <EVENT> payload.
  • Subscription/Codex: transform_for_subscription prepends the system chunk onto the first user message as intended.

Verification

  • pytest tests/sdk/context/condenser/test_llm_summarizing_condenser.py -> 44 passed
  • Broader run (tests/sdk/context/, tests/sdk/conversation/test_condense.py, tests/sdk/agent/test_agent_context_window_condensation.py) -> 470 passed
  • ruff format + ruff check clean
  • End-to-end render check confirms msg[0].role == "system" (instructions) and msg[1].role == "user" (the <EVENT> list), and the condensation produces the summary correctly

Type

  • Bug fix
  • Feature
  • Refactor
  • Breaking change
  • Docs / chore

Notes

The summarizing_prompt.j2 removal was verified with a repo-wide grep (including pyproject.toml/MANIFEST.in which glob *.j2) — no other references remain.


This PR was created by an AI agent (OpenHands) on behalf of @juanmichelini.


🐳 Agent Server images for this PR — GHCR package, pull/run commands, and all pushed tags (click to expand)

• GHCR package: https://github.com/OpenHands/agent-sdk/pkgs/container/agent-server

Variants & Base Images

Variant Architectures Base Image Docs / Tags
java amd64, arm64 eclipse-temurin:17-jdk Link
python-slim amd64, arm64 python-node-runtime Link
python-minimal amd64, arm64 python-node-runtime Link
python amd64, arm64 python-node-runtime Link
golang amd64, arm64 golang:1.21-bookworm Link

Pull (multi-arch manifest)

# Each variant is a multi-arch manifest supporting both amd64 and arm64
docker pull ghcr.io/openhands/agent-server:bb07a23-python

Run

docker run -it --rm \
  -p 8000:8000 \
  --name agent-server-bb07a23-python \
  ghcr.io/openhands/agent-server:bb07a23-python

All tags pushed for this build

ghcr.io/openhands/agent-server:bb07a23-golang-amd64
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-golang-amd64
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-golang-amd64
ghcr.io/openhands/agent-server:bb07a23-golang_tag_1.21-bookworm-amd64
ghcr.io/openhands/agent-server:bb07a23-golang-arm64
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-golang-arm64
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-golang-arm64
ghcr.io/openhands/agent-server:bb07a23-golang_tag_1.21-bookworm-arm64
ghcr.io/openhands/agent-server:bb07a23-java-amd64
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-java-amd64
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-java-amd64
ghcr.io/openhands/agent-server:bb07a23-eclipse-temurin_tag_17-jdk-amd64
ghcr.io/openhands/agent-server:bb07a23-java-arm64
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-java-arm64
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-java-arm64
ghcr.io/openhands/agent-server:bb07a23-eclipse-temurin_tag_17-jdk-arm64
ghcr.io/openhands/agent-server:bb07a23-python-amd64
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-python-amd64
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-python-amd64
ghcr.io/openhands/agent-server:bb07a23-python-node-runtime-amd64
ghcr.io/openhands/agent-server:bb07a23-python-arm64
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-python-arm64
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-python-arm64
ghcr.io/openhands/agent-server:bb07a23-python-node-runtime-arm64
ghcr.io/openhands/agent-server:bb07a23-python-minimal-amd64
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-python-minimal-amd64
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-python-minimal-amd64
ghcr.io/openhands/agent-server:bb07a23-python-node-runtime-minimal-amd64
ghcr.io/openhands/agent-server:bb07a23-python-minimal-arm64
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-python-minimal-arm64
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-python-minimal-arm64
ghcr.io/openhands/agent-server:bb07a23-python-node-runtime-minimal-arm64
ghcr.io/openhands/agent-server:bb07a23-python-slim-amd64
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-python-slim-amd64
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-python-slim-amd64
ghcr.io/openhands/agent-server:bb07a23-python-node-runtime-slim-amd64
ghcr.io/openhands/agent-server:bb07a23-python-slim-arm64
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-python-slim-arm64
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-python-slim-arm64
ghcr.io/openhands/agent-server:bb07a23-python-node-runtime-slim-arm64
ghcr.io/openhands/agent-server:bb07a23-golang
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-golang
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-golang
ghcr.io/openhands/agent-server:bb07a23-golang_tag_1.21-bookworm
ghcr.io/openhands/agent-server:bb07a23-java
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-java
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-java
ghcr.io/openhands/agent-server:bb07a23-eclipse-temurin_tag_17-jdk
ghcr.io/openhands/agent-server:bb07a23-python-minimal
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-python-minimal
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-python-minimal
ghcr.io/openhands/agent-server:bb07a23-python-node-runtime-minimal
ghcr.io/openhands/agent-server:bb07a23-python-slim
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-python-slim
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-python-slim
ghcr.io/openhands/agent-server:bb07a23-python-node-runtime-slim
ghcr.io/openhands/agent-server:bb07a23-python
ghcr.io/openhands/agent-server:bb07a23fbd43948aae7665b99cafd48be66beeba-python
ghcr.io/openhands/agent-server:condenser-system-user-prompt-split-python
ghcr.io/openhands/agent-server:bb07a23-python-node-runtime

About Multi-Architecture Support

  • Each variant tag (e.g., bb07a23-python) is a multi-arch manifest supporting both amd64 and arm64
  • Docker automatically pulls the correct architecture for your platform
  • Individual architecture tags (e.g., bb07a23-python-amd64) are also available if needed

The LLMSummarizingCondenser previously sent its summarization request to the
condenser's own LLM as a single `user` message containing both the steering
instructions and the forgotten-event payload. This puts the event data in the
wrong slot on the OpenAI Responses API path: a `system` message serializes to
the `instructions` field with empty `input` (via Message.to_responses_value /
message_to_responses_dict returning [] for system), so the payload-to-summarize
never lands as a real input item.

Split the single `summarizing_prompt.j2` template into two:
- `summarizing_system.j2` -> sent as a `system` message (steering instructions)
- `summarizing_events.j2` -> sent as a `user` message (the <EVENT> payload)

This is the canonical system-preamble + user-payload shape for both the Chat
Completions and Responses APIs: the steering header goes to `instructions`/
`system` and the event data goes to `input`/`user`, where it belongs. It also
plays well with the subscription/Codex transport, which prepends system chunks
onto the first user message.

The change is isolated to the condenser's separate side-channel call
(self.llm.generate(..., store=False)) and does not touch the agent's main
conversation. Add a shared `_build_summary_messages` helper to dedupe the
sync and async generation paths.

Co-authored-by: openhands <openhands@all-hands.dev>
@github-actions

github-actions Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

REST API breakage checks (OpenAPI) — ✅ PASSED

Result: ✅ PASSED

Action log

@github-actions

github-actions Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Coverage

Coverage Report •
FileStmtsMissCoverMissing
openhands-sdk/openhands/sdk/context/condenser
   llm_summarizing_condenser.py1952985%376, 379–380, 383–384, 389, 394, 396–397, 410–411, 481–482, 487, 495, 511–512, 514–516, 521–525, 528, 532, 534–535
TOTAL437861212472% 

@all-hands-bot all-hands-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This review was posted by an AI agent (OpenHands).

Review summary

The code change itself is correct and well-scoped. _build_summary_messages() produces a canonical system + user pair, both the sync and async condensation paths now share it, the removed summarizing_prompt.j2 had no other references in the tree (verified with a repo-wide grep, including pyproject.toml/MANIFEST.in which glob *.j2), and the split is content-preserving against the old combined template. I confirmed the serialization empirically on the head checkout:

  • Chat Completions: roles ['system', 'user'].
  • Responses API: instructions = the steering text, input = one user item carrying the <EVENT> payload (this is the slot-mapping the linked issue describes).
  • Subscription/Codex: transform_for_subscription prepends the system chunk onto the first user message as intended.
  • Prompt caching: _apply_prompt_caching now marks the system block plus the trailing user block (two breakpoints), which is within Anthropic's limit.

pytest tests/sdk/context/condenser/test_llm_summarizing_condenser.py -> 44 passed, and the broader set (tests/sdk/context/, tests/sdk/conversation/test_condense.py, tests/sdk/agent/test_agent_context_window_condensation.py) -> 470 passed on this head. ruff check/ruff format --check are clean on the changed files.

I found no material correctness, security, or compatibility defect in the diff. Two blocking gaps remain, both process/evidence rather than code.

Blocking findings

1. Required integration evidence is missing for a prompt-template change

AGENTS.md (TESTING) states: "For changes to prompt templates, tool descriptions, or agent decision logic, add the integration-test label to trigger integration tests and verify no unexpected impact on benchmark performance." This PR rewrites the condenser's prompt templates (summarizing_system.j2, summarizing_events.j2) and changes the message roles sent on every condensation, which the review guide's "Agent behavior and evaluation" checkpoint explicitly covers (prompts, condensation).

No integration-test or condenser-test label is present, and no Run Integration Tests workflow ran for head 60ca1cc. The PR description reports only unit-test runs. The required eval evidence is therefore missing, which per the repository review guide is a COMMENT condition. Please add the integration-test (or condenser-test) label and confirm the results on this head before merge.

2. Current-head required check is failing: PR Description Check

The Validate PR description check failed on this head with six errors, including: the first visible line must be HUMAN:; a human-written note is required between HUMAN: and AGENT:; and the ## Why and ## How to Test template sections are absent. It also reports: "Linked issue(s) (#5142) carry neither ready-for-dev nor a pre-rollout creation date." Per the repository's PR_DESCRIPTION_HUMAN_CHECK policy, the HUMAN: section is reserved for a human and must not be filled in by an AI agent, so this needs the author to complete the description in their own words and the issue to be marked ready-for-dev.

Verdict

The implementation is sound and I would approve the code on its own merits. It is not mergeable at this head because the repository-required prompt-template integration evidence is absent and a required check is red. Both need to be resolved before approval.

🔄 CHANGES REQUESTED

@juanmichelini
juanmichelini requested review from all-hands-bot and removed request for all-hands-bot September 24, 2026 18:37
all-hands-bot

This comment was marked as outdated.

@juanmichelini juanmichelini added the integration-test Runs the integration tests and comments the results label Sep 26, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Hi! I started running the integration tests on your PR. You will receive a comment with the results shortly.

@github-actions

Copy link
Copy Markdown
Contributor

🧪 Integration Tests Results

Overall Success Rate: 98.0%
Total Cost: $1.97
Models Tested: 5
Timestamp: 2026-09-26 01:08:09 UTC

📁 Detailed Logs & Artifacts

Click the links below to access detailed agent/LLM logs showing the complete reasoning process for each model. On the GitHub Actions page, scroll down to the 'Artifacts' section to download the logs.

📊 Summary

Model Overall Tests Passed Skipped Total Cost Tokens
litellm_proxy_openai_gpt_5.5 100.0% 10/10 0 10 $0.89 333,465
litellm_proxy_gemini_3.1_pro_preview 100.0% 10/10 0 10 $0.48 348,429
litellm_proxy_anthropic_claude_sonnet_4_6 90.0% 9/10 0 10 $0.59 400,354
litellm_proxy_deepseek_deepseek_v4_flash 100.0% 10/10 0 10 $0.01 421,372
litellm_proxy_minimax_MiniMax_M2.7 100.0% 9/9 1 10 $0.00 444,282

📋 Detailed Results

litellm_proxy_openai_gpt_5.5

  • Success Rate: 100.0% (10/10)
  • Total Cost: $0.89
  • Token Usage: prompt: 328,287, completion: 5,178, cache_read: 201,728, reasoning: 1,777
  • Run Suffix: litellm_proxy_openai_gpt_5.5_d26b34b_gpt_5_5_run_N10_20260926_010549

litellm_proxy_gemini_3.1_pro_preview

  • Success Rate: 100.0% (10/10)
  • Total Cost: $0.48
  • Token Usage: prompt: 343,861, completion: 4,568, cache_read: 144,968, reasoning: 2,708
  • Run Suffix: litellm_proxy_gemini_3.1_pro_preview_d26b34b_gemini_3_1_pro_run_N10_20260926_010548

litellm_proxy_anthropic_claude_sonnet_4_6

  • Success Rate: 90.0% (9/10)
  • Total Cost: $0.59
  • Token Usage: prompt: 394,914, completion: 5,440, cache_read: 281,883, cache_write: 112,972, reasoning: 368
  • Run Suffix: litellm_proxy_anthropic_claude_sonnet_4_6_d26b34b_claude_sonnet_4_6_run_N10_20260926_010542

Failed Tests:

  • t02_add_bash_hello: Shell script is not executable (Cost: $0.06)

litellm_proxy_deepseek_deepseek_v4_flash

  • Success Rate: 100.0% (10/10)
  • Total Cost: $0.01
  • Token Usage: prompt: 415,431, completion: 5,941, cache_read: 352,256, reasoning: 1,550
  • Run Suffix: litellm_proxy_deepseek_deepseek_v4_flash_d26b34b_deepseek_v4_flash_run_N10_20260926_010549

litellm_proxy_minimax_MiniMax_M2.7

  • Success Rate: 100.0% (9/9)
  • Total Cost: $0.00
  • Token Usage: prompt: 438,626, completion: 5,656, cache_read: 347,891
  • Run Suffix: litellm_proxy_minimax_MiniMax_M2.7_d26b34b_minimax_m2_7_run_N10_20260926_010554
  • Skipped Tests: 1

Skipped Tests:

  • t08_image_file_viewing: This test requires a vision-capable LLM model. Please use a model that supports image input.

@all-hands-bot all-hands-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This review was posted by an AI agent (OpenHands).

Review summary

I re-checked this PR fresh at head bb07a23 and found no material correctness, security, compatibility, or acceptance-criterion defect. Both concerns from my earlier reviews are now resolved: the integration-test label is applied and the linked issue #5142 now carries ready-for-dev, and the Validate PR description check is green on this head (the body now has the HUMAN:/AGENT: markers, ## Why, and ## How to Test).

What I verified on this head:

  • Content preservation. Reconstructing the two new templates (summarizing_system.j2 + summarizing_events.j2) yields text identical to the removed combined summarizing_prompt.j2, so no instructions or payload text is lost or duplicated by the split.
  • No stale references. A repo-wide search finds no remaining reference to summarizing_prompt.j2; pyproject.toml (*.j2) and MANIFEST.in (recursive-include openhands-sdk *.j2) still package both new templates, so the rename does not drop them from the distribution.
  • Transport shape (real code paths). _build_summary_messages() returns ['system', 'user']; on the Responses path format_messages_for_responses yields the steering text in instructions and exactly one user input item carrying the <EVENT> payload. transform_for_subscription prepends the system chunk onto the first user message as described.
  • Prompt caching side effect. _apply_prompt_caching now marks the system block plus the trailing user block — two breakpoints, within provider limits, and a strict improvement over the previous single breakpoint.
  • Scope. The change is confined to the condenser's isolated side-channel call; CondensationSummaryEvent and the agent's main history are untouched, and no serialized event, setting, or API contract changes.

Evidence at this head: 52 check runs, 50 success / 2 skipped / 0 failure. pytest tests/sdk/context/condenser/test_llm_summarizing_condenser.py -> 45 passed; the broader set (tests/sdk/context/, tests/sdk/conversation/test_condense.py, tests/sdk/agent/test_agent_context_window_condensation.py) -> 474 passed. ruff check/ruff format --check clean on the changed packages.

Non-blocking observations

  • The label-triggered Python eval run (Run Integration Tests integration-test, run 36207121798) executed on commit d26b34b, which is an ancestor of this head, and reported 98.0% overall success with the workflow concluding success. The condenser files are byte-identical between d26b34b and bb07a23 (the only delta is unrelated grayswan security files), so the run did exercise the changed code. Flagging only for transparency, not as a blocker.
  • The single failed eval case (t02_add_bash_hello, "Shell script is not executable", Claude Sonnet 4.6) is a shell-permission assertion unrelated to condenser message roles; the condenser path is not exercised in that short test, so it is not a regression from this change.

Verdict

The change is correct, minimal, and adequately evidenced. Approving.

✅ APPROVED

@juanmichelini
juanmichelini merged commit 032e0e9 into main Sep 26, 2026
55 checks passed
@juanmichelini
juanmichelini deleted the condenser/system-user-prompt-split branch September 26, 2026 14:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

integration-test Runs the integration tests and comments the results

Projects

None yet

Development

Successfully merging this pull request may close these issues.

condenser: split summarization prompt into system + user messages

3 participants