Skip to content

LCORE-3789: carry image attachments through compacted-mode requests - #2866

Open
max-svistunov wants to merge 4 commits into
lightspeed-core:mainfrom
max-svistunov:lcore-3789-compacted-mode-image-attachments
Open

max-svistunov wants to merge 4 commits into
lightspeed-core:mainfrom
max-svistunov:lcore-3789-compacted-mode-image-attachments

Conversation

@max-svistunov

@max-svistunov max-svistunov commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Description

LCORE-3789. A query with an image attachment on /v1/query or /v1/streaming_query was rejected once its conversation had been compacted. The guard, which raises HTTP 422, came with LCORE-3582 (#2451). In compacted mode the request input is overridden through extra_body with an explicit item list (summaries, recent turns, the new query), and that list is text only. An image attachment reaches the model as an ImageUrl part of the pydantic-ai prompt, and the override replaces the input derived from the prompt, so the image would have been dropped and the model would have answered without seeing it.

This PR sends the images of the current turn in compacted mode and removes the guard, its two call sites and its tests. It closes LCORE-3789.

How it works. OgxResponsesModel gets one step, _carry_prompt_media, run in request() and request_stream(). When the request carries the input override and the user prompt holds media, the step replaces the trailing user message of the override (the new query) with pydantic-ai's own mapping of that message's text followed by the media of the prompt. The trailing message keeps the text the explicit input already had and gains only the media. The items before it and the rest of the model settings are passed on unchanged. The new query then has the form it has outside compacted mode: a test sends the same image prompt with and without the override through the real OpenAI client and compares the last input item of the two serialised request bodies. A text-only prompt leaves the settings object as it is, so a text-only compacted request is sent as before. If the override does not end with a user message that has string content, the media is appended as a user message of its own and nothing is dropped; /v1/query and /v1/streaming_query always build a string there.

Where it departs from the ticket. The ticket's "Approach" says to thread image_attachments into the compaction path (step 1) and to render the trailing message in _query_input_message from an input_text part and one input_image part per attachment (step 2). Neither is done. apply_compaction, _query_input_message and the endpoint handlers are not touched. The reasons:

The price is a call to a private pydantic-ai method, OpenAIResponsesModel._map_user_prompt. pydantic-ai is pinned (pydantic-ai==2.27.1 in pyproject.toml), _model.py already imports from pydantic-ai's private modules (_run_context, _utils), and the tests run the real method, not a mock. If the private call is not acceptable, the alternative is the ticket's approach. Steps 3 and 4 of the ticket are followed as written: the non-compacted path is untouched, and the guard is gone.

The replaced item is in pydantic-ai's form, which has no "type": "message" key, while the history items before it have one. OGX accepted the mixed list in the live run described under Testing.

What changes for users.

  • A compacted turn with an image attachment on /v1/query or /v1/streaming_query is answered. On /v1/query it used to get HTTP 422, "Image attachments are not supported on this conversation". On /v1/streaming_query the same error was raised inside the stream, after the start event, so not as the HTTP status (read from the code on main, not run).
  • Only the images of the current turn are sent. Recent turns are rendered from their text and summaries are text, so an image from an earlier turn is not sent again once the conversation is compacted. That held before this PR and now has a test.
  • The stored turn stays text only. In compacted mode lightspeed-stack appends the turn to the conversation itself, and it stores the text of the query without the image. Outside compacted mode OGX stores the turn, image included. A side effect: an image sent with an empty query is stored as an empty user message, and an empty message is skipped when recent turns are rendered, so the next compacted request carries the answer of that turn without a user message before it.
  • An empty query in a compacted conversation, with or without an image (the last commit). agent_prompt_text skipped an empty trailing message and returned the text of the message before it, which is the last answer or the summary. The first model request of the run was not affected, because its query text comes from the explicit input. Everything that reads the agent prompt was: the shield capabilities judged the earlier text, and the second model request of a client-side tool loop sent it as the user's words. agent_prompt_text now returns the empty string, which is what the prompt holds outside compacted mode. This also applies to a text-only empty query: the shields and the agent prompt now get the empty string there, where they got the earlier text. The fix is in this PR because an image without text is the realistic case of an empty query, and this PR is what lets that request through in compacted mode.

Known limits and what is not covered.

  • Images are not counted in the token estimate that triggers compaction. The estimate reads message text only, so a compacted request with images can exceed the context window the compaction was meant to fit. The ticket asks to check this. It is now stated in the user guide and in the design document and pinned by a test, and it is not fixed here. The follow-up is LCORE-4601.
  • Capabilities that rewrite the prompt text, such as PII redaction, do not reach the query text of a compacted request: that text comes from the explicit input, which is built before the agent run. Compacted text turns behave this way on main. Compacted image turns, rejected until now, behave the same. Not changed here; it is LCORE-4593.
  • That the model's answer reflects the image (the first acceptance criterion) is shown by a local run only: library mode, one model (gpt-4o-mini). Server mode was not run, no e2e scenario is added, and the e2e suite was not run. The automated coverage is unit tests, two of them on the serialised request body over httpx.MockTransport: non-streamed through agent.run, streamed through request_stream called directly, not through the agent.
  • /v1/responses is not touched. The ticket notes that it is unaffected: an input_image part sent there is in the input itself, and _query_input_message passes an item list through unchanged.

Relation to the other open PRs. This PR does not depend on #2796 (LCORE-3908) or #2797 (LCORE-3910), and nothing has to merge first. It conflicts with #2796 in one place: both edit the utils.conversation_compaction import block in src/utils/agents/query.py and src/utils/agents/streaming.py. This PR removes the guard's name from it, #2796 removes store_compacted_turn. Whichever merges second gets a one-line conflict in each of the two files, resolved by dropping both names. A trial merge with the #2796 commit rebased onto the same main (acaa1496) conflicts there and nowhere else; the other files both PRs touch merge automatically. #2797 is stacked on #2796 and meets the same conflict.

Four commits, each passing its tests on its own: the step, with the guard still in place; the guard removed; tests and documentation for the ticket's two "watch out for" points; the empty query.

File Change
src/pydantic_ai_lightspeed/ogx/_model.py _prompt_media(), OgxResponsesModel._carry_prompt_media(), called in request() and request_stream()
src/utils/conversation_compaction.py reject_image_attachments_in_compacted_mode() removed; agent_prompt_text() keeps an empty query empty
src/utils/agents/query.py, src/utils/agents/streaming.py the guard call and its import removed
tests/unit/pydantic_ai_lightspeed/ogx/test_compacted_input_media.py new, 7 tests: the step, and the serialised request body (non-streamed and streamed)
tests/unit/utils/agents/test_query.py, tests/unit/utils/agents/test_streaming.py the image-prompt tests also run in compacted mode
tests/unit/utils/test_conversation_compaction.py guard tests removed; earlier images are not replayed; images are not in the estimate; an empty query stays empty
docs/user_doc/conversation_compaction.md, docs/design/conversation-compaction/conversation-compaction.md image attachments in a compacted conversation; changelog entry

Type of change

  • Refactor
  • New feature
  • Bug fix
  • CVE fix
  • Optimization
  • Documentation Update
  • Configuration Update
  • Bump-up service version
  • Bump-up dependent library [pyproject.toml + uv.lock]
  • Bump-up dependent library [requirements.*.txt for Konflux]
  • Bump-up library or tool used for development (does not change the final image)
  • CI configuration change
  • Konflux configuration change
  • Unit tests improvement
  • Integration tests improvement
  • End to end tests improvement
  • Benchmarks improvement

Tools used to create PR

Identify any AI code assistants used in this PR (for transparency and review context)

  • Assisted-by: Claude Opus 4.8
  • Generated by: Claude Opus 4.8

Related Tickets & Documents

  • Related Issue # LCORE-3789
  • Closes # LCORE-3789

Checklist before requesting a review

  • I have performed a self-review of my code.
  • PR has passed all pre-merge test jobs.
  • If it is a core feature, I have added thorough tests.

Testing

  1. Check on a running service that a compacted turn with an image is answered and that the answer reflects the image. Start the service in library mode with compaction forced by a small context window (inference.context_windows: openai/gpt-4o-mini: 2000; compaction: enabled: true, threshold_ratio: 0.1, token_floor: 100, buffer_turns: 1). Send text queries to /v1/query on one conversation until the response reports context_status: summarized. Then send, on the same conversation, a query with an image attachment to /v1/query and to /v1/streaming_query. Use a PNG that shows a code which appears nowhere in the text:
POST /v1/query
{"query": "Read the code written in the attached image and reply with that code only.",
 "conversation_id": "<id>",
 "attachments": [{"attachment_type": "image", "content_type": "image/png", "content": "<base64 PNG showing PLUM-7391>"}]}

Expect HTTP 200, context_status: summarized and the code in the answer. Without this PR the guard rejects both requests (that state was not run live).

This was run in library mode only, with gpt-4o-mini, and before the last commit. The last commit changes what agent_prompt_text returns for an empty query, and a docstring; the run was not repeated after it. Output of the run (the second image shows KIWI-2086, the third and fourth are sent with an empty query):

turn 1: HTTP 200 context_status=full response='OK.'
turn 2: HTTP 200 context_status=full response='OK.'
turn 3: HTTP 200 context_status=summarized response='atlas-prod.'
/v1/query image turn: HTTP 200 context_status=summarized response='PLUM-7391'
/v1/streaming_query image turn: HTTP 200 events=['compaction', 'end', 'start', 'token', 'turn_complete'] text='KIWI-2086' end.context_status=summarized
empty query + image FIG-5524, fresh conversation: HTTP 200 context_status=full response='It looks like you\'ve provided a text snippet labeling a figure as "FIG-5524." If you have any questions or need further assistance related to this figure, feel free to ask!'
empty query + image LIME-6610, compacted conversation: HTTP 200 context_status=summarized response='LIME-6610'
follow-up text turn: HTTP 200 context_status=summarized response='No.'

To see what reaches the model, point base_url of the remote::openai provider at a recording proxy in front of the OpenAI API. In the same run each image turn carried one image, the earlier image turn was replayed as text, and the follow-up text turn carried no image (one message per line, long texts cut):

/v1/query image turn
  system    You are a helpful assistant
  user      Summary of earlier conversation: **User's Original Question an...
  user      Summary of earlier conversation: **User's Original Question an...
  user      What is the name of my cluster? One word.
  assistant atlas-prod.
  user      [text:'Read the code written in the attached image and reply with that code o', image_url:<9514 chars>]

/v1/streaming_query image turn
  system    You are a helpful assistant
  user      Summary of earlier conversation: **User's Original Question an...
  user      Summary of earlier conversation: **User's Original Question an...
  user      Summary of earlier conversation: User's Question and Environme...
  user      Read the code written in the attached image and reply with tha...
  assistant PLUM-7391
  user      [text:'Read the code written in the attached image and reply with that code o', image_url:<10618 chars>]

empty query + image, compacted conversation
  system    You are a helpful assistant
  user      Summary of earlier conversation: **User's Original Question an...
  user      Summary of earlier conversation: **User's Original Question an...
  user      Summary of earlier conversation: User's Question and Environme...
  user      Summary of earlier conversation: User's Question: The user req...
  user      Read the code written in the attached image and reply with tha...
  assistant KIWI-2086
  user      [text:'', image_url:<8250 chars>]

follow-up text turn
  system    You are a helpful assistant
  user      Summary of earlier conversation: **User's Original Question an...
  user      Summary of earlier conversation: **User's Original Question an...
  user      Summary of earlier conversation: User's Question and Environme...
  user      Summary of earlier conversation: User's Question: The user req...
  user      Summary of earlier conversation: **User's Original Question an...
  assistant LIME-6610
  user      Is an image attached to this message? Answer yes or no.

The stored items after the run hold no image for the compacted conversation (first row). The second row is the fresh conversation of the empty-query check, which is not compacted, so OGX stored its turn itself:

sqlite3 -readonly sql_store.db "select conversation_id, count(*), sum(item_data like '%input_image%'), sum(item_data like '%base64,%') from conversation_items group by conversation_id"
conv_223714ec...|19|0|0
conv_e7972296...|2|1|1

The answer to an image sent with an empty query varied between runs of this check: LIME-6610 above, and MANGO-7145, Unknown. and I don't know who or what this is. in three other runs. Two of the four runs recorded the upstream request, and in both it was the same, an empty text part followed by the image (answers LIME-6610 and I don't know who or what this is.). So the empty-query step shows that the request is accepted and carries the image, not that the model reads an image reliably when no question comes with it. The two image turns with a question returned the code in all three runs that included them; the first of them ran on an earlier revision of the step.

  1. Run the tests for this change:
uv run pytest tests/unit/pydantic_ai_lightspeed/ogx tests/unit/utils/agents tests/unit/utils/test_conversation_compaction.py \
    tests/integration/endpoints/test_compaction_query.py tests/integration/endpoints/test_compaction_streaming_query.py -q   # 269 passed

What was checked: with the call of the step removed from request() or from request_stream(), the matching transport test fails (test_request_body_carries_the_image_as_outside_compacted_mode, test_streamed_request_body_carries_the_image). With the guard restored, the compacted case of test_success_with_image_attachments_sends_multimodal_prompt and of test_streams_with_image_attachments_passes_multimodal_prompt fails with the 422. The two tests of behaviour this PR does not change fail when it is broken on purpose: test_build_explicit_input_does_not_replay_earlier_images when _verbatim_input_message passes the stored content through, test_estimate_does_not_count_image_parts when extract_message_text counts image_url. Before the last commit test_agent_prompt_text_empty_query_stays_empty fails with assert 'prior answer' == ''.

  1. Run the full suites:
uv run pytest tests/unit -q   # 3715 passed, 1 skipped (main at acaa1496: 3707 passed, 1 skipped)
uv run pytest tests/integration --ignore=tests/integration/container_lifecycle -q   # 326 passed (main: 326 passed)
  1. Format and linters:
uv run make format   # 539 files left unchanged

black, ruff, docstyle, pylint and pyright pass. mypy reports what it reports on main and nothing new: one error in src/models/config.py and 49 errors in 10 test files.

@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Next included review available in 59 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration
  • Configuration used: Repository: lightspeed-core/lightspeed-stack/.coderabbit.yaml
  • Review profile: ASSERTIVE
  • Plan: Advanced
  • Run ID: 34a9e39c-a549-44fc-8f3e-8f02ca608667
📥 Commits

Reviewing files that changed from the base of the PR and between 5737342 and 3675be3.

📒 Files selected for processing (10)
  • docs/design/conversation-compaction/conversation-compaction.md
  • docs/user_doc/conversation_compaction.md
  • src/pydantic_ai_lightspeed/ogx/_model.py
  • src/utils/agents/query.py
  • src/utils/agents/streaming.py
  • src/utils/conversation_compaction.py
  • tests/unit/pydantic_ai_lightspeed/ogx/test_compacted_input_media.py
  • tests/unit/utils/agents/test_query.py
  • tests/unit/utils/agents/test_streaming.py
  • tests/unit/utils/test_conversation_compaction.py
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

In compacted mode the request input is overridden through extra_body
with the explicit item list (summaries, recent turns, new query). That
list is text only. Image attachments reach the model as ImageUrl parts
of the pydantic-ai prompt, and the override replaces the prompt-derived
input, so the images never reached the request body.

OgxResponsesModel gets a step, run in request() and request_stream().
When the user prompt carries media, it replaces the trailing user
message of the override with pydantic-ai's own mapping (the private
_map_user_prompt, 2.27.1) of that message's text followed by the media,
so the new query has the form it has outside compacted mode. The text
comes from the override, not from the prompt: for an empty query
agent_prompt_text returns the text of an earlier message. A text-only
prompt leaves the override untouched.

The 422 guard for such turns is still in place, so nothing changes for
users yet. The next commit removes it.

Tests: unit tests for the step, and two tests over httpx.MockTransport.
One sends the same image prompt with and without the override through
agent.run and compares the serialised bodies: no conversation
parameter, text-only history unchanged, last input item equal to the
non-compacted one, base64 payload present once. The other checks the
body of a streamed request.
LCORE-3582 added reject_image_attachments_in_compacted_mode, which
rejected a compacted turn with image attachments with a 422 error
(the HTTP status on /v1/query; on /v1/streaming_query it was raised
after the stream had started), because the explicit input could not
carry the images and the model would have answered without seeing
them. The previous commit sends the images of
the current turn in compacted mode, so the guard has nothing left to
protect against.

This removes the guard, its two call sites in the /v1/query and
/v1/streaming_query agent helpers, the imports that only it used, and
its tests. The agent_prompt_text docstring now says where the images
join the request.

Tests: the two agent-helper tests that check the multimodal prompt for
image attachments are parametrized over compacted mode. Before the
removal the compacted cases failed with the 422.
Two properties of compacted mode held without a test, and the ticket
asks that they stay true now that the current turn's images are sent:

- An image from an earlier turn is not sent again. Recent turns are
  rendered from their text, so every explicit input item is plain text.
- An image part adds nothing to the token estimate that triggers
  compaction. Only message text is counted, so a base64 payload never
  reaches the tokenizer.

This adds one unit test for each. Each test fails when the property is
broken on purpose: passing the stored content through in
_verbatim_input_message, and counting image_url in
extract_message_text.

The design document and the user guide now describe image attachments
in a compacted conversation: the current turn's images are sent,
earlier images are not sent again, images are not counted in the
estimate, and the stored turn holds text only. The design document
also gets a changelog entry for the removed 422.
A request may carry an image and an empty query. In a compacted
conversation agent_prompt_text skipped the empty trailing message of
the explicit input and returned the text of the message before it,
which is the last assistant answer or the summary. The first model
request of the run was not affected, because its query text comes from
the explicit input. Everything else that reads the prompt was: a shield
got the earlier text to judge, and the second model request of a
client-side tool loop sent the earlier text as the user's words next to
the image.

agent_prompt_text now returns the text of the trailing message item
even when it is empty, which is what the prompt holds outside compacted
mode. The warning stays for an explicit list without a message item.
The _carry_prompt_media docstring no longer names the old fallback as
the reason for taking the text from the explicit input.

Tests: one unit test builds the explicit input for an empty query and
expects an empty prompt. Before the change it got the text of the
earlier assistant message.
@max-svistunov
max-svistunov force-pushed the lcore-3789-compacted-mode-image-attachments branch from b3d62e1 to 3675be3 Compare October 7, 2026 22:50

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant