Repository navigation
LCORE-3789: carry image attachments through compacted-mode requests - #2866
Open
max-svistunov wants to merge 4 commits into
Open
max-svistunov wants to merge 4 commits into
max-svistunov wants to merge 4 commits into
Conversation
Contributor
|
Warning Review limit reachedYou've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Next included review available in 59 minutes. View limit detailsLimit details: You’ve used the included review currently available. Review configuration: ⚙️ Run configuration
📒 Files selected for processing (10)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
In compacted mode the request input is overridden through extra_body with the explicit item list (summaries, recent turns, new query). That list is text only. Image attachments reach the model as ImageUrl parts of the pydantic-ai prompt, and the override replaces the prompt-derived input, so the images never reached the request body. OgxResponsesModel gets a step, run in request() and request_stream(). When the user prompt carries media, it replaces the trailing user message of the override with pydantic-ai's own mapping (the private _map_user_prompt, 2.27.1) of that message's text followed by the media, so the new query has the form it has outside compacted mode. The text comes from the override, not from the prompt: for an empty query agent_prompt_text returns the text of an earlier message. A text-only prompt leaves the override untouched. The 422 guard for such turns is still in place, so nothing changes for users yet. The next commit removes it. Tests: unit tests for the step, and two tests over httpx.MockTransport. One sends the same image prompt with and without the override through agent.run and compares the serialised bodies: no conversation parameter, text-only history unchanged, last input item equal to the non-compacted one, base64 payload present once. The other checks the body of a streamed request.
LCORE-3582 added reject_image_attachments_in_compacted_mode, which rejected a compacted turn with image attachments with a 422 error (the HTTP status on /v1/query; on /v1/streaming_query it was raised after the stream had started), because the explicit input could not carry the images and the model would have answered without seeing them. The previous commit sends the images of the current turn in compacted mode, so the guard has nothing left to protect against. This removes the guard, its two call sites in the /v1/query and /v1/streaming_query agent helpers, the imports that only it used, and its tests. The agent_prompt_text docstring now says where the images join the request. Tests: the two agent-helper tests that check the multimodal prompt for image attachments are parametrized over compacted mode. Before the removal the compacted cases failed with the 422.
Two properties of compacted mode held without a test, and the ticket asks that they stay true now that the current turn's images are sent: - An image from an earlier turn is not sent again. Recent turns are rendered from their text, so every explicit input item is plain text. - An image part adds nothing to the token estimate that triggers compaction. Only message text is counted, so a base64 payload never reaches the tokenizer. This adds one unit test for each. Each test fails when the property is broken on purpose: passing the stored content through in _verbatim_input_message, and counting image_url in extract_message_text. The design document and the user guide now describe image attachments in a compacted conversation: the current turn's images are sent, earlier images are not sent again, images are not counted in the estimate, and the stored turn holds text only. The design document also gets a changelog entry for the removed 422.
A request may carry an image and an empty query. In a compacted conversation agent_prompt_text skipped the empty trailing message of the explicit input and returned the text of the message before it, which is the last assistant answer or the summary. The first model request of the run was not affected, because its query text comes from the explicit input. Everything else that reads the prompt was: a shield got the earlier text to judge, and the second model request of a client-side tool loop sent the earlier text as the user's words next to the image. agent_prompt_text now returns the text of the trailing message item even when it is empty, which is what the prompt holds outside compacted mode. The warning stays for an explicit list without a message item. The _carry_prompt_media docstring no longer names the old fallback as the reason for taking the text from the explicit input. Tests: one unit test builds the explicit input for an empty query and expects an empty prompt. Before the change it got the text of the earlier assistant message.
max-svistunov
force-pushed
the
lcore-3789-compacted-mode-image-attachments
branch
from
October 7, 2026 22:50
b3d62e1 to
3675be3
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
LCORE-3789. A query with an image attachment on
/v1/queryor/v1/streaming_querywas rejected once its conversation had been compacted. The guard, which raises HTTP 422, came with LCORE-3582 (#2451). In compacted mode the requestinputis overridden throughextra_bodywith an explicit item list (summaries, recent turns, the new query), and that list is text only. An image attachment reaches the model as anImageUrlpart of the pydantic-ai prompt, and the override replaces the input derived from the prompt, so the image would have been dropped and the model would have answered without seeing it.This PR sends the images of the current turn in compacted mode and removes the guard, its two call sites and its tests. It closes LCORE-3789.
How it works.
OgxResponsesModelgets one step,_carry_prompt_media, run inrequest()andrequest_stream(). When the request carries the input override and the user prompt holds media, the step replaces the trailing user message of the override (the new query) with pydantic-ai's own mapping of that message's text followed by the media of the prompt. The trailing message keeps the text the explicit input already had and gains only the media. The items before it and the rest of the model settings are passed on unchanged. The new query then has the form it has outside compacted mode: a test sends the same image prompt with and without the override through the real OpenAI client and compares the last input item of the two serialised request bodies. A text-only prompt leaves the settings object as it is, so a text-only compacted request is sent as before. If the override does not end with a user message that has string content, the media is appended as a user message of its own and nothing is dropped;/v1/queryand/v1/streaming_queryalways build a string there.Where it departs from the ticket. The ticket's "Approach" says to thread
image_attachmentsinto the compaction path (step 1) and to render the trailing message in_query_input_messagefrom aninput_textpart and oneinput_imagepart per attachment (step 2). Neither is done.apply_compaction,_query_input_messageand the endpoint handlers are not touched. The reasons:_query_input_messagewould be a second mapping of an image to the Responses API, and it would have to stay equal to the one pydantic-ai applies outside compacted mode. The step calls that one.apply_compactionfrom its callers in the endpoints. LCORE-3908: give the turns we store ourselves one owner #2796 and LCORE-3910: count the summarization calls in quota, token counts and metrics #2797 rework the code at those call sites (what is passed toapply_compactionand what is done with its result), and LCORE-4108 moves compaction on these endpoints into the agent run.The price is a call to a private pydantic-ai method,
OpenAIResponsesModel._map_user_prompt. pydantic-ai is pinned (pydantic-ai==2.27.1inpyproject.toml),_model.pyalready imports from pydantic-ai's private modules (_run_context,_utils), and the tests run the real method, not a mock. If the private call is not acceptable, the alternative is the ticket's approach. Steps 3 and 4 of the ticket are followed as written: the non-compacted path is untouched, and the guard is gone.The replaced item is in pydantic-ai's form, which has no
"type": "message"key, while the history items before it have one. OGX accepted the mixed list in the live run described under Testing.What changes for users.
/v1/queryor/v1/streaming_queryis answered. On/v1/queryit used to get HTTP 422, "Image attachments are not supported on this conversation". On/v1/streaming_querythe same error was raised inside the stream, after thestartevent, so not as the HTTP status (read from the code onmain, not run).agent_prompt_textskipped an empty trailing message and returned the text of the message before it, which is the last answer or the summary. The first model request of the run was not affected, because its query text comes from the explicit input. Everything that reads the agent prompt was: the shield capabilities judged the earlier text, and the second model request of a client-side tool loop sent it as the user's words.agent_prompt_textnow returns the empty string, which is what the prompt holds outside compacted mode. This also applies to a text-only empty query: the shields and the agent prompt now get the empty string there, where they got the earlier text. The fix is in this PR because an image without text is the realistic case of an empty query, and this PR is what lets that request through in compacted mode.Known limits and what is not covered.
main. Compacted image turns, rejected until now, behave the same. Not changed here; it is LCORE-4593.gpt-4o-mini). Server mode was not run, no e2e scenario is added, and the e2e suite was not run. The automated coverage is unit tests, two of them on the serialised request body overhttpx.MockTransport: non-streamed throughagent.run, streamed throughrequest_streamcalled directly, not through the agent./v1/responsesis not touched. The ticket notes that it is unaffected: aninput_imagepart sent there is in the input itself, and_query_input_messagepasses an item list through unchanged.Relation to the other open PRs. This PR does not depend on #2796 (LCORE-3908) or #2797 (LCORE-3910), and nothing has to merge first. It conflicts with #2796 in one place: both edit the
utils.conversation_compactionimport block insrc/utils/agents/query.pyandsrc/utils/agents/streaming.py. This PR removes the guard's name from it, #2796 removesstore_compacted_turn. Whichever merges second gets a one-line conflict in each of the two files, resolved by dropping both names. A trial merge with the #2796 commit rebased onto the samemain(acaa1496) conflicts there and nowhere else; the other files both PRs touch merge automatically. #2797 is stacked on #2796 and meets the same conflict.Four commits, each passing its tests on its own: the step, with the guard still in place; the guard removed; tests and documentation for the ticket's two "watch out for" points; the empty query.
src/pydantic_ai_lightspeed/ogx/_model.py_prompt_media(),OgxResponsesModel._carry_prompt_media(), called inrequest()andrequest_stream()src/utils/conversation_compaction.pyreject_image_attachments_in_compacted_mode()removed;agent_prompt_text()keeps an empty query emptysrc/utils/agents/query.py,src/utils/agents/streaming.pytests/unit/pydantic_ai_lightspeed/ogx/test_compacted_input_media.pytests/unit/utils/agents/test_query.py,tests/unit/utils/agents/test_streaming.pytests/unit/utils/test_conversation_compaction.pydocs/user_doc/conversation_compaction.md,docs/design/conversation-compaction/conversation-compaction.mdType of change
pyproject.toml+uv.lock]requirements.*.txtfor Konflux]Tools used to create PR
Identify any AI code assistants used in this PR (for transparency and review context)
Related Tickets & Documents
Checklist before requesting a review
Testing
inference.context_windows: openai/gpt-4o-mini: 2000;compaction: enabled: true, threshold_ratio: 0.1, token_floor: 100, buffer_turns: 1). Send text queries to/v1/queryon one conversation until the response reportscontext_status: summarized. Then send, on the same conversation, a query with an image attachment to/v1/queryand to/v1/streaming_query. Use a PNG that shows a code which appears nowhere in the text:Expect HTTP 200,
context_status: summarizedand the code in the answer. Without this PR the guard rejects both requests (that state was not run live).This was run in library mode only, with
gpt-4o-mini, and before the last commit. The last commit changes whatagent_prompt_textreturns for an empty query, and a docstring; the run was not repeated after it. Output of the run (the second image shows KIWI-2086, the third and fourth are sent with an empty query):To see what reaches the model, point
base_urlof theremote::openaiprovider at a recording proxy in front of the OpenAI API. In the same run each image turn carried one image, the earlier image turn was replayed as text, and the follow-up text turn carried no image (one message per line, long texts cut):The stored items after the run hold no image for the compacted conversation (first row). The second row is the fresh conversation of the empty-query check, which is not compacted, so OGX stored its turn itself:
The answer to an image sent with an empty query varied between runs of this check:
LIME-6610above, andMANGO-7145,Unknown.andI don't know who or what this is.in three other runs. Two of the four runs recorded the upstream request, and in both it was the same, an empty text part followed by the image (answersLIME-6610andI don't know who or what this is.). So the empty-query step shows that the request is accepted and carries the image, not that the model reads an image reliably when no question comes with it. The two image turns with a question returned the code in all three runs that included them; the first of them ran on an earlier revision of the step.What was checked: with the call of the step removed from
request()or fromrequest_stream(), the matching transport test fails (test_request_body_carries_the_image_as_outside_compacted_mode,test_streamed_request_body_carries_the_image). With the guard restored, the compacted case oftest_success_with_image_attachments_sends_multimodal_promptand oftest_streams_with_image_attachments_passes_multimodal_promptfails with the 422. The two tests of behaviour this PR does not change fail when it is broken on purpose:test_build_explicit_input_does_not_replay_earlier_imageswhen_verbatim_input_messagepasses the stored content through,test_estimate_does_not_count_image_partswhenextract_message_textcountsimage_url. Before the last committest_agent_prompt_text_empty_query_stays_emptyfails withassert 'prior answer' == ''.black, ruff, docstyle, pylint and pyright pass. mypy reports what it reports on
mainand nothing new: one error insrc/models/config.pyand 49 errors in 10 test files.