Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -233,6 +233,8 @@ After compaction, lightspeed-stack writes the summary as a marked conversation i

When building context for a compacted conversation, lightspeed-stack fetches the conversation items, reads the active summaries from the summary cache (LCORE-1571) — falling back to the marker texts when no persisting cache is configured — takes the items after the last marker as the recent verbatim buffer, and sends `[summaries] + [recent items] + [new query]` as **explicit input**, **without** the `conversation` parameter. This is necessary because OGX reloads the *full* stored message history whenever the `conversation` parameter is set — there is no marker-based selection hook (verified empirically; see the Changelog). Each completed turn is then appended back to the conversation items by lightspeed-stack, since OGX no longer auto-stores it.

Image attachments (LCORE-3789): on `/v1/query` and `/v1/streaming_query` an image attachment is not part of the text input. It reaches the model as a part of the pydantic-ai user prompt, and the explicit input is text only. Whenever the prompt carries an image, `OgxResponsesModel` therefore replaces the trailing message of the explicit input, the new query, with pydantic-ai's mapping of the same text followed by the images of the prompt, so the new query is sent in the form it has outside compacted mode. Only the images of the current turn are sent. Recent turns are rendered from their text and summaries are text, so an image from an earlier turn is not sent again once the conversation is compacted. Images are not counted in the token estimate that triggers compaction, which reads message text only. The turn that lightspeed-stack appends to the conversation holds the text input, without the image.

This preserves a single continuous conversation identity. The `conversation_id` never changes and the user sees one conversation in the UI. OGX stores the full history, including the summary marker items, and returns it through its Conversations API. lightspeed-stack's own `GET /v1/conversations/{conversation_id}` leaves the marker items out of the chat history it returns (LCORE-3909): they are stored as user messages, but the user never sent them. Stored conversations are not migrated: the markers must stay in storage as the fallback source of truth, so they are filtered when the conversation is read.

## API response changes
Expand Down Expand Up @@ -446,6 +448,13 @@ never holds a marker, and is unchanged. Earlier revisions of this document said
that the Conversations API returns the marker items; that holds for the OGX
Conversations API only.

**2026-10-06 — Image attachments in compacted mode (LCORE-3789).**
A compacted turn on `/v1/query` or `/v1/streaming_query` that carried an image
attachment was rejected by a guard (LCORE-3582; HTTP 422 on `/v1/query`),
because the explicit input is text only. `OgxResponsesModel` now merges the
images of the current turn into the explicit input, and the guard is removed.
Details are in "Changed request flow after compaction".

# Appendix A: PoC Evidence

A proof-of-concept was built and tested.
Expand Down
9 changes: 9 additions & 0 deletions docs/user_doc/conversation_compaction.md
Original file line number Diff line number Diff line change
Expand Up @@ -128,6 +128,15 @@ Compaction triggers when **all** of the following are true:
3. The estimated token count of the conversation exceeds `threshold_ratio × context_window`
4. The estimated token count is at least `token_floor`

### Image attachments

A query with image attachments on `/v1/query` or `/v1/streaming_query` is answered in a compacted conversation as well:

- The images attached to the current query are sent to the LLM, together with the summary and the recent turns.
- Images from earlier turns are not sent again once the conversation is compacted. The LLM gets the text of those turns, or their summary.
- Images are not counted in the token estimate that triggers compaction. Only text is counted.
- The turn that the service stores in the conversation holds the text of the query, without the image.

### Degrading guard

The `buffer_turns` setting specifies a target number of recent turns to preserve. If the selected buffer turns exceed the available budget (`buffer_max_ratio × context_window`), the system reduces the buffer by one turn pair at a time until the budget fits. In extreme cases, the buffer can shrink to zero turns, meaning only the summary and the current query are sent to the LLM.
Expand Down
82 changes: 79 additions & 3 deletions src/pydantic_ai_lightspeed/ogx/_model.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@
from __future__ import annotations as _annotations

from collections import defaultdict
from collections.abc import AsyncIterator, Callable
from collections.abc import AsyncIterator, Callable, Sequence
from contextlib import asynccontextmanager
from typing import Any, Final, Optional, cast

Expand All @@ -31,7 +31,14 @@
from pydantic_ai import UnexpectedModelBehavior
from pydantic_ai._run_context import RunContext
from pydantic_ai._utils import PeekableAsyncStream, Unset, number_to_datetime
from pydantic_ai.messages import ModelMessage, ModelResponse
from pydantic_ai.messages import (
ModelMessage,
ModelRequest,
ModelResponse,
UserContent,
UserPromptPart,
is_multi_modal_content,
)
from pydantic_ai.models import (
ModelRequestParameters,
StreamedResponse,
Expand Down Expand Up @@ -82,7 +89,8 @@ def _model_settings_from_responses_params(
# the wire ``input`` from the prompt alone. Overriding via extra_body
# replaces it with the explicit list, exactly as the non-agent
# /v1/responses path sends it. Dropped again on tool-loop
# continuations — see ``_prepare_compacted_input``.
# continuations — see ``_prepare_compacted_input``. The media of the
# current prompt is merged back in by ``_carry_prompt_media``.
extra_body["input"] = payload["input"]
settings_dict: dict[str, Any] = {}
if extra_body:
Expand All @@ -105,6 +113,28 @@ def _model_settings_from_responses_params(
return cast("OpenAIResponsesModelSettings", settings_dict)


def _prompt_media(messages: Sequence[ModelMessage]) -> list[UserContent]:
"""Return the media items, such as images, of the user prompt being sent.

Parameters:
messages: Model messages for the request.

Returns:
The media items of the last user prompt part, in prompt order. Empty
when there is no user prompt or it is text only.
"""
prompts = [
part
for message in messages
if isinstance(message, ModelRequest)
for part in message.parts
if isinstance(part, UserPromptPart)
]
if not prompts or isinstance(prompts[-1].content, str):
return []
return [item for item in prompts[-1].content if is_multi_modal_content(item)]


class _FilteredResponseStream:
"""Wraps an OpenAI AsyncStream to reorder spurious events from OGX.

Expand Down Expand Up @@ -348,6 +378,7 @@ async def request( # pylint: disable=unused-argument
messages, model_settings
)
model_settings = self._prepare_compacted_input(messages, model_settings)
model_settings = await self._carry_prompt_media(messages, model_settings)
return await super().request(messages, model_settings, model_request_parameters)

def _prepare_conversation_continuation(
Expand Down Expand Up @@ -419,6 +450,50 @@ def _prepare_compacted_input(
new_settings["extra_body"] = new_extra_body
return cast("ModelSettings", new_settings)

async def _carry_prompt_media(
self,
messages: list[ModelMessage],
model_settings: Optional[ModelSettings],
) -> Optional[ModelSettings]:
"""Give the compacted ``input`` override the media of the current prompt.

The explicit item list is text only, while image attachments reach the
model as ``ImageUrl`` parts of the user prompt, so they would be lost
when the override replaces the prompt-derived ``input`` (LCORE-3789).
When the prompt carries media, the trailing user message of the
override (the new query) is replaced by pydantic-ai's own mapping of
its text followed by that media, which is the form the query has
outside compacted mode. The text stays the one the override carries,
so the explicit list is the only source of the query text. A text-only
prompt leaves the override as it was built.

Parameters:
messages: Model messages for the request.
model_settings: Model settings, possibly carrying the override.

Returns:
The settings unchanged, or a copy whose override ends with the
new query and its media.
"""
if not model_settings or not isinstance(model_settings, dict):
return model_settings
extra_body = model_settings.get("extra_body")
if not isinstance(extra_body, dict) or "input" not in extra_body:
return model_settings
content = _prompt_media(messages)
if not content:
return model_settings

history = list(extra_body["input"])
query = history[-1] if history else {}
if query.get("role") == "user" and isinstance(query.get("content"), str):
history.pop()
content = [query["content"], *content]
new_query = await self._map_user_prompt(UserPromptPart(content=content))
new_settings = dict(model_settings)
new_settings["extra_body"] = {**extra_body, "input": [*history, new_query]}
return cast("ModelSettings", new_settings)

@asynccontextmanager
async def request_stream( # pylint: disable=unused-argument
self,
Expand Down Expand Up @@ -446,6 +521,7 @@ async def request_stream( # pylint: disable=unused-argument
messages, model_settings
)
model_settings = self._prepare_compacted_input(messages, model_settings)
model_settings = await self._carry_prompt_media(messages, model_settings)

model_settings_cast = cast("OpenAIResponsesModelSettings", model_settings or {})
response = await self._responses_create(
Expand Down
4 changes: 0 additions & 4 deletions src/utils/agents/query.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,6 @@
)
from utils.conversation_compaction import (
agent_prompt_text,
reject_image_attachments_in_compacted_mode,
store_compacted_turn,
)
from utils.otel_tracing import (
Expand Down Expand Up @@ -282,9 +281,6 @@ async def retrieve_agent_response(
no_tools=no_tools,
)
logger.debug("Starting agent non-streaming response processing")
reject_image_attachments_in_compacted_mode(
responses_params, image_attachments
)
if image_attachments:
prompt = build_multimodal_input(
agent_prompt_text(responses_params),
Expand Down
2 changes: 0 additions & 2 deletions src/utils/agents/streaming.py
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,6 @@
)
from utils.conversation_compaction import (
agent_prompt_text,
reject_image_attachments_in_compacted_mode,
store_compacted_turn,
)
from utils.otel_tracing import (
Expand Down Expand Up @@ -416,7 +415,6 @@ async def agent_response_generator(
rag_id_mapping=context.rag_id_mapping,
turn_summary=turn_summary,
)
reject_image_attachments_in_compacted_mode(responses_params, image_attachments)
if image_attachments:
prompt = build_multimodal_input(
agent_prompt_text(responses_params),
Expand Down
55 changes: 7 additions & 48 deletions src/utils/conversation_compaction.py
Original file line number Diff line number Diff line change
Expand Up @@ -49,15 +49,13 @@
from dataclasses import dataclass
from typing import Any, Optional, cast

from fastapi import HTTPException
from ogx_api.openai_responses import OpenAIResponseMessage
from ogx_client import AsyncOgxClient

from cache.cache import Cache
from cache.cache_error import CacheError
from configuration import configuration
from log import get_logger
from models.api.responses.error import UnprocessableEntityResponse
from models.common.responses.responses_api_params import ResponsesApiParams
from models.common.responses.types import ResponseInput
from models.common.turn_summary import ContextStatus
Expand Down Expand Up @@ -351,7 +349,9 @@ def agent_prompt_text(params: ResponsesApiParams) -> str:
the new user query is the trailing message item. The agent pipeline still
needs a plain string prompt (capabilities and multimodal input operate on
it); the full explicit list reaches the request body separately via the
``extra_body`` input override (LCORE-3582).
``extra_body`` input override (LCORE-3582). Image attachments added to
this prompt are merged into that override by ``OgxResponsesModel``
(LCORE-3789).

Args:
params: Prepared (possibly compaction-rewritten) request parameters.
Expand All @@ -365,9 +365,10 @@ def agent_prompt_text(params: ResponsesApiParams) -> str:
return params.input
for item in reversed(list(params.input)):
if is_message_item(item):
text = extract_message_text(item)
if text:
return text
# The trailing message is the new query. An empty one (an image
# sent without text) stays empty rather than borrowing the text
# of an earlier message.
return extract_message_text(item)
# The wire input still carries the explicit list via the extra_body
# override, so the request itself is well-formed — but capabilities and
# multimodal construction operate on this prompt, and an explicit input
Expand All @@ -380,48 +381,6 @@ def agent_prompt_text(params: ResponsesApiParams) -> str:
return ""


def reject_image_attachments_in_compacted_mode(
params: ResponsesApiParams,
image_attachments: Optional[Sequence[Any]],
) -> None:
"""Reject a compacted turn that carries image attachments (LCORE-3582).

In compacted mode the wire ``input`` is overridden with the explicit item
list built by :func:`_build_explicit_input`, which is text-only: image
attachments are converted to pydantic-ai ``ImageUrl`` parts on the prompt,
and the override replaces the prompt-derived input wholesale, so those
parts never reach the request body.

Answering anyway would return a confident response that never saw the
image, with nothing to tell the caller their attachment was ignored. Fail
explicitly instead until the explicit input can carry ``input_image``
content parts of its own (LCORE-3789).

Args:
params: Prepared (possibly compaction-rewritten) request parameters.
image_attachments: Image attachments for this turn, if any.

Raises:
HTTPException: 422 when the turn is compacted and carries images.
"""
if not image_attachments or not params.omit_conversation:
return
logger.warning(
"Rejecting compacted turn with %d image attachment(s): the explicit "
"input override cannot carry image content parts (LCORE-3789)",
len(image_attachments),
)
response = UnprocessableEntityResponse(
response="Image attachments are not supported on this conversation",
cause=(
"This conversation has been compacted to fit the model's context "
"window, and compacted turns cannot carry image attachments yet. "
"Send the image in a new conversation, or retry without it."
),
)
raise HTTPException(**response.model_dump())


def _query_input_message(original_input: ResponseInput) -> list[Any]:
"""Render the new user query as explicit input items.

Expand Down
Loading
Loading