Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
116 changes: 112 additions & 4 deletions docs/design/conversation-compaction/conversation-compaction.md
Original file line number Diff line number Diff line change
Expand Up @@ -311,10 +311,11 @@ Add `compaction` field to the root `Configuration` class.
|----------------------------------------|----------------------------------------------------------------------------|
| `pyproject.toml` | Add `tiktoken` dependency |
| `src/utils/token_estimator.py` | New module: `estimate_tokens()`, `estimate_conversation_tokens()` |
| `src/utils/compaction.py` | New module: summarization logic, partitioning, additive summary management |
| `src/utils/compaction.py` | New module: summarization logic, partitioning, additive summary management; reports the usage of each LLM call it makes — LCORE-3910 |
| `src/models/config.py` | Add `CompactionConfiguration` (near `ConversationHistoryConfiguration`) |
| `src/configuration.py` | Add `compaction_configuration` property to `AppConfig` singleton |
| `src/utils/conversation_compaction.py` | New module: `apply_compaction()` / `apply_compaction_blocking()`, `needs_compaction_path()`, marker helpers, per-conversation lock |
| `src/utils/pending_turn.py` | `PendingTurn`, the owner of every turn the endpoints store themselves: decides whether the turn is ours, stores it against the input as it arrived, and stores it once — LCORE-3908 |
| `src/models/common/responses/responses_api_params.py` | `omit_conversation` flag — drops the `conversation` parameter from the request body in compacted mode |
| `src/app/endpoints/query.py` | Call `apply_compaction_blocking()` after preparing params; store the turn in compacted mode |
| `src/app/endpoints/streaming_query.py` | Compaction-aware SSE path that emits the `compaction` event before summarizing (R12) |
Expand All @@ -331,9 +332,11 @@ reusable unit in `src/utils/conversation_compaction.py` that each endpoint
calls after its params are prepared:

- `apply_compaction_blocking(client, params, inference_config, compaction_config)`
returns a `CompactionResult` (possibly-rewritten params, a `summarized`
flag, and the `original_input`). Non-streaming `/v1/query`, A2A, and
`/v1/responses` use this.
returns a `CompactionResult` (possibly-rewritten params, a `compacted`
flag, the `original_input`, and the `summarization_usage`). Non-streaming
`/v1/query`, A2A, and `/v1/responses` use this. The endpoints also pass
`endpoint_path` and, where there is a quota, `charge`; see
[Who pays for the summarization calls](#who-pays-for-the-summarization-calls).
- `apply_compaction(..., emit_events=True)` is the async-generator variant that
yields a `CompactionStartedEvent` before the summarization LLM call; the
native `/v1/streaming_query` SSE path uses it to satisfy R12.
Expand All @@ -346,6 +349,67 @@ When compaction is active, the endpoint builds explicit input, the
`conversation` parameter is omitted (via `ResponsesApiParams.omit_conversation`),
and the completed turn is appended to the conversation items afterward.

That append has one owner, `PendingTurn` in `src/utils/pending_turn.py`
(LCORE-3908). An endpoint creates it from the parameters the request is sent
with and the `original_input` of the `CompactionResult`, and reports how the
turn ended: `store_completed`, `store_blocked`, `store_interrupted` or `drop`.
The first report settles the turn and every later one does nothing, so a turn is
stored once even when several paths of a request want to store it (the end of a
stream, the cancellation handler, the interrupt callback). `PendingTurn` also
covers the other turns OGX does not store: a request a shield blocked, an
interrupted stream, a continuation from `previous_response_id`.

A compacted request that loses its turn through a change in the code fails.
Building a `PendingTurn` for compacted parameters without the original input
raises `ValueError`, and `ensure_settled()` raises `TurnNotStoredError` when a
compacted request reaches the end of its handler and nobody tried to store its
turn or dropped it on purpose. `/v1/query` runs inside the `pending_turn()`
scope, which makes that check when it is left; the streaming paths,
`/v1/responses` and A2A make it after their write. Before this, a lost write
was silent: the conversation stopped growing and nothing failed (LCORE-3883).

The check does not cover three endings, which store nothing and are rows of the
table below: a stream the client stops reading, a `/v1/responses` stream without
a final response, and a failed write where the failure is logged.

The shield capabilities (question validity, Granite Guardian) are the one writer
outside the owner. They store the turn they rejected from inside the agent run,
and only when the model was handed the conversation, so never in compacted mode.

What is stored, per endpoint and per way a turn can end. "OGX" means the
`conversation` parameter was sent and OGX stores the turn itself. The last
column says what a failed write does to the request.

| Endpoint | Turn ended | Not compacted | Compacted | Failed write |
|---|---|---|---|---|
| `/v1/query` | completed | OGX | stored | request fails |
| | blocked by a shield | stored | not stored (LCORE-3788) | request fails |
| | model call failed | not stored | not stored | |
| | run did not finish with success | OGX | not stored | |
| `/v1/streaming_query` | completed, also when the run did not finish with success | OGX | stored when the stream ends | logged |
| | blocked by a shield | stored before the stream starts | stored when the stream ends | request fails / logged |
| | interrupted by the client | stored, with the answer so far | stored, with the answer so far | logged |
| | blocked by a shield, then interrupted | the refusal, stored once before the stream; the interrupt adds nothing | stored, with the answer so far | logged |
| | client stopped reading | OGX | not stored | |
| `/v1/responses` | completed, incomplete or failed | OGX; stored when continuing from `previous_response_id` | stored | request fails; a stream ends before `[DONE]` |
| | blocked by a shield | stored | stored | request fails |
| | stream without a final response | not stored | not stored | |
| | `store: false` | not stored | never compacted | |
| A2A | completed | OGX | stored | logged |
| | agent run failed | not stored | not stored | |

`tests/integration/endpoints/test_turn_persistence.py` pins the table. It runs
the real handlers and compares the conversation item by item after the request.
The compacted column is covered row by row; the other column for the rows where
lightspeed-stack stores the turn, and for a completed turn on each endpoint. The
last column is covered for a completed turn on each endpoint, for a blocked and
for an interrupted stream.

The row "blocked by a shield, then interrupted" is the one place where this
change alters what is stored. The interrupt used to store the turn a second
time, with the interruption notice for an answer. It still records the turn in
the database.

## Fetching conversation history

Use the same pattern as `conversations_v1.py:240-246`:
Expand Down Expand Up @@ -385,6 +449,50 @@ Example config files go in `examples/`.

Compaction adds latency only on the trigger turn. In PoC testing, compaction turns took 14-40 seconds vs 9-20 seconds for normal turns (gpt-4o-mini).

## Who pays for the summarization calls

The provider bills the summarization call and the fold call like any other, so
lightspeed-stack counts them (LCORE-3910). Before that, their usage was
discarded: a user could go over the quota without it ever showing, and by more
the longer the conversation was.

`summarize_chunk()` and `recursively_resummarize()` report each call they make
to a `count_call` callback, with the model and the usage the provider reported.
They do it as soon as the response arrived and before they look at it, so a
call that returned no text is reported too. `apply_compaction()` takes the
path of the endpoint and a `charge` callback, and counts the calls with
`SummarizationCalls` (`src/utils/compaction_usage.py`): for each call it
records the token and call metrics under that endpoint, calls `charge`, and
adds the usage up. The sum is handed to the endpoint as
`CompactionResult.summarization_usage`.

| Endpoint | Quota | Counts the client sees |
|------------------------|---------|----------------------------------------------------------|
| `/v1/query` | charged | `input_tokens` / `output_tokens` include summarization |
| `/v1/streaming_query` | charged | the same, in the `end` event |
| `/v1/responses` | charged | `usage` as the provider reported it, the answer alone |
| `/a2a` | none | none; the endpoint has no quota, the metrics are recorded |

The endpoints with a quota pass `consume_summarization_tokens()`
(`src/utils/query.py`), bound to the user, as `charge`. A call is therefore
charged when it returned, and not together with the turn: the summary is
written before the model is asked for the answer and is kept whatever becomes
of the turn, so the call is a cost of the conversation. It is charged also
when the turn is blocked, fails or is interrupted, and when compaction itself
fails after the call (the marker cannot be written, the fold fails). `/a2a`
passes no `charge`.

`/v1/query` and `/v1/streaming_query` add `summarization_usage` to the usage of
the turn in what they report to the client and set on the request span
(`llm.usage.input_tokens`, `llm.usage.output_tokens`). The sum is not stored
with the turn; the token usage history receives the charges one by one.
`/v1/responses` passes the `usage` object of the response through unchanged, to
stay a drop-in for clients of the OpenAI Responses API.

Two limits. A call that fails, or is cancelled before its response arrived,
has no usage to count, whatever the provider bills for it. And a call the
provider reports no usage for is counted as a call and charges nothing.

# Open Questions for Future Work

- **Compaction-proof instructions**: Allow "pinned" messages that always survive compaction (inspired by Claude Code's CLAUDE.md pattern). Not needed for v1.
Expand Down
3 changes: 2 additions & 1 deletion docs/design/prompt-guardrails/prompt-guardrails.md
Original file line number Diff line number Diff line change
Expand Up @@ -442,7 +442,8 @@ product need justifies it.
version-specific prompt.
- **Guardian token usage:** whether guardian calls should count against user
quota or be tracked as service overhead. Compaction's summarization calls
raise the same question.
are charged to the user's quota (LCORE-3910); the same choice is open for
guardian calls.
- **Streaming checkpoint sizing:** defaults for LCORE-3391 (spike Decision
T4, 70% confidence); tune with real latency data.
- **Cheap classifier tier for `tool`:** Prompt Guard 2-class, and its
Expand Down
8 changes: 4 additions & 4 deletions docs/devel_doc/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -274,9 +274,9 @@ The system defines 30+ actions that can be authorized. Examples (see `docs/user_
- Update quota counters

3. **On Error:**
- If LLM call fails, no tokens are consumed
- Quota remains unchanged
- User can retry the request
- If LLM call fails, no tokens are consumed for that call
- Quota remains unchanged, with one exception: the calls conversation compaction made to summarize older turns for the request were charged when they were made
- User can retry the request; the summary is kept, so the retry does not pay for it again

---

Expand Down Expand Up @@ -391,7 +391,7 @@ The compaction system is split into two layers:
- Marker persistence (`[lightspeed:compaction-summary]` sentinel in conversation items)
- `CompactionStartedEvent` emission for streaming progress indicators
- `apply_compaction()` (async generator) — Main entry point used by all endpoints
- `store_compacted_turn()` — Appends user query + LLM output when in compacted mode
- `PendingTurn` (`utils/pending_turn.py`) — Owns the turns the endpoints store themselves: appends user query + LLM output when in compacted mode, once per request

**Data Flow:**

Expand Down
5 changes: 3 additions & 2 deletions docs/devel_doc/query_endpoint.md
Original file line number Diff line number Diff line change
Expand Up @@ -308,7 +308,7 @@ Cancels an in-progress streaming query.
| `interrupted` | boolean | Whether an active stream was interrupted (`false` if already completed) |
| `message` | string | Human-readable status message |

When a stream is interrupted, any partial response is persisted to conversation history and token consumption is skipped.
When a stream is interrupted, any partial response is persisted to conversation history and token consumption for the answer is skipped. The summarization calls conversation compaction made for the request were charged when they were made (see [Quota and Token Counting](#quota-and-token-counting)). A request a shield had blocked is the exception: its refusal turn is stored before the stream starts, so an interrupt adds nothing to the conversation.

---

Expand Down Expand Up @@ -361,8 +361,9 @@ If the server configuration sets `disable_query_system_prompt` to `true`, reques

- **Pre-request:** `check_tokens_available()` verifies the user/cluster has available quota (429 if not)
- **Post-response:** `consume_query_tokens()` deducts `input_tokens` and `output_tokens` from configured quota limiters
- **Compaction:** each LLM call that summarizes older turns is charged by `consume_summarization_tokens()` as soon as it returned, so also when the turn is then blocked, fails or is interrupted. The `input_tokens` and `output_tokens` the response reports include these calls
- **Available quotas:** Remaining balances per limiter are included in the response (`available_quotas` field in sync, `end` event in streaming)
- **Stream interruption:** Token consumption is skipped for interrupted streams
- **Stream interruption:** Token consumption is skipped for interrupted streams, except for the summarization calls, which were charged before the answer started

---

Expand Down
37 changes: 36 additions & 1 deletion docs/user_doc/conversation_compaction.md
Original file line number Diff line number Diff line change
Expand Up @@ -144,6 +144,37 @@ Compaction acquires a per-conversation lock to prevent concurrent requests on th

Over very long conversations, multiple compaction summaries may accumulate. When the total size of cached summaries approaches the context window threshold, they are recursively folded into a single summary using a dedicated re-summarization prompt. This prevents summaries from themselves exceeding the context window.

### Token usage and quota

Summarizing older turns takes an LLM call of its own, and so does folding the
summaries. The provider bills these calls, so they are counted:

| Endpoint | Quota | Reported token counts | Metrics |
|---|---|---|---|
| `POST /v1/query` | Charged to the user. | `input_tokens` and `output_tokens` include the summarization calls. | Counted under `/v1/query`. |
| `POST /v1/streaming_query` | Charged to the user. | `input_tokens` and `output_tokens` of the `end` event include the summarization calls. | Counted under `/v1/streaming_query`. |
| `POST /v1/responses` | Charged to the user. | `usage` is what the provider reported for the response, so it covers the answer alone. `available_quotas` shows the quota left after both. | Counted under `/v1/responses`. |
| `POST /a2a` | Not charged. The endpoint does not use quotas. | None. | The summarization calls are counted under `/a2a`. The call that answers the request is not recorded in these metrics on this endpoint. |

The metrics are `ls_llm_calls_total`, `ls_llm_token_sent_total` and
`ls_llm_token_received_total`, labelled with the endpoint.

Only a request that makes a summarization call pays for it: one whose
estimated input crosses the threshold. The other requests are served from the
stored summary and cost what they cost without compaction. With the default
settings a summarization happens once every several turns; with a small
context window or a low `threshold_ratio` it can happen on every request.

Each summarization call is charged as soon as it returned, before the model is
asked for the answer. It is therefore charged also when the request is blocked
by a shield, fails afterwards, or is interrupted by the client. The summary is
kept in these cases, so the next request on the conversation does not pay for
it again. A summarization call that itself fails, or is cut off before it
returned, reports no usage and is not charged.

On `/v1/responses`, the difference between the quota consumed and the `usage`
of the response is the summarization.

## When compaction is disabled

When compaction is disabled (the default), requests that cause the conversation history to exceed the model's context window will fail with HTTP 413 (Prompt Too Long). Clients must manage conversation length themselves, for example by starting new conversations or deleting old ones.
Expand All @@ -158,7 +189,11 @@ Compaction summarizes older turns, so fine-grained details from early in the con

**Does compaction use extra tokens?**

Yes. The summarization step requires an additional LLM call, which consumes tokens. These tokens are counted against the user's quota. The trade-off is that the conversation can continue instead of failing with HTTP 413.
Yes. The summarization step requires an additional LLM call, which consumes tokens. On `/v1/query` and `/v1/streaming_query` these tokens are counted against the user's quota and are included in the `input_tokens` and `output_tokens` of the request that triggered the summarization. `/v1/responses` charges the quota and leaves `usage` as the provider reported it; `/a2a` has no quota. See [Token usage and quota](#token-usage-and-quota). The trade-off is that the conversation can continue instead of failing with HTTP 413.

**Why did one request use many more tokens than the ones before it?**

On `/v1/query` and `/v1/streaming_query`, that request triggered a summarization. Its `input_tokens` include the older turns that were sent to the LLM to be summarized, and its `output_tokens` include the summary. `context_status` is `"summarized"` on that request, and also on the requests after it that are served from the stored summary and do not pay for it again.

**Can I use compaction with all LLM providers?**

Expand Down
Loading
Loading