Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
e2c88e6
feat(python): separate Foundry Responses agent history and background
eavanvalkenburg Sep 25, 2026
be003ff
docs(python): clarify Foundry provider background recovery
eavanvalkenburg Sep 28, 2026
57e2628
fix(python): stabilize Responses hosting CI checks
eavanvalkenburg Sep 28, 2026
4664826
fix(python): protect hosted Responses provider state and options
eavanvalkenburg Sep 28, 2026
e81632c
fix(python): gate unsafe Foundry Responses steering until SDK fix
eavanvalkenburg Sep 28, 2026
f79162b
fix(python): preserve provider finish reasons in Responses updates
eavanvalkenburg Sep 28, 2026
b2af0b8
docs(python): explain Foundry Responses history ownership by mode
eavanvalkenburg Sep 28, 2026
f45d3b0
refactor(python): name Foundry Responses history and background sources
eavanvalkenburg Sep 28, 2026
4f704f0
Merge remote-tracking branch 'upstream/main' into foundry-responses-a…
eavanvalkenburg Sep 28, 2026
bf94a54
Python: Protect service conversations and provider background polls
eavanvalkenburg Sep 29, 2026
25be8e8
Merge remote-tracking branch 'upstream/main' into foundry-responses-a…
eavanvalkenburg Sep 29, 2026
51e7a3d
Merge remote-tracking branch 'upstream/main' into foundry-responses-a…
eavanvalkenburg Sep 29, 2026
8857366
Merge remote-tracking branch 'upstream/main' into foundry-responses-a…
eavanvalkenburg Sep 29, 2026
8ffdbb3
Python: Preserve provider background output through recovery
eavanvalkenburg Sep 30, 2026
7613e87
Merge remote-tracking branch 'upstream/main' into foundry-responses-a…
eavanvalkenburg Sep 30, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
161 changes: 113 additions & 48 deletions python/packages/foundry_hosting/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,28 +29,105 @@ The Responses host continues regular agents through its existing session store a
existing checkpoint store. A callable does not make arbitrary instance fields persistent; state needed by later
requests must remain in the supported stores.

## Conversation history

`ResponsesHostServer` uses AgentServer response history as the model's conversation history by default:
## Responses agent history and storage

The caller's `POST /responses` **`store` flag** controls whether the *outer* response is retrievable and whether
MAF session and approval state is saved. The host's `history_source` independently selects who supplies model history:

| `history_source` | What the model receives | Inner storage on caller `store=True` |
| --- | --- | --- |
| `"agent_server"` (default) | Prior outer Responses transcript plus new input | Disabled. Storing clients run with `store=False`; non-storing clients receive no storage option. |
| `"service"` | New input only | Enabled. The service-issued `AgentSession.service_session_id` is saved privately under the outer response ID or conversation. |
| `"agent"` | New input only; the agent chooses how to load history | Developer-owned: `HistoryProvider` with `default_options={"store": False}` loads from host-persisted `AgentSession.state`, **or** a storing client with `default_options={"store": True}` uses downstream service history. The existing behavior is unchanged. |

For example, suppose the first stored response answers **"My name is Ada"**, then the caller sends
**"What is my name?"** with `previous_response_id` set to that response's **outer** `response.id`:

- **`"agent_server"`:** The model receives the first user input, the first assistant output, and the new
question. The host reconstructs that transcript from the Responses store; the inner service
does not retain it.
- **`"service"`:** The model receives only the new question as *request input*, along with the
private `AgentSession.service_session_id` from the first turn. The downstream service retrieves
its own transcript. A second branch from the first response cannot safely reuse that service
thread and is rejected.
- **`"agent"`:** The model receives the new question. With `InMemoryHistoryProvider` and agent
default `store=False`, the provider adds earlier messages from the saved MAF session; with agent
default `store=True`, the downstream service owns the prior transcript instead.

All three still return **outer** Responses IDs for retrieval and background polling. `store=False`
requests are one-shot: they do not write host-managed state or ask the inner client to store, so
they cannot establish a persistent provider thread. Neither an outer `response.id` nor
`agent_session_id` should be used as an inner `service_session_id`.

Choose **one** mode when constructing each host; do not reuse the same `Agent` instance across hosts.
For example, to use downstream service history:

```python
server = ResponsesHostServer(agent)
server = ResponsesHostServer(agent=agent, history_source="service")
```

In this mode, the configured AgentServer response provider supplies the prior transcript. Hosting rejects
`HistoryProvider` instances with `load_messages=True` and agents configured with a default `conversation_id`,
`previous_response_id`, or `conversation`, adds a transient in-memory provider for function-call loops, and clears
restored downstream service IDs. For clients that advertise `STORES_BY_DEFAULT=True`, hosting forces downstream
`store=False`; for other clients it removes an explicit agent-level `store` option and does not forward one. These
safeguards ensure the model receives the transcript once without sending unsupported storage options.

AgentServer history requires a framework `RawAgent` whose client declares the boolean `STORES_BY_DEFAULT` capability;
the agent's runtime options then let hosting enforce downstream storage behavior. Custom `SupportsAgentRun`
implementations must use `history_source="agent"` because that protocol does not accept runtime chat options.

`ResponsesHostServer` owns a supplied agent instance and may add hosting-specific context providers. Do not reuse that
instance with another host or invoke it directly after constructing the server. An agent returned by a callable belongs
to that request.
Omitting `history_source` selects `"agent_server"`. To use a provider in `"agent"` mode,
configure that agent with `store=False` as shown in
[agent_history.py](../../samples/04-hosting/foundry-hosted-agents/responses/basic/agent_history.py).

`"agent_server"` and `"service"` reject a load-enabled `HistoryProvider` alongside their own history source;
`"agent_server"` also rejects default downstream continuation IDs. Those modes require a `RawAgent` with a client declaring
`STORES_BY_DEFAULT`; `"service"` requires a storing client that returns a private continuation ID. Hosting may
add a transient in-memory provider to support function-call loops, but **never edits `agent.default_options`**.
The host owns the provided agent instance and any providers it adds; do not reuse it with another host. A factory
creates an independent agent for each request.

The existing `history_source="agent"` still preserves the agent's own provider **or** service storage defaults
on stored requests, including `default_options={"store": True}`. Custom `SupportsAgentRun`
implementations can continue using that mode for stored requests; it is not a forced-provider mode.
The outer storage-backend constructor argument is now `response_store=`. The old `store=` backend
argument remains an alias with its own once-per-host deprecation warning; supplying both is an error.
Neither constructor argument sets the caller's per-request `store` flag.

`store=False` returns a one-shot response without **writing** host-managed session, conversation, or approval state;
it also disables downstream service storage regardless of the developer's defaults. Unsafe custom agents, external
history providers that store messages, and fixed downstream continuation defaults fail with an actionable error
instead of silently persisting. An unstored service-mode request cannot resume a private service thread.
`history_source="agent"` also rejects an unstored continuation if its restored session uses downstream storage. Application-owned
tools and external services may still have their own side effects. `background=True` requires outer `store=True`.

Outer background work always uses the caller-visible `response.id` for polling; it does not enable provider-native
background automatically. `background_source="agent_server"` (default) uses only the outer background worker.
`background_source="provider"` is a separate opt-in for `history_source="service"` with a storing
Responses client. Its private continuation token is saved under the outer ID and never returned to the caller.
Use `ResponsesServerOptions(resilient_background=True)` to permit recovery from a **saved** token; a crash before
the token is saved cannot safely restart the inner job. Completed polling output, including local function calls,
results, and usage, is saved together with the next token in the private response-ID snapshot before emission.
Outer output checkpoints record which saved batches have been emitted, so recovery restores their usage and
replays only uncheckpointed output. Once the final output is saved, recovery can finish from that snapshot without
calling the provider again; a later turn drops it from its working session. Shutdown during initial submission
fails rather than replaying a job whose acceptance is unknown. Cancelling an in-flight submission does not prove
the remote provider stopped it.
Each poll retains the caller's generation options and `background=True`, so a tool-loop follow-up requests
another background response and saves its next token. A crash after a local tool side effect but before that
next token is saved can still repeat the tool on recovery; use idempotent tools or avoid provider background
for side-effecting local tools. This mechanism does not provide exactly-once tool execution.
Provider background and steering cannot be combined.
Regular agent runs without this opt-in are not crash-replayable. **Steering is temporarily unavailable:**
`steerable_conversations=True` fails during host construction, before enabling the process-wide TaskManager.
The current AgentServer SDK retains unbounded futures for rejected turns when its steering queue fills. Do not
use a queue-length precheck: another worker can append before it. The guard can be removed only after
[Azure/azure-sdk-for-python#49233](https://github.com/Azure/azure-sdk-for-python/pull/49233) ships in an
official `azure-ai-agentserver-core` wheel, the minimum dependency and `uv.lock` are updated, and a
concurrent queue-overflow regression proves rejected turns leave no pending futures. No future SDK
version is assumed. Non-steerable background polling and legacy `WorkflowAgent` dispatch are unchanged.

Native CreateResponse generation fields become MAF runtime options (notably `max_output_tokens` -> `max_tokens` and
`parallel_tool_calls` -> `allow_multiple_tool_calls`). Flattened OpenAI `extra_body` fields overlay translated keys
**last**. A sync or async `prepare_options(request: HostedResponseRequest, options: dict)` hook can remove or replace
*caller* options before `Agent.run`; removed values fall back to the developer's unchanged agent defaults. Hosting
filters caller platform IDs and private continuation/storage controls from model options and rejects attempts to
reintroduce them through the hook. Nested `extra_body` transport overrides are rejected for both caller
input and developer hooks, because they could override the host's `store=False` after the OpenAI SDK merges
the body. Developer defaults also cannot use this transport channel for host-controlled fields on
explicit history modes or unstored requests. A custom agent cannot accept MAF runtime options: choose
`unsupported_options` as `"ignore"`, `"warn"` (default), or `"error"` for that case. See the
[agent history and options samples](../../samples/04-hosting/foundry-hosted-agents/responses/basic/).

### OAuth consent origin allowlist

Expand All @@ -71,30 +148,6 @@ An omitted allowlist preserves existing behavior and does not restrict the HTTPS
the gate, so an empty sequence rejects every consent link. Entries are normalized as origins, so paths and query strings
belong on the emitted consent link, not in the configuration.

To preserve the agent's regular history and service-storage behavior, select the agent as the history source:

```python
server = ResponsesHostServer(agent, history_source="agent")
```

Hosting then passes only current request input, allows load-enabled history providers, and does not override the
agent's downstream `store` option. For example, `InMemoryHistoryProvider` stores messages in `AgentSession.state`, which
the default `FoundryAgentSessionStore` persists in Foundry:

```python
agent = Agent(
client=client,
context_providers=[InMemoryHistoryProvider()],
default_options={"store": False},
)
server = ResponsesHostServer(agent, history_source="agent")
```

The `store` argument remains independent: it selects the AgentServer response provider used for Responses API
persistence and retrieval. Omitting it or passing `None` selects the environment default. With
`history_source="agent_server"`, that response provider also supplies model history; with `history_source="agent"`, it
does not.

## Computer use

The Responses host emits native `computer_call` and `computer_call_output` items, including ordered `actions`,
Expand Down Expand Up @@ -183,12 +236,24 @@ durably. By default they use `FoundryAgentSessionStore`, backed by Foundry stora
and file-based storage locally. Responses sessions use the `agent_sessions` logical store;
Invocations sessions use the separate `invocation_sessions` store.

Loaded MAF sessions are saved with an ETag condition. A competing turn that has
already advanced the same conversation causes a visible persistence failure instead
of silently overwriting its state. New hosted session keys are created only if absent;
turns using `previous_response_id` write their own new response ID, without applying
the predecessor's ETag to a different key. Local callers can still upsert directly
without first loading a session.
Each stored agent turn saves a snapshot under its **own** outer `response.id`. For a named `conversation`, a
separate mutable conversation-head key is also updated. Loaded MAF sessions use PR1's ETag condition for that key:
a competing turn that advanced the head causes a visible conflict rather than a stale overwrite. A superseded
steered turn saves its response snapshot but skips the head update. In `"service"` or `"agent"` history mode, a
stored named-conversation turn first claims that head with a conditional write *before* calling the inner agent.
Another request that read the old head loses the CAS; one that reads the claim fails before touching the provider.
The claim is cleared when the winning turn successfully commits the new head. If a dispatched turn fails or is
cancelled, the claim remains: the provider may already have changed its thread, so start a new conversation
instead of retrying this one blindly. A recovered provider-background turn must still own the same claim.
The committed head also records its completing outer response ID. If a crash occurs after the head write but before
outer completion, that response can recover its saved final output without reclaiming or rewriting the head.
A different in-flight claim or completing response is not accepted as ownership.
When continuing by `previous_response_id`, the prior response is claimed with a conditional write **after**
the input is validated, so an invalid approval response does not consume a usable parent. A second branch
cannot reuse the same downstream service thread; attempting to fork a named service conversation is also rejected.
New hosted keys are created only if absent. Custom store providers must provide equivalent scoped conditional
writes for concurrent turns, including the pre-dispatch claim. Local callers can still upsert directly without
first loading a session.

See the [custom storage provider sample](../../samples/04-hosting/foundry-hosted-agents/responses/custom_storage/)
for an example that uses an in-memory session store locally and Azure Cosmos DB when hosted.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@

if TYPE_CHECKING:
from ._invocations import InvocationsHostServer
from ._request import HostedResponseRequest
from ._responses import ResponsesHostServer
from ._scope import FoundryRequestScope
from ._state_store import (
Expand Down Expand Up @@ -34,6 +35,7 @@
"FoundryFunctionApprovalStore": "._state_store",
"FoundryRequestScope": "._scope",
"FoundryToolbox": "._toolbox",
"HostedResponseRequest": "._request",
"FunctionApprovalStore": "._state_store",
"FunctionApprovalStoreProvider": "._state_store",
"InvocationsHostServer": "._invocations",
Expand All @@ -52,6 +54,7 @@
"FoundryToolbox",
"FunctionApprovalStore",
"FunctionApprovalStoreProvider",
"HostedResponseRequest",
"InvocationsHostServer",
"ResponsesHostServer",
"StoreProvider",
Expand Down
Loading
Loading