diff --git a/crates/adaptive/README.md b/crates/adaptive/README.md
index 7dd1a1ced..540f35727 100644
--- a/crates/adaptive/README.md
+++ b/crates/adaptive/README.md
@@ -18,7 +18,8 @@ SPDX-License-Identifier: Apache-2.0
`nemo-relay-adaptive` is the Rust companion crate for adaptive NeMo Relay
runtime behavior. Use it with `nemo-relay` when an agent runtime should learn
-from observed executions, inject runtime hints, or persist adaptive state.
+from observed executions, inject runtime hints, persist adaptive state, or
+cache repeated LLM responses.
Adaptive behavior is installed through the same plugin system used by the core
runtime, so applications can enable it without changing their orchestration
@@ -41,9 +42,11 @@ framework.
- **`AdaptiveConfig`**: A canonical config contract for the top-level
`adaptive` plugin component.
- **Built-in component settings**: Typed config helpers for telemetry,
- adaptive hints, tool parallelism, and the Adaptive Cache Governor.
-- **State backends**: In-memory state by default and Redis-backed state behind
- the `redis-backend` feature.
+ adaptive hints, tool parallelism, the Adaptive Cache Governor, and response
+ caching.
+- **Storage backends**: In-memory backends for adaptive state and opt-in
+ response caching, with Redis-backed storage behind the `redis-backend`
+ feature.
- **Learning primitives**: Runtime helpers and learners built on NeMo Relay
events.
- **Adaptive Cache Governor (ACG) module surface**: The canonical
@@ -58,7 +61,8 @@ Install the published crate alongside the core runtime:
cargo add nemo-relay nemo-relay-adaptive
```
-Enable Redis-backed state only when the application needs shared persistence:
+Enable Redis-backed storage when adaptive state or the response cache needs
+shared persistence:
```bash
cargo add nemo-relay-adaptive --features redis-backend
diff --git a/docs/configure-plugins/adaptive/about.mdx b/docs/configure-plugins/adaptive/about.mdx
index e6570839d..3bdf9b30c 100644
--- a/docs/configure-plugins/adaptive/about.mdx
+++ b/docs/configure-plugins/adaptive/about.mdx
@@ -22,6 +22,7 @@ The plugin can coordinate:
- Adaptive hints injected into outgoing model requests.
- Tool-parallelism observation or scheduling behavior.
- Adaptive Cache Governor (ACG) prompt-cache planning.
+- Opt-in response caching for repeated LLM calls.
- Component-local validation policy.
## Use Adaptive When
@@ -51,6 +52,8 @@ If instrumentation is not in place yet, start with
cache planning accomplishes.
- [Adaptive Hints](/configure-plugins/adaptive/adaptive-hints) explains request hint injection and how
downstream model paths can consume the hints.
+- [Response Cache](/configure-plugins/adaptive/response-cache) explains the opt-in LLM response cache:
+ turning it on, what gets cached, and how savings are reported.
State, telemetry, tool parallelism, and policy are whole-plugin configuration
areas. They are documented on [Adaptive Configuration](/configure-plugins/adaptive/configuration) rather
diff --git a/docs/configure-plugins/adaptive/configuration.mdx b/docs/configure-plugins/adaptive/configuration.mdx
index 570729b34..d85391b2e 100644
--- a/docs/configure-plugins/adaptive/configuration.mdx
+++ b/docs/configure-plugins/adaptive/configuration.mdx
@@ -1,7 +1,7 @@
---
title: "Adaptive Configuration"
sidebar-title: "Configuration"
-description: "Configure the built-in Adaptive component, including state, telemetry, hints, and ACG."
+description: "Configure the Adaptive component, including state, telemetry, hints, ACG, and response caching."
position: 2
---
{/* SPDX-FileCopyrightText: Copyright (c) 2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved.
@@ -11,9 +11,10 @@ SPDX-License-Identifier: Apache-2.0 */}
Use this page when you want to configure the built-in Adaptive plugin component
as a whole. The component kind is `adaptive`.
-Adaptive plugin configuration uses the generic NeMo Relay plugin document shape.
-Field names stay `snake_case` in every binding and in `plugins.toml`, even when
-language helper functions use language-native naming conventions.
+Adaptive plugin configuration uses the generic NeMo Relay plugin document
+shape. Canonical plugin documents and `plugins.toml` use `snake_case`.
+Binding helpers can use language-native names and serialize them to canonical
+keys.
For plugin file discovery, precedence, merge behavior, editor controls, and
gateway conflict rules, refer to
@@ -32,10 +33,12 @@ The top-level adaptive object contains:
| `adaptive_hints` | Request hint-injection behavior. |
| `tool_parallelism` | Tool scheduling observation or scheduling behavior. |
| `acg` | Adaptive Cache Governor prompt-cache planning. |
+| `response_cache` | Opt-in LLM response cache for repeated managed calls. Requires a non-empty trust-domain `namespace`. |
| `policy` | Adaptive-local handling for unknown fields and unsupported values. |
-The requested area pages cover [Adaptive Cache Governor (ACG)](/configure-plugins/adaptive/acg) and
-[Adaptive Hints](/configure-plugins/adaptive/adaptive-hints). State, telemetry, tool parallelism, and
+Dedicated pages cover [Adaptive Cache Governor (ACG)](/configure-plugins/adaptive/acg),
+[Adaptive Hints](/configure-plugins/adaptive/adaptive-hints), and
+[Response Cache](/configure-plugins/adaptive/response-cache). State, telemetry, tool parallelism, and
policy remain whole-plugin settings:
- Use `state.backend.kind = "in_memory"` for local experiments.
@@ -271,9 +274,9 @@ finally:
-The Node.js binding exposes the built-in adaptive runtime through the adaptive
-plugin component helpers. Use the Plugin Configuration example above when
-activating adaptive behavior from Node.js.
+The Node.js binding exposes `adaptive.AdaptiveRuntime` for manual lifecycle
+ownership. Construct it from the same `adaptiveConfig`, then call `register()`,
+`waitForIdle()`, and `shutdown()`.
diff --git a/docs/configure-plugins/adaptive/response-cache.mdx b/docs/configure-plugins/adaptive/response-cache.mdx
new file mode 100644
index 000000000..b9299bc0f
--- /dev/null
+++ b/docs/configure-plugins/adaptive/response-cache.mdx
@@ -0,0 +1,427 @@
+---
+title: "Response Cache"
+description: "Configure exact-match response caching for managed LLM calls."
+position: 5
+---
+{/* SPDX-FileCopyrightText: Copyright (c) 2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved.
+SPDX-License-Identifier: Apache-2.0 */}
+
+
+Use the response cache when the same LLM request is made more than once and the
+repeat should be served from a store instead of calling the provider again. An
+eligible request that matches an unexpired cache entry can be served from the
+store without calling the provider. Buffered hits preserve the stored response
+shape and usage fields; streaming hits replay an equivalent provider-native
+stream.
+
+The cache is an optional `response_cache` section of the
+[Adaptive plugin](/configure-plugins/adaptive/configuration), not a standalone plugin
+kind. It is off until the section is present, applies to
+[managed LLM calls](/instrument-applications/instrument-llm-call) without
+changing the execution API. By default, only requests with an explicit numeric
+`temperature = 0` are eligible; set `cache_nondeterministic = true` to opt
+sampled requests into caching. Runtime backend errors fail open to a normal
+live call, while invalid configuration is rejected during validation.
+
+`namespace` is required and defines one trusted cache-sharing domain. Do not
+use one namespace across mutually untrusted tenants or upstreams.
+`header_allowlist` does not replace this trust boundary.
+
+## When to Use It
+
+- Development and test loops that replay the same prompts — eligible repeats
+ can reuse stored responses.
+- CI suites that exercise real models — cached repeats can avoid additional
+ provider calls.
+- Demos and workshops — reuse stable stored answers while reducing provider
+ calls.
+- A shared team cache — callers in one trusted sharing domain can use the same
+ Redis store, namespace, and key prefix across processes and machines.
+- Gateway deployments — enable caching for an agent without touching its code.
+
+Reuse requires the request to match exactly after normalization, so prompts
+that embed volatile content (timestamps, random IDs) will not repeat. The TTL
+is the maximum entry age, not guaranteed residency; a backend can evict or
+delete an entry sooner. Use `bypass_rate` to re-run a sample of cacheable calls
+live and refresh an entry when the new response is stored.
+
+## `plugins.toml` Example
+
+```toml
+version = 1
+
+[[components]]
+kind = "adaptive"
+enabled = true
+
+[components.config]
+version = 1
+
+[components.config.response_cache]
+ttl_seconds = 3600 # maximum reuse age
+namespace = "dev-harness" # identifies one trusted cache-sharing domain
+bypass_rate = 0.0 # 0.1 would re-run 10% of cacheable calls live
+cache_nondeterministic = false # set true to cache nondeterministic requests
+
+[components.config.response_cache.backend]
+kind = "in_memory" # or "redis" for a shared cache
+```
+
+With this configuration the first eligible occurrence of a request runs the
+provider and attempts to store a complete answer. A later eligible, identical
+request within the TTL can be served from the store when that write succeeds
+and the entry remains present. The request's provider surface is auto-detected
+from its shape, so there is nothing else to configure. For file discovery and
+precedence rules, refer to
+[Plugin Configuration Files](/configure-plugins/plugin-configuration-files).
+
+## Plugin Configuration
+
+Use plugin configuration when the application should let NeMo Relay own the
+cache lifecycle.
+
+
+
+```python
+import asyncio
+
+import nemo_relay
+
+adaptive_config = nemo_relay.adaptive.AdaptiveConfig(
+ response_cache=nemo_relay.adaptive.ResponseCacheConfig(
+ ttl_seconds=3600,
+ namespace="dev-harness",
+ ),
+)
+
+plugin_config = nemo_relay.plugin.PluginConfig(
+ components=[nemo_relay.adaptive.ComponentSpec(adaptive_config)]
+)
+
+report = nemo_relay.plugin.validate(plugin_config)
+if any(diagnostic["level"] == "error" for diagnostic in report["diagnostics"]):
+ raise RuntimeError(report["diagnostics"])
+
+def call_model(_request):
+ return {
+ "model": "gpt-4o",
+ "choices": [{
+ "message": {"role": "assistant", "content": "NeMo Relay instruments agent calls."},
+ "finish_reason": "stop",
+ }],
+ }
+
+async def main():
+ await nemo_relay.plugin.initialize(plugin_config)
+ try:
+ # Managed calls need no changes. The first call runs the provider;
+ # an identical repeat can be served from the cache.
+ request = nemo_relay.LLMRequest(
+ {},
+ {
+ "model": "gpt-4o",
+ "messages": [{"role": "user", "content": "What is NeMo Relay?"}],
+ "temperature": 0,
+ },
+ )
+ await nemo_relay.llm.execute("openai", request, call_model)
+ await nemo_relay.llm.execute("openai", request, call_model)
+ finally:
+ nemo_relay.plugin.clear()
+
+asyncio.run(main())
+```
+
+
+
+```js
+const adaptive = require("nemo-relay-node/adaptive");
+const plugin = require("nemo-relay-node/plugin");
+
+const adaptiveConfig = adaptive.defaultConfig();
+adaptiveConfig.responseCache = adaptive.responseCacheConfig({
+ ttlSeconds: 3600,
+ namespace: "dev-harness",
+});
+
+const pluginConfig = plugin.defaultConfig();
+pluginConfig.components = [adaptive.ComponentSpec(adaptiveConfig)];
+
+const report = plugin.validate(pluginConfig);
+if (report.diagnostics.some((diagnostic) => diagnostic.level === "error")) {
+ throw new Error(JSON.stringify(report.diagnostics));
+}
+
+void (async () => {
+ await plugin.initialize(pluginConfig);
+ try {
+ // Eligible deterministic repeats can be served from the cache.
+ } finally {
+ plugin.clear();
+ }
+})().catch((error) => {
+ console.error(error);
+ process.exitCode = 1;
+});
+```
+
+
+
+```rust
+use nemo_relay::plugin::{
+ clear_plugin_configuration, initialize_plugins, validate_plugin_config, PluginConfig,
+};
+use nemo_relay_adaptive::plugin_component::{register_adaptive_component, ComponentSpec};
+use nemo_relay_adaptive::{AdaptiveConfig, ResponseCacheConfig};
+
+#[tokio::main]
+async fn main() -> Result<(), Box> {
+ let mut adaptive = AdaptiveConfig::default();
+ adaptive.response_cache = Some(ResponseCacheConfig {
+ ttl_seconds: 3600,
+ namespace: "dev-harness".into(),
+ ..ResponseCacheConfig::default()
+ });
+
+ let mut plugin_config = PluginConfig::default();
+ plugin_config.components.push(ComponentSpec::new(adaptive).into());
+
+ register_adaptive_component()?;
+ let report = validate_plugin_config(&plugin_config);
+ assert!(!report.has_errors());
+
+ let _active = initialize_plugins(plugin_config).await?;
+ // Eligible deterministic repeats can now be served from the cache.
+
+ clear_plugin_configuration()?;
+ Ok(())
+}
+```
+
+
+
+
+
+Canonical plugin documents and `plugins.toml` use `snake_case`. Python
+response-cache fields follow that convention, while Node.js uses `camelCase` at
+its response-cache helper surface. `ComponentSpec` and `AdaptiveRuntime`
+serialize those fields to canonical keys. Python uses
+`BackendSpec.redis(url, key_prefix=...)`; Node.js uses
+`adaptive.redisBackend(url, keyPrefix)`. Both Redis helpers default the prefix
+to `"nemo_relay:"`, overriding the value in the fields table. Configure the
+same `key_prefix` in every process and binding that should share entries.
+
+
+## Manual API
+
+Use the manual runtime API when an integration needs to own the adaptive
+lifecycle directly instead of activating the top-level plugin component.
+
+
+
+```python
+import asyncio
+
+import nemo_relay
+
+adaptive_config = nemo_relay.adaptive.AdaptiveConfig(
+ response_cache=nemo_relay.adaptive.ResponseCacheConfig(namespace="dev-harness"),
+)
+
+async def main():
+ runtime = nemo_relay.adaptive.AdaptiveRuntime(adaptive_config.to_dict())
+ await runtime.register()
+ try:
+ # Run instrumented application work here.
+ runtime.wait_for_idle()
+ finally:
+ await runtime.shutdown()
+
+asyncio.run(main())
+```
+
+
+
+```js
+const adaptive = require("nemo-relay-node/adaptive");
+
+const adaptiveConfig = adaptive.defaultConfig();
+adaptiveConfig.responseCache = adaptive.responseCacheConfig({
+ namespace: "dev-harness",
+});
+
+const runtime = new adaptive.AdaptiveRuntime(adaptiveConfig);
+void (async () => {
+ await runtime.register();
+ try {
+ // Run instrumented application work here.
+ runtime.waitForIdle();
+ } finally {
+ await runtime.shutdown();
+ }
+})().catch((error) => {
+ console.error(error);
+ process.exitCode = 1;
+});
+```
+
+
+
+```rust
+use nemo_relay_adaptive::{AdaptiveConfig, AdaptiveRuntime, ResponseCacheConfig};
+
+#[tokio::main]
+async fn main() -> Result<(), Box> {
+ let mut adaptive = AdaptiveConfig::default();
+ adaptive.response_cache = Some(ResponseCacheConfig {
+ namespace: "dev-harness".into(),
+ ..ResponseCacheConfig::default()
+ });
+
+ let mut runtime = AdaptiveRuntime::new(adaptive).await?;
+ runtime.register().await?;
+
+ // Run instrumented application work here.
+
+ runtime.wait_for_idle();
+ runtime.shutdown().await?;
+ Ok(())
+}
+```
+
+
+
+
+## What Gets Cached
+
+Only complete, replayable answers are stored:
+
+- A response with a non-null `error` or a `status` such as `failed`,
+ `cancelled`, `incomplete`, or `in_progress` is never stored.
+- Stateful requests bypass the cache entirely: a request that opts into
+ server-side persistence (a non-null `store` value other than `false`, or an
+ OpenAI Responses call without an explicit `store = false`), continues a
+ stored interaction (`previous_response_id`), or references server-side state
+ (`conversation`, `container`). Their answers depend on state the cache key
+ cannot see.
+- Streaming calls are cached too. On a miss the live chunks are forwarded to
+ the consumer while being assembled into one aggregate response — the same
+ shape a buffered call stores, so buffered and streaming calls share one
+ keyspace. On a hit the stored answer is replayed as provider-native chunks
+ (OpenAI Chat deltas, OpenAI Responses lifecycle events, or Anthropic
+ Messages events), so strict streaming clients parse it like a live stream.
+ Provider-terminal token-limited Chat (`finish_reason = "length"`) and
+ Anthropic (`stop_reason = "max_tokens"`) answers can be stored. OpenAI
+ Responses answers with `status = "incomplete"`, streams without a terminal
+ event, and streams whose content cannot be replayed faithfully (for example
+ thinking blocks) are never stored. A streaming request whose surface cannot
+ be inferred runs live, uncached.
+
+Streaming publication is write-behind: Relay can report end-of-stream before
+the backend write finishes. `wait_for_idle()` does not wait for cache writes.
+An immediate repeat can therefore run live, and concurrent identical misses
+can each call the provider because the cache does not coalesce them.
+
+A buffered hit returns the stored response unchanged. A streaming hit replays
+an equivalent provider-native stream. Available saved-token and estimated-cost
+values are reported on the `response_cache` mark, never by editing a buffered
+response body.
+
+## Cache Keys
+
+Two requests hit the same entry when they are the same request after
+normalization, under the default `key_strategy = "exact_request"`:
+
+- The request is decoded to its normalized form and fingerprinted with
+ SHA-256, so provider-shaped differences that mean the same thing collapse to
+ one key. When a request cannot be decoded faithfully, the raw body is
+ fingerprinted instead — that fallback can only cost a miss, never a wrong
+ hit.
+- Field order and whitespace never matter (RFC 8785 canonicalization).
+- Only the built-in noise fields `stream`, `user`, `metadata`, and `store` are
+ dropped. All other normalized or raw provider controls remain in the key,
+ including `service_tier`; there is no configurable skip list.
+- OpenAI Chat requests preserve whether the caller used `max_tokens` or
+ `max_completion_tokens`, because providers can treat the two fields
+ differently.
+- Tool-call IDs are normalized, so randomly generated per-call IDs do not
+ fragment the keyspace.
+- Request headers stay out of the key unless named in `header_allowlist`.
+ Allowlist every trusted, non-secret response-affecting header; an omitted
+ header does not partition the key. Known auth headers are rejected, but
+ validation cannot recognize every custom credential name, so never allowlist
+ credentials.
+- The provider name and required namespace partition every key. A
+ Switchyard-selected backend ID is also partitioned automatically. Without
+ Switchyard, a provider name cannot distinguish upstreams that reuse that
+ name, so use separate configurations and namespaces for different trusted
+ upstream domains.
+- Requests containing integers outside the exactly representable RFC 8785
+ range (less than `-2^53` or greater than `2^53`) bypass the cache.
+
+## Observability
+
+Every cache decision emits a `response_cache` mark with
+`data.status` set to one of:
+
+| Status | Meaning |
+|---|---|
+| `hit` | Served the stored answer; the provider was skipped. |
+| `miss` | No entry was served; the call ran live. After an ordinary lookup miss, Relay attempts to store a cacheable result. |
+| `bypass` | The request is not cacheable, or the `bypass_rate` sampler chose to run live. |
+
+Mark attributes use `nemo_relay.response_cache.*`: `backend`, `surface`,
+`key_hash` (the `sha256:…` fingerprint), `ttl_ms`, and `age_ms` as applicable;
+`saved_tokens` and `saved_cost_usd` appear on hits when they can be derived. A
+`reason` appears on bypasses and store-error misses (for example `sampled`,
+`stateful_store`, `store_error`, or `stream_no_codec`). Cache marks never
+include prompts, answers, or credentials.
+
+`nemo-relay doctor` reports the cache state: `not configured` when the section
+is absent, `configured but disabled (adaptive plugin disabled)` when the
+adaptive component is off, `on; backend '' reachable` when healthy, and
+a failure when the config is invalid or the backend is unreachable.
+
+## Fields
+
+| Field | Default | Notes |
+|---|---|---|
+| `ttl_seconds` | `3600` | Maximum entry age in seconds; a backend can evict an entry sooner. |
+| `namespace` | `""` (unconfigured) | Required non-empty trust-domain partition folded into every key. Empty or whitespace-only values are rejected. |
+| `priority` | `50` | LLM execution intercept priority. Lower values run earlier. |
+| `bypass_rate` | `0.0` | Probability in `[0.0, 1.0]` of running a cacheable call live. A sampled call attempts to refresh the entry when its result is cacheable and the write succeeds. |
+| `cache_nondeterministic` | `false` | Only requests with an explicit numeric `temperature = 0` are eligible. Set `true` to cache and reuse sampled responses. |
+| `key_strategy` | `"exact_request"` | The only supported strategy: reuse requires the same normalized request. |
+| `header_allowlist` | `[]` | Trusted, non-secret response-affecting headers folded into the key (case-insensitive). Known auth headers are rejected. |
+| `backend.kind` | `"in_memory"` | `"in_memory"`, or `"redis"` (requires building with the `redis-backend` feature). |
+| `backend.config.max_bytes` | 256 MiB | In-memory size budget; the oldest entries are evicted first. |
+| `backend.config.url` | — | Redis connection URL. Required for the `redis` backend. |
+| `backend.config.key_prefix` | `"nemo-relay:llm-cache:"` | Prefix for keys in Redis. |
+
+For a gateway that uses Switchyard, `switchyard.priority` must be lower than
+`response_cache.priority`. To derive keys before ACG rewrites requests, set
+`response_cache.priority` lower than `acg.priority`. With all three components,
+priorities of `0`, `40`, and `50`, respectively, satisfy both orderings.
+
+
+Cached responses are stored unredacted. PII sanitize guardrails rewrite emitted
+telemetry, never payloads, so the store holds full response bodies. Cache
+entries can also store provider and model diagnostics plus the key fingerprint;
+they do not store full request bodies or headers. A shared Redis backend must be
+trusted and access-controlled. Use a separate configuration and namespace for
+each mutually untrusted tenant or upstream domain.
+
+
+## Common Validation Failures
+
+- `namespace` is empty or whitespace-only, `ttl_seconds` is `0`, or
+ `bypass_rate` is outside `[0.0, 1.0]`.
+- `key_strategy` is not `"exact_request"`.
+- `header_allowlist` names an auth header such as `authorization` or
+ `x-api-key`.
+- `in_memory` `max_bytes` is zero or is not an integer.
+- `backend.kind` is unknown; or `redis` has a missing, non-string, or
+ whitespace-only `backend.config.url`, uses a non-string `key_prefix`, or is
+ unavailable because Relay was built without the `redis-backend` feature.
+- Gateway Switchyard priority is equal to or greater than
+ `response_cache.priority`.
diff --git a/docs/reference/migration-guides.mdx b/docs/reference/migration-guides.mdx
index 8e11c37ea..91d95bdf2 100644
--- a/docs/reference/migration-guides.mdx
+++ b/docs/reference/migration-guides.mdx
@@ -27,6 +27,10 @@ for 0.7 before using those surfaces. NeMo Relay does not adapt synchronous Rust
middleware callbacks or one-argument LLM sanitizer callbacks.
+Rust code that constructs `AdaptiveConfig` with an exhaustive struct literal
+must add `response_cache: None`. Prefer `..AdaptiveConfig::default()` when the
+literal should remain compatible with new optional fields.
+
### Migrate Middleware Callbacks
The following callback families are now asynchronous: conditional execution