Catalog of AuthBridge pipeline plugins — every plugin with a Go
implementation that calls plugins.RegisterPlugin(). For the config
convention, session-event contract, and lifecycle interfaces plugins
implement, see plugin-reference.md. For
writing a new plugin, see plugin-tutorial.md.
"Production ready?" reflects whether the plugin is carried by a shipped
profile in scripts/profile-tags versus available only on
explicit request or requiring a separate binary. Every plugin is opt-in
(-tags include_plugin_<name>); a build with no tags registers none. It is
a packaging signal, not a claim about test coverage or operational maturity.
"Direction" is inbound (caller → this agent) or outbound (this agent → callee); "both" means the plugin evaluates on both pipelines. "Default config?" marks whether the plugin is enabled in Rossoctl's default AuthBridge pipeline YAML, not whether it is compiled into the binary (see "Production ready?" above for that).
| Name | Description | Production ready? | Direction | Default config? |
|---|---|---|---|---|
a2a-parser |
Parses A2A messages into pctx.Extensions.A2A for downstream plugins. |
Beta | Inbound | No |
context-guru |
Compacts the outbound LLM request context before forwarding. | Opt-in | Outbound | No |
cpex |
APL DSL + named CPEX plugins (Cedar, PII, audit, …) over a single chain step. | Opt-in | Outbound | No |
ibac |
LLM-judge intent-based access control for outbound tool calls. | Alpha | Outbound | No |
inference-parser |
Parses LLM completions into pctx.Extensions.Inference. |
Alpha | Outbound | No |
inference-router |
Sends each coding agent's new sessions to the inference server chosen for it; a session stays on the server it started on. | Alpha | Outbound | No |
jwt-validation |
Inbound JWT validation (signature, issuer, audience) against JWKS. | Ready | Inbound | YES |
lineage-telemetry |
Emits two facts-only OTel lineage spans per HTTP exchange, parented across pods through one tracestate member. |
Alpha | Both | No |
litellm-budget-track |
Tracks x-litellm-response-cost (with -original fallback) and enforces a daily budget limit. Place on whichever chain carries LLM traffic — inbound when fronting the LLM endpoint, outbound when hosting an agent via authbridge exec. |
Alpha | Both | No |
mcp-parser |
Parses MCP tool calls/results into pctx.Extensions.MCP. |
Beta | Outbound | No |
opa |
OPA policy enforcement for inbound and outbound requests. | Alpha | Both | No |
sparc |
Pre-tool reflection: blocks ungrounded/hallucinated tool calls. | Alpha | Outbound | No |
static-inject |
Swaps a placeholder credential for a real static credential on outbound requests. | Alpha | Outbound | No |
session-budget |
Enforces per-session token, call, and duration budgets via Redis. | Alpha | Outbound | No |
token-broker |
Exchanges incoming tokens against a configured IdP via a broker service. | Alpha | Outbound | No |
token-exchange |
RFC 8693 outbound token exchange per route. | Ready | Outbound | YES |
tool-prune |
Removes unused tool definitions from inference requests. | Alpha | Outbound | No |
Parses A2A JSON-RPC 2.0 request bodies into pctx.Extensions.A2A
(method, session ID, message parts, role) for downstream guardrails.
No configuration — registered as a bare plugin name, no config: block.
Compacts an agent's outbound LLM request context before forwarding,
using the embedded context-guru engine. OnResponse is currently a
pass-through; model-driven expand/restore is a later integration.
Opt-in at build time (-tags include_plugin_contextguru) because its
engine pulls a large transitive dependency set.
paths([]string) — inference request paths to compact. Default:/v1/chat/completions,/v1/completions,/v1/messages.model(object) — optional "cheap" LLM endpoint for model-backed components (summarize, extract:code); omitted means those degrade to deterministic/no-op.base_url— OpenAI-compatible endpoint base.model— model name to call.api_key— optional bearer token.max_tokens— completion cap, default 4096.timeout_ms— per-call timeout, default 150000.
engine(object) — native context-guru config (preset / pipeline / per-component / store), passed through verbatim. Default:preset: balanced.
Bridges AuthBridge hooks to the CPEX
framework (a policy enforcement runtime for AI agents): an APL DSL plus named
CPEX policy plugins (Cedar, PII, audit, …). Requires the separate
cortex-cpex binary (-tags cpex, CGO_ENABLED=1, links a pinned
libcpex_ffi.a). Full details in cpex-plugin.md;
see also the plugin's README.
hooks.on_request/hooks.on_response([]string) — CPEX hook names to fire on each phase, in order (AuthBridge classifies traffic onto thecmf.*hooks — see Hook chains).config(string) — inline CPEX runtime YAML (plugins:/global:/plugin_settings:); mutually exclusive withconfig_file.config_file(string) — path to a file with the CPEX runtime YAML; mutually exclusive withconfig.fail_open(bool) — allow traffic through if CPEX itself errors/panics. A CPEX policy deny is always honored regardless. Defaultfalse.worker_threads(int) — size of CPEX's tokio worker pool;0= automatic.bypass_hosts/bypass_paths([]string) — globs skipped entirely (outbound only for hosts); default to Keycloak/SPIRE/observability infra.
LLM-judge intent-based access control: judges outbound tool calls against recorded inbound user intent. Full details, including the prompt-injection threat model, in ibac-plugin.md.
The "LLM-judge service" is any OpenAI-compatible chat-completions
endpoint (a local Ollama/vLLM, or a hosted provider) — AuthBridge ships
no judge of its own. The plugin POSTs the recorded user intent plus the
proposed action to {judge_endpoint}/v1/chat/completions and parses an
allow/deny verdict from the reply — see
Request Flow. For guidance on which
model to point it at, see
Choosing a Judge Model.
judge_endpoint(string) — base URL of the LLM-judge service ({endpoint}/v1/chat/completions).judge_model(string) — model name passed to the judge.judge_bearer(string) — optional bearer token; empty for unauthenticated local LLMs.system_prompt(string) — override the built-in judge system prompt.timeout_ms(int) — per-call timeout; values below 100 rejected. Default 5000.judge_max_tokens(int) — cap on judge reply length. Default 1024.judge_json_mode(*bool) — forceresponse_format: json_object. Defaulttrue.judge_inference(bool) — also judge outbound LLM-reasoning traffic (high cost). Defaultfalse.agent_llm_host(string) — the agent's own LLM host; auto-added tobypass_hosts.bypass_hosts/bypass_paths([]string) — globs skipped without judging.no_intent_policy(string) — behavior when an action has no recorded intent:allow(default) ordeny.unclassified_policy(string) — behavior when no parser claimed the request:passthrough(default) orjudge.
Parses outbound LLM inference requests/responses into pctx.Extensions.Inference for
downstream policy plugins, and prices the finished response — it is the one place
tokens become dollars.
It reads two dialects, chosen by how the request path ends, under any prefix: a path ending
in /completions is OpenAI chat completions (or legacy completions), and one ending in
/v1/messages is Anthropic Messages. That covers providers that mount the same API under a
prefix of their own — IBM Bob's /inference, OpenCode Zen's /zen, OpenRouter's /api,
Groq's /openai, Azure's /openai/deployments/<d> — with no code change. A body is taken
for inference only if it carries a messages array, or a prompt for legacy completions.
A body that fails that check is recorded the way an unrecognised path is: with no inference
record, and still priced from a gateway's cost header. Other dialects — the Responses API,
Gemini's native API, Bedrock's native API — are not parsed.
Costing lives here because this is the only component that knows when usage is final: it
owns the three response-finalization paths and the assembled-usage handling (Claude Code's
?beta=true path reports cache counts on message_delta, not message_start). It also means a
priced request can no longer be missing its record: previously the figure came from
litellm-budget-track, so a pipeline without that plugin showed tokens and no money, with
the same field silently meaning "modelled" rather than "authoritative" depending on
configuration.
A record is not the same as a price. Where no rate resolves for the model, the record is
still published — carrying the token counts, any avoided cost, and no dollar figure — and the
gap is named in /v1/usage's unpricedBy so an operator knows which pricing: entry to
add. An absent figure is reported as absent, never as $0.00.
The arithmetic and the gateway header semantics are in core/cost/settle, not in the parser:
a provider-shaped body parser has no business knowing one gateway's header names. The result
is published as a cost record keyed cost on the session event (see
Cost records).
No configuration — no config struct, does not implement Configurable. Rates arrive by
injection from the top-level pricing: section; with none configured the parser
still parses and reports the traffic as unpriced.
Sends a coding agent's new sessions to the inference server chosen for it — a
LiteLLM gateway, say, with its own URL and key — and keeps every session on the
server it started on. Routing is opt-in per agent: an agent not listed under
agents is not routed, and its requests, keys included, pass untouched. Manage it
with agentop server;
the config stays hand-editable. Built for the laptop: it needs a listener that
honors a redirect, so cortex-envoy and a config with an mtls: block refuse it
at startup, and in-cluster use has not been examined.
servers(map, required) — inference servers by name: lowercase letters, digits,.,_and-. Each has:url(string, required) —scheme://host[:port],httporhttps, no path, query or fragment. A port must be 1–65535; the scheme's default is dropped, sohttps://x:443isx. No two servers may share a host, compared without port or case: a session's server is named from the host its requests went to. A plain-httpserver on another machine logs a WARN at load, since routed requests would cross the network decrypted.key(string, required) — the API key sent to this server in place of the client's, in the header the client used:X-Api-Keywhen it sent one,Authorization: Bearerotherwise, both when it sent both. Those two are the only headers replaced: a credential the client sends in any other header, such as one Claude Code'sANTHROPIC_CUSTOM_HEADERSsets, reaches the server unchanged. Printable ASCII, no spaces./configand/v1/pipelineredact it.opus,sonnet,haiku(strings) — this server's own model for each of Claude Code's model families, for a server that does not serve Claude Code's names. All three or none: a partial set is refused, naming the families given and missing. Leave them out for a server that serves Claude Code's names, and every name passes through unchanged.
agents(map) — agent name, as agentop shows it (claude-code,opencode), to a server name. Each value must name a listed server.unknown, the name for requests with no User-Agent, cannot be routed.
- name: inference-router
config:
servers:
ete:
url: https://ete-litellm.example.com
key: sk-…
glm:
url: https://glm-litellm.example.com
key: sk-…
opus: glm-5.3 # this server's names for Claude Code's three families
sonnet: glm-5.3
haiku: glm-5.3
agents:
claude-code: glm # only agents listed here are routedWhat it does to a request. Only a request the forward proxy re-sends — a plain
proxied one, or one decrypted by the TLS bridge — can be routed. A CONNECT or a
transparently redirected connection is dialed where the client chose, and is
skip/not_redirectable. A request to a host no server has is
skip/not_an_inference_server. On a server's host every path is handled,
/v1/models and count_tokens included, and the session decides:
-
The first request the router sees from a session pins it, in the process store, to where that request went: a server when the redirect took effect, or "not routed" when the request stayed where the client sent it — because the agent is not routed, or because the router runs under
on_error: observe, even for an agent with a server. A session started under observe therefore stays unrouted after a switch to enforce. A first request whose redirect failed, or that was refused for its model, pins nothing, and the next one decides again. Every request renews the pin, which lapses 30 days after the last. -
Where that request goes is decided by the session's history, for a routed agent: a session already running when routing was first configured, quiet while it was, has no pin but is not new. The latest earlier inference request that agent sent in the session to a server's host, as the session store holds it, keeps the session on that server, from then with that server's key. A session with no such request is new and goes to its agent's current server. A request to any other host is no evidence, since the router never routes one: an OpenCode session that used another provider first is new when it addresses a server. The history is read from the
inferencerecordinference-parserputs on each request, so the rule needsinference-parserin the outbound chain; without it, a session the router has not pinned is always treated as new. -
A pin is its agent's. A request from another agent filed under the same session id — by process attribution, which files a command an agent runs under that agent's session and is on by default on a laptop; by the active-session fallback with client affinity off; or by a header id two clients share — is decided as if the session were unpinned, from that agent's own history and server, and leaves the pin alone.
-
Changing
agentstherefore moves only new sessions while the proxy runs: the pins and the history survive a hot reload. A restart loses both — the router reads the store's history in memory, not the session archive on disk — so a running session's next request after a restart is treated as a new session's, and moves if its agent's server is another. The store also drops quiet sessions while the proxy runs —session.max_sessions(100 by default) evicts the least recently used, and a configuredsession.ttlexpires them — and a session dropped before the router pinned it is treated as new too; a pin outlives that. A request with no session follows its agent's current server, unpinned, and so does one filed under a synthetic session — thedefaultbucket or apending:<agent>id — since each holds many conversations, not one. -
Not routed:
skip/not_routed, and the request is left as the client sent it. -
Pinned to a server since removed:
deny/pinned_server_removed, a 503 asking for a new session. The conversation is never moved to another server. -
Otherwise the request is redirected to the server, and the key is replaced only once the redirect has taken effect, so it goes to the server and nowhere else. Every routed request is redirected, even one that already names the server's host: the
Hostheader is the client's word, and a TLS-bridged request is otherwise dialed to the host the clientCONNECTed to, which need not be the same. Each therefore carries the framework'smodify/redirectedrecord — withfromequal totowhen the agent's base URL (ANTHROPIC_BASE_URL, for Claude Code) already names the server — followed by the router'smodify/routedwithserverandpin(new,existingornone). -
Models. On a server with
opus,sonnetandhaiku, a routed request'smodelis mapped by family — the family is a word of the name it asks for, soclaude-opus-5-5andclaude-haiku-4-5-20251001map with no list of ids — throughpctx.SetRequestModel, which changes that one JSON value and nothing else. The timeline gains the framework'smodify/body_rewrittenandmodify/model_rewrittenbetweenmodify/redirectedandmodify/routed, and the inference record'smodelbecomes the server's, withrequestedModelkeeping Claude Code's: settlement prices the server's model, and agentop's detail pane shows both. Every routed agent's request is decided in this order, Claude Code's and OpenCode's alike:- A name with one family word is mapped to the server's model for that family,
even when it is also one of the server's own models — so a server whose models
are Claude's own, a downgrader with
opus: claude-sonnet-5say, works. The cost is that a server model whose own name holds a family word is read as that family. - Otherwise a name that is exactly one of the server's three models — one picked
from its model list with
/model, say — goes as it is. - Otherwise a Claude model name, one with
claudeas a word —claude-fable-5-1, say, or one naming two families — isdeny/no_model_for_familywith the requested name asmodel, a 400 sayingglm has no model for claude-fable-5-1. It asked for a Claude model the server has none for, and nothing is guessed. - Any other name goes as it is, the client's to choose and the server's to
answer: OpenCode asking a GLM server for
glm-4.6is served.
A body that names a model the rewrite cannot read — an empty one, one that is not a string, or
modelnamed twice or only in another letter case — isdeny/model_rewrite_failed, a 400 too, since the fix is the client's. A refused request carries no server key, and a session whose first request is refused pins nothing. Its denied row names the server's host, withrequestedHostthe one the client asked for, because the listener applies the redirect before it answers the refusal:/v1/usagecounts the denial under the server, and agentop's detail pane shows aredirected:line for it. A request whose body names no model —GET /v1/models, a body that is not JSON, or JSON with nomodelkey — is routed as it is. - A name with one family word is mapped to the server's model for that family,
even when it is also one of the server's own models — so a server whose models
are Claude's own, a downgrader with
-
A redirect the listener refuses:
deny/redirect_failed, a 503. -
Under
on_error: observenothing moves, the client's key stays and no model is mapped or refused. The timeline shows two rows: the framework's shadowmodify/redirected, and the router'sobserve/would_route.
Put it last in the outbound chain, where agentop server add puts it. A
redirect moves pctx.Host, so the plugins before the router decide on the host
the client asked for, and a plugin after it would see the server's instead; put
nothing after it that keys on the host. Session events, usage and cost follow the
server's host, and each event's requestedHost keeps the one asked for when it
differs. The router also rewrites the request body to map models, so the framework
holds it after every body reader: a chain with a parser after the router is
refused, at startup and on reload.
Validates inbound JWTs: signature via JWKS, issuer, and audience.
issuer(string) — expectedissclaim; required.jwks_url(string) — JWKS endpoint; derived from Keycloak URL/realm or issuer when omitted.keycloak_url/keycloak_realm(string) — used to derivejwks_urlwhen omitted.audience(string) — expectedaudclaim; one ofaudience/audience_file/audience_mode=per-hostrequired.audience_file(string) — file to read expected audience from. Default/shared/client-id.txt.audience_mode(string) —static(default) orper-host(derived from theHostheader).allowed_audiences([]string) — extra audience values accepted (OR semantics).bypass_paths([]string) — path globs skipped. Default/healthz,/readyz,/livez,/metrics,/.well-known/*.placeholder_mode(bool) — replace the validated inbound token with an opaque placeholder before forwarding, for later outbound resolution. Defaultfalse.placeholder_ttl(string) — how long the real token is retained. Default1h.
Emits two facts-only OpenTelemetry spans per HTTP exchange — a request
span on sight and a response span at stream end, paired by
lineage.exchange.id — carrying direction, protocol, endpoints, outcome
and, optionally, the parsed payload. Cross-pod parenting rides one
tracestate member, lineage-parent; a request that arrives with no
valid traceparent is forwarded with one naming the request span, and a
valid one is never modified. The wire format is
lineage-wire-contract.md. Place it after
the protocol parsers (declared in RequiresAny) and after
jwt-validation when the principal facts are wanted; a request-phase
denial by a plugin ordered before it emits no spans.
otel_endpoint(string) — OTLP gRPC target:host:port,http://host:portorhttps://host:port; any other scheme is refused. Defaultlocalhost:4317.otel_tls(bool) — dial the collector with TLS, verified against the system roots orotel_ca_file. Anhttps://endpoint implies it;https://withotel_tls: false, andhttp://withotel_tls: trueorotel_ca_file, are refused. A plaintext dial to a non-loopback collector is allowed and logged as a WARN at start. Defaultfalse.otel_ca_file(string) — PEM bundle to verify the collector's certificate against (a private CA, e.g. cert-manager issued). Impliesotel_tls; with an explicitotel_tls: falseit is refused; an unreadable file or one with no certificate refuses to start. Default: system roots.capture_io(bool) — attach the parsed request/response content asinput.value/output.value. Defaultfalse.max_payload_bytes(int) — cap on those two values, cut on a UTF-8 boundary with a…[truncated]marker;0or unset takes the default,-1attaches whole, any other negative is refused at start. Default4096.max_attr_bytes(int) — cap on every variable-content string attribute (url.path,lineage.peer.host,mcp.tool, …) and the span name — except the two identity factslineage.self.idandlineage.self.namespace, which are operator configuration and never truncated; same0/-1/ negative semantics asmax_payload_bytes. Default256.mint_traceparent(bool) — forward atraceparentnaming this request span when the request carried no valid one;false= a pure observer that writes notraceparent. Defaulttrue.bypass_paths([]string) — path globs (path.Match, query stripped, path normalized — the shared bypass matcherjwt-validationandsparcuse) that produce no spans. Default/.well-known/*,/healthz,/readyz,/health. Setting either bypass key replaces its default list rather than extending it, as inibac/sparc/cpex; an entry matching everything is refused at start.bypass_hosts([]string) — outbound host globs (path.Match, port stripped, case folded) that produce no spans; ignored inbound, whereHostis caller-controlled. Defaultotel-collector,otel-collector.*,jaeger,jaeger.*,zipkin,zipkin.*,prometheus,prometheus.*.self_id(string) — this workload's identity, emitted aslineage.self.id(a SPIFFE ID reduced to its last path segment); a blank value, or one with no non-empty/-segment (/), is refused at start.self_id_file(string) — read whenself_idis empty. Until it is readable and carries an identity the plugin is not ready and skips every exchange (no span, no header), re-reading the file in the background while/readyznames it — the same handlingjwt-validationgives this path, so a late Secret mount never fails the sidecar (a pod probing/readyzstays out of rotation until the file lands). Refused at start only whenself_idis also empty. Default/shared/client-id.txt.namespace(string) — this workload's Kubernetes namespace (an RFC 1123 DNS label), emitted aslineage.self.namespaceon every span: the other half of its identity, sinceself_idis the last segment of a SPIFFE ID and the same segment in two namespaces is two workloads. Required (ornamespace_file) — absent, blank, or not a DNS label is refused at start; never derived from the SPIFFE path. The attach kit writes itsNAMESPACE; a sidecar older than this key rejects a config that carries it, so image and config flip together.namespace_file(string) — read once at start whennamespaceis empty; meant for/var/run/secrets/kubernetes.io/serviceaccount/namespace, the file the kubelet projects from the pod's own metadata — the one source that is right in every copy of a ConfigMap shared across namespaces (the platform'sauthbridge-runtime-config), where an inline literal would be wrong everywhere but one. Absent, blank, or not a DNS label refuses at start; no default, no poller.
Keeps the daily spend ledger and enforces a spend budget, from the cost record
inference-parser publishes. Full details in
litellm-budgettrack-plugin.md.
Provider-specific: x-litellm-response-cost is emitted only by
LiteLLM, so this plugin works only when
LiteLLM is the inference provider in front of the model. Against a
provider that doesn't set the header (raw OpenAI, Ollama, vLLM, …), no
cost is ever accumulated and the budget never trips.
spend_file(string) — path to the JSON spend ledger file; required. The ledger is a small JSON file the plugin creates and rewrites, holding the current UTC date plus the cumulative spend and call count for that day (it resets automatically at midnight UTC) — see Ledger Format.max_budget(float64) — daily budget in USD; required, must be > 0.- No rate options, and no pricing at all. This plugin bills a figure it does not compute:
inference-parsersettles the cost and publishes the record, and this plugin adds the day's total, enforces the cap, and warns when the rate table disagrees with what the gateway charged. Rates live in the top-levelpricing:section. - Requires
inference-parserLATER in the chain (RequiresLater). The response passes walk the chain in reverse, so the parser must sit at a higher index to fold each frame before this plugin settles the cost. A chain without it — or with it earlier — fails to build.
Parses MCP tool calls/results into pctx.Extensions.MCP for downstream
policy plugins.
This plugin makes no decisions of its own — it exists to feed others. The
plugins that consume pctx.Extensions.MCP are:
| Consumer | How it uses the MCP extension |
|---|---|
ibac |
Reads the tool name and arguments to judge the call against user intent. Declares mcp-parser in RequiresAny. |
sparc |
Extracts the tool name/arguments to reflect on, and returns clarifications as MCP results. Declares mcp-parser in RequiresAny; required in enforcement: mcp mode. |
cpex |
Converts the parsed call/result into a CMF message for the cmf.tool_pre_invoke / cmf.tool_post_invoke hooks. Declares mcp-parser in RequiresAny. |
opa |
Exposes the parsed call as input.mcp for policy (add mcp.params to include for arguments). |
Place mcp-parser before these plugins on the outbound chain;
without it they see no MCP data and pass the traffic through
unclassified.
paths([]string) — URL path globs treated as MCP endpoints (for body-less transport detection: SSE GET, session-terminate DELETE). Default["/mcp"].
Evaluates OPA (Open Policy Agent) policy bundles against inbound and outbound requests, using an embedded OPA engine and four fixed decision paths. Full details in the plugin's README.
bundle_url(string) — base URL of the Rossoctl Bundle Server, the in-cluster service that serves per-agent OPA policy bundles keyed by SPIFFE ID (see how it works); required.agent_id_file(string) — path to the agent's client-ID file. Default/shared/client-id.txt.agent_id(string) — inline agent ID; overridesagent_id_filewhen set.polling_min_delay/polling_max_delay(int) — bundle polling interval bounds in seconds. Defaults 10 / 120.include([]string) — optional field groups exposed in the OPA input document (e.g.mcp.params,a2a.content,inference.messages); default lean/empty.
Pre-tool reflection: sends proposed tool calls to a
SPARC reflection service — a companion
HTTP service wrapping the SPARCReflectionComponent from the
agent-lifecycle-toolkit
(ALTK) package — and enforces the configured policy on the result. It
must be deployed once per cluster before enabling this plugin. Full
details in sparc-plugin.md.
reflector_endpoint(string) — base URL of the SPARC reflection service ({endpoint}/reflect); required.reflector_bearer(string) — optional bearer token.enforcement(string) —mcp(gate outbound MCPtools/call, default) orinference(gate/rewrite LLM completions).track(string) — reflection track:fast_track(default),slow_track,syntax,spec_free,transformations_only.timeout_ms(int) — per-call timeout; values below 100 rejected. Default 30000.on_reject_action(string) —observe(log only),reflect(default, return clarification), ordeny(hard block).deny_score_threshold(float64) — escalate a reject to hard deny when the grounding score is at or below this value.0disables escalation.fail_policy(string) — behavior when SPARC is unreachable:open(default, allow + record) orclosed(block).skip_tools/reflect_tools([]string) — tool-name globs to exclude from, or restrict, reflection.bypass_hosts/bypass_paths([]string) — globs skipped without reflecting; default to Keycloak/SPIRE/otel/etc.
Swaps a placeholder credential for a real static credential on outbound requests, so the workload never holds the real secret.
source(string) —secret_dir(read one file per key fromsecret_dir) ormappings(inline map; tests/dev only).secret_dir(string) — directory of per-key credential files.mappings(map[string]string) — inline key-to-credential map; not for real secrets.key_by(string) —host(default, use the outbound destination host) orstatic(always usekey).key(string) — lookup key used whenkey_by=static.placeholder(string) — if set, the inbound bearer must exactly equal this value before injection proceeds.inject_header(string) — header to inject the credential into. DefaultAuthorization(writesBearer <value>); any other value writes the raw credential and drops the inboundAuthorizationheader.
Enforces per-session token, call-count, and duration budgets via Redis. Opt-in at build time (-tags include_plugin_sessionbudget). Full details in session-budget-plugin.md.
redis_url(string) — Redis/Valkey connection URL; required.max_tokens(int64) — cumulative token ceiling per session.0= no limit.max_input_tokens(int64) — per-kind ceiling for uncached prompt tokens.0= no limit.max_cache_read_tokens(int64) — per-kind ceiling for prompt tokens served from cache.0= no limit.max_cache_write_tokens(int64) — per-kind ceiling for prompt tokens written to cache.0= no limit.max_output_tokens(int64) — per-kind ceiling for generated completion tokens.0= no limit.max_reasoning_tokens(int64) — per-kind ceiling for reasoning-only output tokens (a subset of output).0= no limit.max_calls(int64) — max inference calls per session.0= no limit.max_duration_seconds(int64) — wall-clock session lifetime.0= no limit.on_exceed(string) —deny(default, block),observe(log only), orpause(HITL webhook approval).pause_webhook(string) — URL to POST for approval whenon_exceed=pause. Required in pause mode.pause_timeout(string) — how long to wait for webhook response. Default30s.pause_timeout_action(string) — fallback on timeout/error:deny(default) orallow.pause_grace_period(string) — suppress repeated webhooks after approval. Default5m.session_ttl_seconds(int) — Redis key TTL; must be ≥max_duration_secondswhen the latter is set (enforced at Configure time). Default 7200.refresh_interval(string) — how often the local cache syncs from Redis. Default5s.redis_unavailable(string) — onlyfail_open(default) is implemented;fail_closedis rejected at Configure time.default_session_fallback(bool) — pool sessionless traffic into a shareddefaultbucket. Single-workload only: one caller exhausting the budget denies the rest. Defaultfalse.
At least one of max_tokens, the five per-kind ceilings, max_calls, or
max_duration_seconds must be > 0, or Configure fails.
Cold-cache behavior is mode-dependent; see session-budget-plugin.md for details.
Exchanges incoming tokens against a configured IdP through an external
token broker service, per host-based routing rules. An alternative to
token-exchange, not a complement — both replace the
outbound Authorization header, so use one or the other on a given
chain. Full details in token-broker-plugin.md.
broker_url(string) — base URL of the token broker service; required.default_policy(string) — behavior when no route matches:passthrough(default) orbroker.routes.file(string) — path to aroutes.yamlfile; merged with inline rules.routes.rules(list) — inline route entries; each has:host— glob pattern to match the target host.action—broker(default) orpassthrough.authorization_endpoint/token_endpoint— per-route OAuth endpoint overrides sent to the broker.
RFC 8693 outbound token exchange per route. Supports Keycloak, Entra
ID, Okta, and any RFC 8693-compliant IdP. For the IdPProvider
interface each IdP implements, see
idp-plugin-contract.md.
token_url(string) — OAuth token endpoint; required unless derived fromprovider+provider_url(+provider_realm), or the deprecatedkeycloak_url/keycloak_realm.provider(string) — IdP selector for endpoint derivation and client auth:keycloak,generic.provider_url/provider_realm(string) — IdP base URL and realm/tenant, meaning varies by provider.keycloak_url/keycloak_realm(string) — deprecated aliases forprovider_url/provider_realmwithprovider=keycloak.default_policy(string) — behavior when no route matches:passthrough(default) orexchange(empty-audience client-credentials) for hosts explicitly configured inauthproxy-routes.no_token_policy(string) — behavior for outbound requests with no bearer token:client-credentials,allow, ordeny(default).identity.type(string) —spiffe(JWT-SVID assertion) orclient-secret; required.identity.client_id/identity.client_id_file— OAuth client ID, inline or from file (default/shared/client-id.txt).identity.client_secret/identity.client_secret_file— client secret, inline or from file (default/shared/client-secret.txt).identity.jwt_audience(string) — audience claim minted on the JWT-SVID assertion; required whentype=spiffe.identity.assertion_type(string) — client-assertion URN:jwt-spiffe(default) orjwt-bearer(Okta).routes.file(string) — path toroutes.yaml. Default/etc/authproxy/routes.yaml.routes.rules(list) — inline route entries (host,target_audience,token_scopes,token_url,action), combined with file-loaded routes.audience_from_host(bool) — derive audience from host for unrouted requests. Defaultfalse.resolve_placeholders(bool) — resolve an inbound placeholder-prefixed bearer to its real token before exchange; unresolvable placeholders are denied. Defaultfalse.
Removes unused tool definitions from the outbound inference manifest, so the tokens for tools an agent never calls are not billed on every turn. The manifest is assembled by the client, so the proxy is the only place to trim it without changing every client.
Requires inference-parser earlier in the chain, and must sit after any
body-reading plugin (it rewrites the request body). Declares
WritesRequestBody only, so response streaming is unaffected.
-
remove([]string) — tool names to delete from the manifest. The complete verdict: no learning, no state, no storage. Names absent from a given request are ignored. An empty list is the off switch — the plugin is inert until a name is added, which is how it ships in the local install. -
paths([]string) — request paths to act on, matched exactly or by suffix. Defaults to/v1/chat/completions,/v1/completions,/v1/messages. -
No rates and no pricing. This plugin reduces tokens; pricing the reduction belongs to whoever owns cost. It publishes what it removed — tool names, byte delta, and whether the prune was applied or only measured — and reads the priced figure back off the cost record to fill its
$ savedmetric. A figure derived from the bundled table is still labelledbundled, and a model with no rate anywhere is still counted in arequests unpricedrow rather than charged at another model's rate. There is still no output rate: pruning only shrinks the prompt.The saving has to be priced on the response side, not here: the dollar amount depends on which prompt-cache tier the removed tokens came out of — 1x, 1.25x or 0.1x of the same rate — and only the response reveals that. It is inherently a request-fact times a response-fact.
Generate the list from local transcripts with agentop tools scan, which
proposes only tools it recognises as Claude Code built-ins and never proposes
one it has seen called. --days N sets the recency window (30 by default) and
--all drops it; widening is the cautious direction, since a longer window
finds more tools in use and so proposes fewer for removal. With --write it
refuses when it observed no tool calls at all, because "tools you have not
called" would then mean every tool it knows. See
tool-prune-plugin.md for the measure-then-enforce
rollout, the metrics readout, and what the saving does and does not change.
Moved. The cost session-event record — its fields, and the rule that nothing in
avoided is spend — is in pricing.md.
Like pricing: below, cost is core rather than plugin: settle.Settle and the ledger,
aggregator and /v1/usage endpoint all live in core/, and a plugin only feeds them.
Moved. Model rates, gateway discounts, the shipped defaults and how to override them are
in pricing.md.
It lives there because pricing: is a top-level config section rather than a plugin — it
registers no plugin and this file catalogs "every plugin with a Go implementation that
calls plugins.RegisterPlugin()". The plugins that consume those rates
(inference-parser, litellm-budget-track, tool-prune) are still catalogued above.