diff --git a/.changeset/harness-ag-ui-bridge.md b/.changeset/harness-ag-ui-bridge.md new file mode 100644 index 0000000000..d4c9c853bc --- /dev/null +++ b/.changeset/harness-ag-ui-bridge.md @@ -0,0 +1,13 @@ +--- +'@tanstack/ai-harness': minor +--- + +New subpath `@tanstack/ai-harness/ag-ui`: the AG-UI bridge. A harness session already streams AG-UI events, so this is a thin, documented seam for AG-UI clients. + +- `sessionEventsToAgUi(events, options?)` normalizes a session's `SessionEvent` stream into a pure AG-UI event stream: run usage is surfaced under `metadata.tanstack.usage` (`normalizeUsage` handles both the AG-UI spec array and the TanStack prompt/completion shapes), interrupt/approval waits stay as `RUN_FINISHED` with `outcome.type === 'interrupt'`, and subagent attribution is preserved. It can optionally drop the harness-native `CUSTOM` control events or emit interim `tanstack.spend` `CUSTOM` ticks for a live spend meter. +- `createAgUiHandler(options)` is a `fetch` handler a bare `@ag-ui/client` `HttpAgent` can point at. `POST` a `RunAgentInput` to run a prompt, or one with `resume` entries to answer the last turn's interrupts. It streams AG-UI events as SSE via `@ag-ui/encoder`'s `EventEncoder` (protobuf framing when the client's `Accept` prefers it). The stream is strict AG-UI by default (harness control events dropped so it begins with `RUN_STARTED`); the harness-native approval path and the relay dashboard are unchanged. +- `operationToAgUiRun(entries, options)` coalesces one harness operation — which may span several model turns (each a `RUN_STARTED`/`RUN_FINISHED` pair) — into a single valid AG-UI run: one `RUN_STARTED` (synthesized when a resumed run leads with a tool result), one terminal `RUN_FINISHED` carrying the operation outcome and summed usage, and `RUN_FINISHED.usage` conformed to the AG-UI `SpecTokenUsage[]` shape. This is what makes the stream consumable by a strict `@ag-ui/client` without "run already finished" / usage-shape errors. `createAgUiHandler` uses it per request. + +The AG-UI wire is pinned: built and tested against `@ag-ui/core@1.0.0`, `@ag-ui/encoder@1.0.0`, and `@ag-ui/client@1.0.0`. + +Generated with Claude Code. diff --git a/.changeset/harness-p10-goal.md b/.changeset/harness-p10-goal.md new file mode 100644 index 0000000000..bf27c56dca --- /dev/null +++ b/.changeset/harness-p10-goal.md @@ -0,0 +1,7 @@ +--- +'@tanstack/ai-harness': minor +--- + +Add the `goal({ judge })` plugin to `@tanstack/ai-harness/plugins`. `/goal ` sets a goal and starts a turn. After each turn, the `judge` model reads the goal and the end of the transcript and decides if the goal is met. If it is not met, the plugin starts the next turn, up to `maxRounds` turns (20 by default). The loop also stops when a turn waits for approval or fails, and when the user sends a message. `/goal` shows the status, `/goal stop` ends the goal, and `/goal resume` continues it. The goal is plugin state, so it survives a restart. The `GoalMet` event tells other plugins and clients that the goal is met. + +A plugin's `ctx.session.prompt(text)` now returns the turn, so the plugin can cancel it before it starts. The turn always waits in the queue, also on a harness with `busy: 'reject'`. diff --git a/.changeset/harness-p11-session-view.md b/.changeset/harness-p11-session-view.md new file mode 100644 index 0000000000..845ea5b33b --- /dev/null +++ b/.changeset/harness-p11-session-view.md @@ -0,0 +1,7 @@ +--- +'@tanstack/ai-harness': minor +--- + +Add `createSessionView(session | client)` at `@tanstack/ai-harness/view`. It keeps a live TanStack Store of everything a UI shows: messages with streaming text, tool calls, and child agents, plus approvals, questions, sign-ins, status, background agents, commands, settings, tools, and plugin state. It has actions (`send`, `command`, `setConfig`, `cancel`, `approve`, `reject`) and typed events (`view.on('approval', ...)`, `view.on(GoalMet, ...)`). It works with any UI library through the TanStack Store adapters or `store.subscribe`, and it works in the browser with `createHarnessClient`. + +New reads for it: `session.transcript()`, `session.describe()`, `snapshot().plugins`, and on the client `transcript()`, `describe()`, `answer()`, `command()`, `setConfig()`, and `events({ onConnection })`. `createHarnessHandler` serves `GET .../transcript` and `GET .../describe`. `selectGoal(state)` reads the goal from a view. diff --git a/.changeset/harness-p12-cli.md b/.changeset/harness-p12-cli.md new file mode 100644 index 0000000000..148499ae0c --- /dev/null +++ b/.changeset/harness-p12-cli.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-harness-cli': minor +--- + +The CLI has no UI library now: it does not depend on `ink` or `react`. An interactive terminal uses line mode, which now prints from a session view and opens sign-in links in the browser. `runCli(harness, { ui })` runs your own screen with any TUI library: `ui` gets a ready `createSessionView` view and resolves when the user quits. An Ink screen is in `examples/harness-cli/src/tui.tsx`. diff --git a/.changeset/harness-p13-agent-middleware.md b/.changeset/harness-p13-agent-middleware.md new file mode 100644 index 0000000000..ea64d5500d --- /dev/null +++ b/.changeset/harness-p13-agent-middleware.md @@ -0,0 +1,8 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-harness': minor +--- + +`@tanstack/ai`: the chat middleware context has `subagentName`, the agent name of a child that runs through `ctx.chat`, and `parentSubagentRunId`, the id of the child that started a nested child. `chat()` takes both as options. A child that calls `ctx.chat({ subagents })` now gives its own children the host middleware (`subagents.binding.chatMiddleware` and `generationMiddleware`) before their own, also when the binding has no budget. + +`@tanstack/ai-harness`: plugins can contribute `agentMiddleware`, chat middleware for every agent run: subagents the lead model calls (coding agents too), background agents, and their children. It does not run in the lead turn. Put the same middleware in `middleware` and `agentMiddleware` to see every model call. Run plugins' `generationMiddleware` now also reaches the agents of their turn. The first-party `usage()` plugin now counts the model calls of every agent, not only the lead turn. diff --git a/.changeset/harness-p14-mcp-server.md b/.changeset/harness-p14-mcp-server.md new file mode 100644 index 0000000000..91a5dceab1 --- /dev/null +++ b/.changeset/harness-p14-mcp-server.md @@ -0,0 +1,8 @@ +--- +'@tanstack/ai-mcp': minor +'@tanstack/ai-harness-cli': minor +--- + +`@tanstack/ai-mcp`: add `createHarnessMcpServer` at `@tanstack/ai-mcp/harness`. It serves a harness as an MCP server, so Claude Code, Claude Desktop, Cursor, or another agent can use it. The tools are `chat`, `steer`, `cancel`, `approve`, `reject`, `resolve`, `answer`, and `status`, plus `agent_` for each exposed agent and `command_` for each plugin command. Every tool takes an optional `threadId`. With `approvals: 'ask'` (the default), the client asks the user about each approval when it supports elicitation. Otherwise the approvals come back in the result. `approvals: 'auto'` approves every tool call. Each interrupt in a result has a `kind`: `approval`, `client-tool`, or `generic`. `resolve` answers every kind in one call: `approved` for an approval, and `payload` for the others. + +`@tanstack/ai-harness-cli`: `--mcp` serves the harness as an MCP server over stdio, and `--yes` approves every tool call in MCP mode. `--serve` also serves MCP at `/mcp`, behind the same bearer token. Both need `@tanstack/ai-mcp`, which is an optional peer dependency. diff --git a/.changeset/harness-p9-coding-agents.md b/.changeset/harness-p9-coding-agents.md new file mode 100644 index 0000000000..62a57e4b40 --- /dev/null +++ b/.changeset/harness-p9-coding-agents.md @@ -0,0 +1,13 @@ +--- +'@tanstack/ai-sandbox': minor +'@tanstack/ai-harness': minor +'@tanstack/ai-harness-cli': minor +'@tanstack/ai-acp': minor +'@tanstack/ai-dashboard': minor +--- + +`@tanstack/ai-sandbox/harness` adds `codingAgents({ sandbox, agents, workspace? })`: a harness plugin that gives the lead model one tool per coding agent (Claude Code, Codex, Grok Build, or any ACP agent). Each agent runs in the sandbox, keeps its own session per thread (also after a restart), and starts read-only in the harness `plan` mode. `workspace: 'shared'` (default) runs one agent at a time in one sandbox per thread. `'per-agent'` gives each agent its own sandbox. `/fresh [agent]` starts new sessions. + +`@tanstack/ai-harness` plugins can contribute `subagents`: agents the model can call as tools. + +Child agent work is now visible: the CLI shows each child's tool calls and a finish line with the start of its answer, ACP editors get the child's tool calls, and the dashboard shows a block per child. diff --git a/.changeset/harness-system-preamble.md b/.changeset/harness-system-preamble.md new file mode 100644 index 0000000000..728e4c9180 --- /dev/null +++ b/.changeset/harness-system-preamble.md @@ -0,0 +1,17 @@ +--- +'@tanstack/ai-harness': minor +--- + +Add an additive `systemPreamble` to the `prompt` input op. When present, its +strings are prepended (ahead of the harness's own `systemPrompts`) as +system/developer messages for that one run — a place for a trigger to attach +per-run context (e.g. operational memory) without the agent author doing +anything. + +Also make the server-tool execution `context` always carry the live `threadId` +and `runId` (merged over the harness's static `context`) so a tool invoked +in-band (from a model turn) can resolve the calling thread, matching the flat +`{ threadId, runId, signal }` already passed to out-of-band `{ op: 'tool' }` +invocations. + +Generated with Claude Code. diff --git a/.changeset/harness-tool-op.md b/.changeset/harness-tool-op.md new file mode 100644 index 0000000000..2cc4cadbf3 --- /dev/null +++ b/.changeset/harness-tool-op.md @@ -0,0 +1,12 @@ +--- +'@tanstack/ai-harness': minor +--- + +Out-of-band tool invocation: a new `{ op: 'tool', name, args?, meta? }` session input runs one registered tool with no model turn. This is the provisional harness surface the agent-dashboard "injection" work builds on (schedules, run-now, webhooks), following the `@tanstack/ai-harness/ag-ui` precedent of shipping dashboard-facing seams here. + +- `session.tool(name, args?, meta?)` (and the `{ op: 'tool' }` client input via `applyInput`) resolves the tool from the harness + session-plugin tools, validates `args` against its input schema, and invokes its server executor directly — modeled on the existing `command` op. The tool's lifecycle is published into the session feed as a normal AG-UI run (`RUN_STARTED`, `TOOL_CALL_START`/`ARGS`/`END`, `TOOL_CALL_RESULT`, `RUN_FINISHED`), so watchers render and persist it exactly like a tool call the model made. `meta` (e.g. an injection trigger) is echoed onto the run as a `tanstack.injection` `CUSTOM` event. Injected tools are fire-and-forget (no approval interrupt, like commands). +- Tool **visibility**: a new `toolVisibility?: Record` on `defineHarness`. Tools default to `private` — only tools named `public` may be invoked out-of-band; unknown or private tools are rejected (`unknown_tool` / `not_public`). Visibility is reported per tool in the capabilities document (`capabilitiesOf().tools.items[].visibility`). + +Additive and backward compatible: existing ops, tools, and streams are unchanged. + +Generated with Claude Code. diff --git a/PODS-PROTOCOL-PROPOSAL.md b/PODS-PROTOCOL-PROPOSAL.md new file mode 100644 index 0000000000..f4bd7998c3 --- /dev/null +++ b/PODS-PROTOCOL-PROPOSAL.md @@ -0,0 +1,151 @@ +# Pods/Teams — harness protocol proposal (for @AlemTuzlak) + +**Status:** proposal · **Date:** 2026-09-26 · **Author:** Jack + Claude Code +**Basis:** `~/Downloads/pods-design-doc.md` (v2) and `pods-implementation-plan.md` + +This is a **discussion doc, not a change**. No `packages/ai-harness` or +`packages/ai-dashboard` (relay) code has been touched. Phase 1 of the teams +reframe ships entirely in `examples/agent-dashboard` (dashboard-local +collections). Everything below is the net-new **harness/relay protocol surface** +that Phases 2+ need, grounded in what the current code does and doesn't do, so we +can agree on the shape before writing any of it. + +The design doc says "pod"; the dashboard ships the noun **"team"**. Same concept. + +## Why this doc exists + +Three capabilities the design depends on are not expressible in the harness as it +stands today. Each is core protocol surface, so it needs your review before +implementation. Findings were verified against the current tree (head at the top +of the harness PR stack). + +--- + +## 1. Out-of-band tool invocation (`injectToolCall`) + +**Design need.** The dashboard's injection model (design §5.4) calls a single +named tool deterministically, with **no model turn and zero tokens** +(`injectToolCall(pod, tool, args)`), and posts the result to the stream. This is +the *default* trigger path (timers, webhooks, "run now"). + +**Current reality.** Every tool call originates from a model `chat()` turn. The +input ops are fixed: + +``` +INPUT_OPS = ['prompt','steer','followUp','resolve','agent','cancel','command','answer','config'] + // packages/ai-harness/src/protocol.ts:37 +``` + +There is no op that runs a registered tool directly. The closest precedent is +plugin **commands** (`session.command(name, input)`, `session.ts:579`), which run +user-initiated actions outside a model turn and return a `Receipt` — this is the +shape to copy. + +**Proposal.** +- Add input op `{ op: 'tool', name: string, args: unknown }` to `INPUT_OPS` + (`protocol.ts:37`) and a case in `applyInput()` (`protocol.ts:119`) that + resolves the tool, validates `args` against its schema, invokes + `tool.execute(args)` **without** opening a `chat()` stream, and returns a + `Receipt` (mirror `executeCommand`). +- Publish the result into the feed under a fresh `operationId` so it renders in + the stream like any tool result (reuse `OperationImpl`/`SessionFeed`). + +**Open questions for you.** Should this reuse the command machinery outright +(register injectable tools *as* commands) rather than a parallel op? What runs the +tool's `needsApproval` gate on the injection path — does an injected tool that +needs approval still raise an interrupt? + +## 2. Tool registry + visibility (`public` vs `private`) + +**Design need.** A pod tool registry the dashboard can query, with visibility: +`public` tools are pod-invocable (dashboard, other agents, schedules); `private` +tools run only in the owning agent's own runs (design §5.2). + +**Current reality.** Tools are static per harness (`HarnessConfig.tools`, +`define.ts:36`), collected per-run from harness + plugins (`session.ts:1080`). +`expose` exists but only for **agents**, not tools: + +``` +expose?: { agents?: ReadonlyArray<...> } // define.ts:58 +``` + +`capabilitiesOf()` lists harness-level tool names only (`protocol.ts:197`), with +no visibility concept and no plugin/discovered tools. + +**Proposal.** +- Add `expose?: { tools?: ReadonlyArray }` to `HarnessConfig` + (`define.ts`), defaulting to private (owner-only). +- Add `session.tools()` mirroring `session.commands()` (`session.ts:562`). +- Have `capabilitiesOf()` (`protocol.ts:187`) return the visibility-filtered set, + so the dashboard registry query and the §1 injection path share one source of + truth. + +## 3. Injection to offline hosts vs. the relay constraint + +**Design need.** "Injected tool calls execute on the owning agent's host; offline +hosts use the relay's existing queued-input behavior rather than silent drops" +(design §5.4). + +**Current reality — this is a real conflict, flagging it explicitly.** The +offline queue exists, but it is **private** to the relay: + +``` +const sendToHost = (host, envelope): 'sent' | 'queued' => { + if (host.streams.size === 0) { host.queue.push({ envelope, expiresAt: ... }); return 'queued' } + ... +} // packages/ai-dashboard/src/server.ts:160 +``` + +`sendToHost()` is called only from `/api/sessions/:host/:thread/input` and +`/open`. There is **no external API to enqueue** for an offline host, and the +queue is drained only on the host's `/api/host/stream` reconnect +(`server.ts:253`). So: + +- Adding an injection route that reuses the queue **requires modifying the relay** + (add a route that calls `sendToHost()`), which collides with the hard "don't + modify the relay" constraint. +- Duplicating the queue outside the relay **does not work**: the relay flushes + only its own queue on reconnect, so externally-queued frames would never be + delivered. + +**Proposal / decision needed from you.** Pick one: +- **(a)** Accept a minimal, additive relay change: one new frame type + (`harness.inject`) enqueued through the existing `sendToHost()` path. Smallest + possible surface; keeps offline semantics correct. (Recommended — the + constraint exists to protect the relay's protocol, and this is protocol we'd be + co-designing with you rather than a unilateral edit.) +- **(b)** Keep injection **online-only** for now (dashboard injects only to hosts + with a live stream; offline injection is deferred). No relay change; weaker + guarantee. + +## 4. Per-event causal metadata (loop-TTL bounce protection) + +**Design need.** Bounce protection (design §6.3) leads with a **causal depth +TTL**: every event carries its causal chain and the pod caps reaction depth. That +requires per-event provenance. + +**Current reality.** Only operation-level lineage exists — `parentRunId` on turns +for agent resume chains (`session.ts:163`), passed to `chat({ parentRunId })`. +`SessionEvent` itself (`feed.ts`) carries `{ cursor, operationId, event }` — no +`parentEventId`, no `causedByInputId`, no `depth`. + +**Proposal.** +- Extend the feed/`SessionEvent` with `parentEventId?`, `causedByInputId?`, and a + monotonically-increasing `depth` set when an operation is spawned in reaction to + another event. +- Enforcement (TTL cap, cycle detection, quarantine) stays in the + dashboard/harness layer per the design; this proposal is only about **carrying** + the provenance so enforcement is possible later. + +**Open question.** Is `parentRunId` enough to derive depth for the agent-to-agent +case, or do we genuinely need event-level parentage (I believe we do, because a +single run reacts to many upstream events)? + +--- + +## Phasing implication + +Phases 2–7 of the implementation plan are all gated on §§1–4. Phase 1 (the +dashboard-local teams reframe) is done and needs none of this. Recommend a short +review pass on §§1–3 first (they unblock the injection + registry work in Phase +2), with §4 reviewed alongside Phase 4 (bounce protection). diff --git a/STATUS.md b/STATUS.md new file mode 100644 index 0000000000..f99fbf7d42 --- /dev/null +++ b/STATUS.md @@ -0,0 +1,437 @@ +# Agent Dashboard — build status + +Spec: `~/Downloads/agent-dashboard-spec.md` (original four phases), then the +**teams reframe** (`~/Downloads/pods-design-doc.md` + `pods-implementation-plan.md`), +then the **Reddit pod** (`~/Downloads/teams-reddit-pod.md`). +Everything below is committed and verified; nothing is pushed. + +> **Latest work: DM membership + team-page memory fix** (`cf119c4`) — DMs were +> created with no members (Zod stripped `members` from `pod.channel_create`), so +> messages dropped; now "New DM with X" is a working 1:1 with that agent. Pod +> memory moved from the devtools onto the team page (one panel per agent), so +> remembered entries are visible directly. Before that, the **product +> team-composition UI + default subscriptions** +> (`bb4ff79`) — the home page is now an agents table with **Add to team**, the +> team page has **+ Add agent** and a per-agent **🔧 run-tool** dialog, and the +> seeded-team launchers moved into the Demo Controls devtools panel. Agents +> carry default subscriptions by harness, so a hand-composed team reacts like a +> seeded one. See the "Product team-composition UI" section below. Before that, +> the **Reddit pod** (`3a30d2f`) — the first pod wired to a real external +> service (`reddit.search_react_news` reads Reddit's public RSS) and a real LLM +> (`sentiment/react` on Anthropic). The demo-controls devtools panel (`26734dc`) +> and everything under it +> still stand. + +## Where this lives + +- **Worktree**: `~/projects/tanstack/ai-workingtrees/ai-dashboard-work` +- **Branch**: `feat/agent-dashboard` +- **Base**: `origin/feat/harness-p8-code-mode` (top of the harness PR stack + #1509→#1524, head `50ab70e3e`) — the real `TanStack/ai` monorepo, **not** the + standalone `tanstack-agent-dashboard` POC. +- **Do not touch** the stacked `feat/harness-p*` branches or the relay in + `packages/ai-dashboard` (spec §6). + +## Commits (atop the base) + +| Commit | What | +|--------|------| +| `7da107f` | `feat(ai-harness)`: the `@tanstack/ai-harness/ag-ui` bridge subpath | +| `d6c538a` | `fix(ai-harness)`: coalesce multi-turn operations + conform usage for strict AG-UI clients | +| `78a5ba8` | `feat(examples/agent-dashboard)`: live session view + approval queue (Phase 2) | +| `d28d9de` | `feat(examples/agent-dashboard)`: control plane — history, spend, config (Phase 3) | +| `0940975` | `feat(examples/agent-dashboard)`: meta-chat over live agent state (Phase 4) | +| `063245e` | `docs`: add STATUS.md summarizing the agent-dashboard build | +| `9ec1ece` | `feat(examples/agent-dashboard)`: teams reframe (Phase 1) + Alem protocol proposal | +| `9b1db82` | `feat(ai-harness)`: out-of-band tool invocation (`{ op: 'tool' }`) + tool visibility | +| `720660d` | `feat(examples/agent-dashboard)`: teams Phase 2 — tool registry + injection | +| `04639fd` | `feat(ai-harness)`: `systemPreamble` on the prompt op + in-band tool thread id | +| `4101e64` | `feat(examples/agent-dashboard)`: teams Phase 3 — system tools, channels, pod memory | +| `66dd71b` | `feat(examples/agent-dashboard)`: persist agent + team state on the server | +| `26734dc` | `feat(examples/agent-dashboard)`: move demo controls into a devtools panel | +| `3a30d2f` | `feat(examples/agent-dashboard)`: Reddit pod — real service + real LLM | +| `bb4ff79` | `feat(examples/agent-dashboard)`: product team-composition UI + default subscriptions | +| `cf119c4` | `fix(examples/agent-dashboard)`: DM membership + surface pod memory on the team page | + +## Phase status + +- **Phase 1 — AG-UI bridge** ✅ `packages/ai-harness/src/ag-ui.ts` + - `sessionEventsToAgUi()` (normalizer), `operationToAgUiRun()` (coalesces one + harness operation → one valid AG-UI run), `createAgUiHandler()` (SSE via + `@ag-ui/encoder`). Consumable by a bare `@ag-ui/client` `HttpAgent`. + - Interrupts stay as `RUN_FINISHED.outcome`, usage → `metadata.tanstack.usage` + (+ conformed to `SpecTokenUsage[]`), subagent attribution preserved, optional + `tanstack.spend` ticks. + - 22 unit tests (`tests/ag-ui.test.ts`), changeset, `docs/harness/ag-ui.md` + (kiira passes, registered in `docs/config.json`). AG-UI pinned at `1.0.0`. +- **Phase 2 — Dashboard shell + live session view** ✅ `examples/agent-dashboard` + - TanStack Start; routes hosts → sessions → session detail. + - Session view consumes AG-UI SSE via `@ag-ui/client`, projected into TanStack + DB `localOnly` collections read with `useLiveQuery` (no polling). + - Approval queue: approve → AG-UI resume, deny → harness control tier, edit → + edited tool args. +- **Phase 3 — Control plane** ✅ + - History (`/history`, `runs.listByThread`) + session replay (`/api/replay`). + - Spend (`/spend`): per-session token rollups (live query) + budgets/alerts. + - Config (`/config`): form from `ConfigOption` schemas → writes via the harness + protocol (`op: 'config'`). Endpoints stamped `HARNESS_PROTOCOL_VERSION`. +- **Phase 4 — Meta-chat** ✅ `/chat`, harness `dashboard/meta` + - Tools over live state: `list_agents`, `list_sessions`, `query_runs`, + `get_agent_config`, `set_agent_config`, `summarize_session`. + - Runs on the same host (harness registry); its tool calls are visible inline + and its runs appear in History like any other agent. + +## Teams reframe (Phase 1) ✅ `9ec1ece` + +A product-model reframe on top of the four-phase build: `hosts → sessions` +becomes **teams → channels**. A team is a group of agents sharing a chat, a tool +registry, and a permission boundary (the design doc calls it a "pod"; the UI noun +is **"team"**). One agent looks like a plain chat; a **second member reveals the +team** — roster appears, per-agent attribution shows up, and both agents' AG-UI +streams merge into one channel. **Dashboard-local only — zero harness/relay +changes.** + +- **Key trick:** each member owns its own thread, so `agentId == threadId`. + Namespacing every projected row id by `agentId` is therefore per-thread, which + lets the old threadId-keyed `/sessions/$threadId` and `/chat` routes keep + working **unchanged** while the new `/teams/$teamId` route queries by + `channelId` and gets the multiplex. +- `src/db/collections.ts` — `teams`/`channels`/`memberships` localOnly + collections; `channelId`/`agentId` (+ raw `interruptId`) on the existing rows. +- `src/lib/session-controller.ts` — `project(ctx)` namespaces ids + (`${agentId}:${raw}`) so N members share one channel with no collisions; + `joinChannel`/`subscribeMember`, `createTeam`/`addAgentToChannel`, + `hydrateMember`; back-compat thread-keyed wrappers; interrupt ids de-namespaced + for resume/control. +- `src/components/channel-view.tsx` + `member-list.tsx`; route + `src/routes/teams.$teamId.tsx`. Team chrome renders only at ≥2 members. +- `index.tsx` Teams section + "New team"; `__root.tsx` nav Hosts→Teams. Meta is a + **cross-team operator** — renders inside a channel, keeps its global tools. +- **Part B — `PODS-PROTOCOL-PROPOSAL.md`** (worktree root, for @AlemTuzlak): the + net-new harness surface Phases 2+ need — out-of-band tool op, injection vs the + relay's private offline queue (`server.ts:160`, a real conflict flagged), and + per-event causal metadata. Proposal only; no harness code changed. +- **Verified:** `tsc` + `oxlint` clean; **6 Playwright e2e pass** (the 5 existing, + unchanged, + new `e2e/team.spec.ts`: second agent reveals the team, both share + one channel). +- **Known cosmetic:** `/` now runs a live query, so SSR falls back to client + rendering (`useLiveQuery` has no `getServerSnapshot`) — same class as the other + live-query routes; page works, tests green. +- **Deferred:** system tools + channels + memory (Phase 3), bounce protection + (Phase 4), polyglot tool face (Phase 5), bridging (Phase 6), install model + (Phase 7), hibernation/policy/pricing/versioning (Phase 8). + +## Teams Phase 2 — tool registry + injection ✅ `9b1db82`, `720660d` + +The dashboard invokes work **deterministically** — scheduled timers, run-now, and +webhooks — with the structured result streaming into the team channel and **zero +LLM tokens** on the trigger path. "The dashboard owns the clock," made concrete. + +Note the constraint change from Phase 1: the Alem review is now a **heads-up, not +a gate**, so provisional harness changes ship on `feat/agent-dashboard` (the +`7da107f` AG-UI-bridge precedent). + +- **Harness (`9b1db82`, additive):** a new `{ op: 'tool', name, args?, meta? }` + input runs one registered tool with no model turn — `session.tool()` / + `executeTool()` modeled on the `command` op. It publishes + `RUN_STARTED`/`TOOL_CALL_*`/`RUN_FINISHED` into the feed, so injected results + ride the existing projection path and persist for replay. Tool **visibility** + (`toolVisibility` on `defineHarness`, default private) — only `public` tools may + run out-of-band; unknown/private are rejected. 180 harness tests pass. **Rebuild + the dist after harness edits** (the example imports the built package). +- **Architecture:** trigger via the control plane, observe via a **live feed + tail** (`/api/tail`, replays then follows live). `channel-view` opens one tail + per member as the single projector; `runAgent` is trigger-only. Interactive + runs, injected tools, timers, and webhooks all render through one + `project()`/`ToolCard` path. Back-compat `/sessions` and `/chat` are unchanged. +- **Server (owns the clock):** `server/scheduler.ts` (1s interval, lazy-boot), + `server/injection.ts` (`runInjection` → the tool op; server-owned schedule + + webhook registries; idempotent jobs; 64KB result cap), `server/cron.ts` (5-field + cron + `everySeconds`). Routes: `api.inject`, `api.schedules`, `api.webhooks` + (+ `api.webhooks.$token` ingress — a single-segment param so it doesn't shadow + `/api/webhooks`), `api.tail`, `api.tools` (public-only registry, enforced at + read time), `api.dev.offline`. Public demo tool `fetch_stats` on triage. +- **Offline** hosts are **simulated** dashboard-side (a pending queue + dev + toggle) because the example embeds the host — the relay and its private queue + are untouched. +- **UI:** an Automations panel (public tools + run-now, schedule table, + send-test-webhook, offline toggle + "host offline — N queued" banner); injected + tool cards carry a distinct trigger badge. +- **Deviation from the Phase 2 doc §4.1:** schedules/webhooks are **server** state + (not `localOnly`) because the server owns the clock, so it owns the table. +- **Verified:** `tsc` + `oxlint` clean; **10 Playwright e2e pass** (6 prior + 4 + new: run-now/private-hidden, timer, webhook, offline queue+flush). + +## Teams Phase 3 — system tools, channels, pod memory ✅ `04639fd`, `4101e64` + +The design doc's **"the pod learns" loop** made concrete: a webhook opens a per-PR +channel, a subscribed security agent reviews it, the human corrects it, the +correction is persisted to **pod memory**, and the *next* PR is handled better — +every step a message or a tool call in the stream, **no hidden state**. + +- **Harness (`04639fd`, additive):** the `prompt` op gains `systemPreamble?: + string[]`, prepended ahead of the harness's own system prompts for one run — the + seam a trigger uses to attach per-run memory. The server-tool execution + `context` now always carries the live `threadId`/`runId`, so an **in-band** + `pod.*` call can resolve its caller (out-of-band already got a flat + `{ threadId }`). 184 harness tests pass; changeset included. +- **System tools (`pod.*`), dual citizens:** `pod.channel_create` / + `pod.message_post` / `pod.memory_write` / `pod.memory_read` — public, callable + in-band during a run *and* out-of-band via `{ op: 'tool' }`. channel_create / + message_post are **structured intents** the dashboard realizes when it observes + them on the tail (create the channel / post the message, routed by the result's + channel id); memory_write mutates a server-side, thread-keyed store + (`server/memory.ts`, `server/systools.ts`). +- **Memory delivery:** `/api/run` is the memory-attaching run trigger — every run + (interactive prompt, subscription dispatch, prompt-mode webhook) prepends the + thread's pod memory as a `systemPreamble`. The platform attaches it; the agent + author does nothing. `/api/memory` backs the Memory panel. +- **Dynamic channels + subscriptions:** `channels` gains `kind`/`topic`/ + `createdBy`; new `channelMembers` (per-channel opt-in) vs the durable team + roster; `memberships` gains `teamId` + `subscriptions`. `project()` routes + system-tool results into channels/messages; client-side dispatch joins + subscribers and triggers a review on `channel_created`. +- **UI:** channel sidebar, `channel_created` / joined system cards, a per-agent + Memory panel, and an "N memory entries attached" run badge. +- **Demo:** `ops/pr-watcher` + `security/review` scripted harnesses and a seeded, + stateful `github.check_pr`; "+ PR-watcher demo" and "Send PR webhook" entry + points. The security model flags public-internet exposure **unless** its + attached memory says the deployment is intranet-only. +- **Server learns each thread's harness:** the client passes `harness` on the + tail/run/inject/webhook calls (`noteThread` adopts it), so a team can mix + harnesses (watcher, reviewer, meta) instead of defaulting everything to triage. +- **Deviations (deliberate):** the watcher opens a review channel it does **not** + join (it posts via `pod.message_post`, routed by channel id); pod memory is keyed + by `threadId` (unique per member, so `teamId` is redundant for the store); + `pod.*` are excluded from the run-now registry (plumbing, not automations). +- **Verified:** `tsc` + `oxlint` clean; **13 Playwright e2e pass** (10 prior + the + full §7 loop, DM creation, memory panel add/remove). + +## Persistence — durable across a restart ✅ `66dd71b` + +The dashboard was all in-process memory; a restart wiped every team. Now state is +durable so you can close the tab, restart the server, and check in on each team. + +- **`src/server/store.ts`** — one JSON snapshot file, written through on every + mutation (debounced, atomic tmp+rename) and replayed on boot. `filePersistence()` + keeps the in-memory reference backend (so all store-contract semantics are + inherited), replays the snapshot into it, then mirrors the four chat state stores + (messages, runs, interrupts, metadata) back to disk. A `FileMap` subclass gives + the side tables (threads, pod memory, schedules, webhooks) write-through with no + route changes. +- **Team roster** — teams/channels/memberships are client-only `localOnly` + collections, so they reset on reload even though the underlying threads persist. + Persisted as a server blob (`GET|POST /api/roster`), POSTed after each roster + mutation, seeded on app load (`hydrateRoster` in `__root`). +- **Not persisted:** the in-flight injection queue, the dev offline toggle, and + inbox/credentials (secrets do not belong in a plaintext file). Single-process, + single-file; a multi-node dashboard swaps a real DB behind the same seam. +- **Verified:** state round-trips across a fresh process; e2e isolates its state to + a wiped throwaway file; `tsc` + `oxlint` clean, **13 e2e pass**. + +## Demo controls — moved to a devtools panel ✅ `26734dc` + +The demo-only scaffolding was interleaved with the real UI, so it wasn't obvious +what drives the demo vs. what an operator would actually use. It now lives in a +custom **Demo Controls** TanStack DevTools panel; the channel view keeps only the +product experience (timeline, roster, approval cards, message box). + +- **Moved:** Start triage demo, + Add agent / + Add operator, the Automations + panel (run-now, schedules, webhook tester, offline sim), and the pod Memory + panel — all now in `src/components/demo-controls.tsx`, registered as a devtools + plugin in `__root.tsx`. +- **State-management stress test (the point):** the panel renders from the + devtools render root, **outside the route tree**, so it takes no props from the + route. `ChannelView` publishes the active channel to a new `uiState` `localOnly` + row; `DemoControls` re-derives the channel + primary member from the same live + TanStack DB collections the channel view reads. Two independent readers over one + source of truth. +- **Two wiring fixes it surfaced:** (1) `TanStackDevtools` moved **inside** + `QueryClientProvider` — the plugin is portaled but follows the React tree, and + its panels use `useQuery` ("No QueryClient set" otherwise). (2) The open panel + is a fixed bottom overlay that covers bottom-of-page controls (message box, + roster "run"); `VITE_E2E=1` opens it on load for specs, and `e2e/devtools.ts` + `openDemo`/`closeDemo` toggle it around those clicks. +- **Verified:** `tsc` + `oxlint` + `build` clean; **14 Playwright e2e pass** (13 + prior + the toggling in `team`/`pr-watcher`). + +## Reddit pod — real service, real AI ✅ `3a30d2f` + +The design doc's Reddit pod (`~/Downloads/teams-reddit-pod.md`): the graduation +from synthetic demo agents to **real external I/O + a real LLM**. Two new +harnesses, one new subscription event, no new UI. + +- **`reddit/fetcher`** (procedural, no LLM) — carries one real tool, + `reddit.search_react_news` (`src/server/reddit.ts`), which reads Reddit's + public **RSS (Atom)** feed with a real `User-Agent` (see finding #3 for why + RSS, not `.json`). Read-only by construction; public, so it's a run-now / + schedule / webhook target. The Atom parser is pure and unit-tested against a + recorded fixture (`reddit.test.ts`, `reddit-fixture.ts`) — the example gains a + `test:lib` target (vitest) for it. +- **`sentiment/react`** (real LLM) — `anthropicText('claude-haiku-4-5')`, gated + on `ANTHROPIC_API_KEY`. **Design decision (Jack): real LLM only** — no product + mock fallback. Without a key it posts a clear "set the key" message instead of + a digest; under `VITE_E2E` it uses a deterministic double so the e2e loop is + hermetic (no key, no network). This is the doc's "one documented manual step." +- **New wiring — the `tool_result` subscription:** `Subscription` grows one + event, `{ event: 'tool_result', tool, action: 'trigger' }`. When the fetcher's + tool result lands, `dispatchToolResult` (session-controller) runs the + subscribed sentiment agent with the capped (≤10-item) batch as a prompt — + memory attached via `/api/run`, exactly like `channel_created` trigger. The + fetcher's result also resolves the tool registry by explicit `harness` param + (`/api/tools`) so run-now lists the right tool without a note-thread race. +- **Loop:** timer → `injectToolCall` → result card → subscription → + `injectMessage` → LLM digest. Trigger path spends zero tokens. +- **Bounce (doc §7):** chat messages don't match the `tool_result` subscription, + and replays don't re-dispatch (`dispatchedResult` guard) — so no cycle. Verify + by running several cycles. +- **Demo entry point:** **+ React-news demo** button (`index.tsx`) → + `createReactNewsTeam`, mirroring **+ PR-watcher demo**. +- **Verified:** unit **4/4** pass; `tsc` + `build` clean; `oxlint` clean on + changed files; **15 Playwright e2e pass** (14 prior + `react-news.spec.ts`; + `meta-chat` count bumped 4 → 6 agents). Only my files touched (`pnpm format` + reformats unrelated committed files, so it was reverted off them). + +### Live test findings (with a real `ANTHROPIC_API_KEY`) + +Ran it against Jack's key + live Reddit. Two findings, one blocking-for-live: + +1. **Real LLM digest works ✅.** `sentiment/react` on `claude-haiku-4-5` + produced a proper per-item sentiment table + digest (overall vibe, hottest + thread, surprising) and wrote **2 pod-memory entries** via the `remember` + tool — the memory loop, live. +2. **Dotted `pod.*` tool names 400 on real providers 🐛 (fixed).** Anthropic + (and OpenAI) require tool names to match `^[a-zA-Z0-9_-]{1,128}$`; `pod. + memory_write` etc. 400 with `tools.0.custom.name: String should match…`. + The scripted mocks never validated names, so this only surfaced against a + real provider. **Fix:** `sentiment/react` carries a single provider-safe + `remember` tool (writes the same memory store) instead of `...podTools`; its + digest is plain text, so it needs no channel/message tools. + **Broader implication (follow-up):** any real-LLM agent carrying `pod.*` + tools hits this — sanitize tool names in the adapter's tool-converter, or + rename the `pod.*` tools to `pod_*`. Out of scope for this example. +3. **Reddit `.json` is IP-blocked here → switched to RSS ✅.** From this + machine's egress (Tailscale `utun5`), `reddit.com/.../*.json` returns **403 + for every User-Agent** (not the 429 the doc §3 anticipated) — an IP/datacenter + block, unfixable by UA changes. But the **Atom feed** (`/r/reactjs/new.rss`) + serves `200`, so the tool now reads RSS instead of JSON. Verified live from + this network: a real inject returned 5 current headlines with authors. RSS + carries title/link/author/timestamp/body — no score/comment counts, so those + are dropped from `RedditPost`. Reddit still rate-limits bursts (a rapid retry + `429`s); the 30-min schedule is well clear. E2E uses the recorded RSS fixture. + +## Product team-composition UI + default subscriptions ✅ `bb4ff79` + +The demo-era scaffolding graduated to real product controls, and hand-composed +teams now behave like the seeded demos (raised while dogfooding a Reddit team +built by hand). + +- **Home is an agents table** (`index.tsx`): name · description · **Add to + team** → start a new team with that agent, or add it to an existing one. +- **Team page** (`channel-view.tsx`): **+ Add agent** picks any available agent + (`/api/hosts`); each roster agent has a **🔧 tools** button opening a + **RunToolDialog** (`run-tool-dialog.tsx`) — pick a public tool + (`/api/tools?harness=`), pass JSON params, `runInjection`. +- **Seeded launchers moved to the Demo Controls panel** (`demo-controls.tsx`): + `+ New team` / `+ React-news demo` / `+ PR-watcher demo`. The triage-demo + button is gated to triage teams, so it no longer shows on other pods. +- **Default subscriptions by harness** (`session-controller.ts`, + `DEFAULT_SUBSCRIPTIONS`): `sentiment/react` → `reddit.search_react_news` + `tool_result`; `security/review` → `channel_created` join+trigger. + `addAgentToChannel` applies them when none are passed, so the demos and manual + composition share one source of truth. This is what makes a hand-composed team + react without wiring. +- **Two demo-era bugs fixed:** the roster **▶ run** sent a hardcoded triage + ticket prompt to *any* agent (→ neutral "Please proceed."); the **🔔 + subscribe** toggle set a hardcoded `channel_created` sub (→ toggles the + harness's real default triggers, shown only for reactive agents). +- **Note:** RunToolDialog takes params as **JSON**, not per-field inputs — + `/api/tools` doesn't expose the input schema and exposing it means touching + the published `ai-harness` protocol. Per-field inputs are a follow-up. +- **Verified:** `tsc` + `oxlint` + `build` clean; **16 Playwright e2e pass** + (15 prior + a hand-composed Reddit team reacting with no manual wiring; `team` + and `react-news` rewritten around the new controls; the moved-button specs + open the devtools panel first). The dev database was wiped for a clean run. + +### DM membership + team-page memory fix ✅ `cf119c4` + +Two bugs found while dogfooding a hand-composed team: + +- **Empty DMs.** `createDm` passed `members` to `pod.channel_create`, but the + tool's Zod schema didn't declare it → Zod stripped it → the tool never echoed + participants → the projector built a channel with no members, so it had no + `primary` and the operator's messages silently dropped. **Fix:** + `pod.channel_create` declares + echoes `members` (`systools.ts`), and **New DM + with X** now seats just the target agent — a working 1:1 that routes the + operator's messages to it (pod memory attached via `/api/run`). The DM e2e now + sends a message and asserts the agent replies. +- **Invisible pod memory.** The Memory panel lived only in the Demo Controls + devtools panel and rendered only the primary agent (e.g. `reddit/fetcher`, + which never writes memory), so a `sentiment/react` agent's `remember` entries + had nowhere to surface. **Fix:** a **Memory** section on the team page — + one panel per agent member (`channel-view.tsx`), removed from the demo panel. + +## Run it + +```bash +# from the worktree root +pnpm install +pnpm --filter @tanstack/ai-harness build # dashboard imports the built dist + +pnpm --filter agent-dashboard dev # http://localhost:3002 (no API key needed) +pnpm --filter agent-dashboard test:e2e # Playwright +``` + +Demo: **New team**, then open the **Demo Controls** devtools panel (trigger at +bottom-left) → **Start triage demo** → approve the drafted reply mid-run → +**+ Add agent** to reveal the team roster, then **▶ run** the second member and +watch both streams share one channel. In the panel's **Automations** section, +**run `fetch_stats` now**, **add a 1s schedule**, **send a test webhook**, or +**simulate the host offline** and watch jobs queue then flush. Then try +**Meta-chat**, **History**, **Spend**, **Config**. + +For Phase 3: **+ PR-watcher demo** → **Send PR webhook** (Demo Controls panel). A +`#pr-…` channel opens, +the security agent joins and flags the PR; reply **"not a security problem — we're +intranet-only here"**, watch it write pod memory, then **Send PR webhook** again — +the next PR is reviewed cleanly with the memory attached. + +For the Reddit pod: **+ React-news demo** → open the **Demo Controls** panel → +**▶ run `reddit.search_react_news`** (or add a 30-min schedule). A real batch of +React headlines lands as a result card, and `sentiment/react` posts a digest +unprompted. Set `ANTHROPIC_API_KEY` before `dev` for a live digest (without it, +the agent asks you to set the key). + +## Verification + +- `@tanstack/ai-harness`: **184 unit tests pass**; `tsc` / `oxlint` / + `publint --strict` clean. +- `examples/agent-dashboard`: **16 Playwright e2e pass** (approval mid-run, config + read/write, spend, history + replay, meta-chat, teams via the agents table + + +Add agent, injection ×4, the PR-watcher loop, DM creation, memory panel, + persistence reload, the Reddit pod loop, a hand-composed Reddit team reacting); + `tsc --noEmit` + `build` clean; parser **unit test** (vitest) pass. +- Base verified before starting: `pnpm build:all` (73/73), example agent runs and + emits AG-UI ndjson, `--serve` + relay boot. + +## Deliberate decisions / open questions (spec §8 — flagged for @AlemTuzlak) + +- **Chart**: spend uses an SSR-safe SVG bar chart, **not** `@tanstack/react-charts`. + Its 0.18 auto-sizing loops a ResizeObserver and freezes the main thread (the + POC hit this too and switched to visx). +- **Location**: `examples/agent-dashboard` (no `apps/` dir in the repo). +- **Auth**: local single-user (`authorize` returns a fixed principal). Multi-user + auth beyond pairing/host-token is a follow-up. +- **Spend**: tokens-first with per-session budgets as data; no canonical price + table yet (cost is a pluggable follow-up). + +## Known cosmetic issue + +- `@ag-ui/client` logs a warning stripping a nonstandard `/toolName` field the + harness core adds to `TOOL_CALL_START`. Warning only — the run is unaffected. + +## Not done + +- Not pushed to any remote (no push was requested). +- Deferred per spec §5: agent versioning, evals/red-teaming, prompt playground, + multi-host aggregation beyond the relay. diff --git a/docs/chat/subagents.md b/docs/chat/subagents.md index 7916787a8f..9da0c131fb 100644 --- a/docs/chat/subagents.md +++ b/docs/chat/subagents.md @@ -387,6 +387,8 @@ A child is its own `chat()` call. Put the child's middleware in that call. The p Inside a child, the middleware context has `subagentRunId`. Use it to tell a child run from a top-level run, and to link a child trace to its card. +When the child calls `ctx.chat`, the context also has `subagentName`, the name of the agent. For a nested child, `parentSubagentRunId` is the id of the child that started it. + On a routed turn where a child runs and main does not, the parent's middleware does not run. Only `withPersistence` records that turn. A turn where main runs, including a handoff, runs the parent's middleware as usual. ## Persistence diff --git a/docs/config.json b/docs/config.json index 920ebe6b79..fbfb0a95a6 100644 --- a/docs/config.json +++ b/docs/config.json @@ -165,7 +165,7 @@ "label": "Subagents", "to": "chat/subagents", "addedAt": "2026-09-21", - "updatedAt": "2026-09-26" + "updatedAt": "2026-09-28" }, { "label": "Agentic Cycle", @@ -879,7 +879,8 @@ { "label": "Run in the terminal", "to": "harness/cli", - "addedAt": "2026-09-26" + "addedAt": "2026-09-26", + "updatedAt": "2026-09-28" }, { "label": "Durable sessions", @@ -889,12 +890,14 @@ { "label": "Write a plugin", "to": "harness/plugins", - "addedAt": "2026-09-26" + "addedAt": "2026-09-26", + "updatedAt": "2026-09-28" }, { "label": "Build a coding agent", "to": "harness/coding-agent", - "addedAt": "2026-09-26" + "addedAt": "2026-09-26", + "updatedAt": "2026-09-28" }, { "label": "Auth and connectors", @@ -904,17 +907,40 @@ { "label": "Use MCP servers", "to": "harness/mcp", - "addedAt": "2026-09-26" + "addedAt": "2026-09-26", + "updatedAt": "2026-09-28" + }, + { + "label": "Use from any MCP client", + "to": "harness/mcp-server", + "addedAt": "2026-09-28" }, { "label": "Code mode in a harness", "to": "harness/code-mode", "addedAt": "2026-09-26" }, + { + "label": "Delegate to coding agents", + "to": "harness/coding-agents", + "addedAt": "2026-09-28" + }, + { + "label": "Work until a goal is met", + "to": "harness/goal", + "addedAt": "2026-09-28" + }, + { + "label": "Build your own UI", + "to": "harness/custom-ui", + "addedAt": "2026-09-28", + "updatedAt": "2026-09-28" + }, { "label": "Run agents from a harness", "to": "harness/subagents", - "addedAt": "2026-09-26" + "addedAt": "2026-09-26", + "updatedAt": "2026-09-28" }, { "label": "Deploy a harness", @@ -925,6 +951,11 @@ "label": "Self-host the dashboard", "to": "harness/dashboard", "addedAt": "2026-09-26" + }, + { + "label": "Stream a harness over AG-UI", + "to": "harness/ag-ui", + "addedAt": "2026-09-26" } ] }, diff --git a/docs/harness/ag-ui.md b/docs/harness/ag-ui.md new file mode 100644 index 0000000000..0239f59800 --- /dev/null +++ b/docs/harness/ag-ui.md @@ -0,0 +1,146 @@ +--- +title: Stream a harness over AG-UI +id: harness-ag-ui +order: 13 +description: "Serve a harness session as a pure AG-UI event stream a bare @ag-ui/client can consume, with approvals as interrupts and token usage as metadata." +keywords: + - tanstack ai + - harness + - AG-UI + - SSE + - interrupts + - usage +--- + +A harness session already speaks AG-UI: `session.events()` yields events whose payloads are AG-UI protocol events. The `@tanstack/ai-harness/ag-ui` subpath is the thin, documented seam an AG-UI client consumes — a normalizer and an SSE handler a bare [`@ag-ui/client`](https://docs.ag-ui.com) `HttpAgent` can point at. + +This does not replace the [harness protocol](./connect.md) or the [relay dashboard](./dashboard.md). AG-UI is the *run stream*; the harness protocol carries the *control plane* (pairing, snapshots, history, config). Use both: AG-UI for the conversation, the harness protocol for everything around it. + +## Serve a session over AG-UI + +`createAgUiHandler` is one `fetch` handler. `authorize` is required, so no endpoint is open by accident. + +```ts group=harness-ag-ui +import { defineHarness, createHarnessHost } from '@tanstack/ai-harness' +import { createAgUiHandler } from '@tanstack/ai-harness/ag-ui' +import { memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' + +const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), +}) + +const host = createHarnessHost({ persistence: memoryPersistence() }) + +export const handler = createAgUiHandler({ + host, + harness: assistant, + authorize: (request) => + request.headers.get('authorization') === + `Bearer ${process.env.HARNESS_TOKEN}` + ? { id: 'user-1' } + : null, + canAccess: (principal, threadId) => threadId.startsWith(principal.id), +}) +``` + +`POST` a `RunAgentInput` to it. A body with a trailing user message runs a prompt; a body with `resume` entries answers the last turn's interrupts and continues. The response is `text/event-stream` of AG-UI events, encoded with `@ag-ui/encoder` so a protobuf-accepting client gets binary framing for free. The stream closes when the run reaches a terminal state. + +## Consume it from a client + +Point a bare `@ag-ui/client` `HttpAgent` at the handler's URL: + +```ts group=harness-ag-ui-client +import { HttpAgent } from '@ag-ui/client' + +const agent = new HttpAgent({ + url: 'https://example.com/agent', + threadId: 'user-1/main', + headers: { authorization: `Bearer ${process.env.HARNESS_TOKEN}` }, +}) + +agent.addMessage({ id: 'u1', role: 'user', content: 'Summarize incidents' }) +await agent.runAgent() + +console.log(agent.messages.at(-1)) // the assistant's reply +``` + +The stream is strict AG-UI by default: it begins with `RUN_STARTED`, as a bare client requires. (The harness-native `CUSTOM` control events — `harness.operation.*`, `harness.question`, and friends — are dropped from the AG-UI stream; read them over the harness protocol's `/events` tier instead. Pass `stream: { includeHarnessEvents: true }` to keep them for a lenient consumer.) + +## Event mapping + +| Harness concept | AG-UI event | +| --- | --- | +| run start / end | `RUN_STARTED` / `RUN_FINISHED` | +| assistant text | `TEXT_MESSAGE_START` / `TEXT_MESSAGE_CONTENT` / `TEXT_MESSAGE_END` | +| tool call | `TOOL_CALL_START` / `TOOL_CALL_ARGS` / `TOOL_CALL_END` / `TOOL_CALL_RESULT` | +| approval / wait | `RUN_FINISHED` with `outcome.type === 'interrupt'` | +| resume | a new run whose `RunAgentInput.resume` answers the interrupts | +| subagent | `SUBAGENT_STARTED` / `SUBAGENT_FINISHED` (with `subagentRunId`) | +| token usage | `metadata.tanstack.usage` on `RUN_FINISHED` | +| live spend (interim) | a `tanstack.spend` `CUSTOM` event | + +## Approvals as interrupts + +A turn that stops for approval finishes with a `RUN_FINISHED` whose `outcome.type === 'interrupt'`. `@ag-ui/client` collects these on `agent.pendingInterrupts`: + +```ts group=harness-ag-ui-client +await agent.runAgent() + +for (const interrupt of agent.pendingInterrupts) { + console.log(interrupt.id, interrupt.message) // "Approval required to run remove" +} +``` + +Answer by starting a new run whose `resume` entries reference the interrupt ids: + +```ts group=harness-ag-ui-resume +await fetch('https://example.com/agent', { + method: 'POST', + headers: { + 'content-type': 'application/json', + authorization: `Bearer ${process.env.HARNESS_TOKEN}`, + }, + body: JSON.stringify({ + threadId: 'user-1/main', + runId: 'run-2', + messages: [], + tools: [], + context: [], + resume: [{ interruptId: 'approval_call_1', status: 'resolved', payload: true }], + }), +}) +``` + +The harness-native approval path (`session.resolve`, the CLI, and the relay dashboard) keeps working; AG-UI interrupt/resume is the AG-UI-shaped view of the same wait. + +## Token usage and spend + +`RUN_FINISHED` carries usage under `metadata.tanstack.usage` as a `NormalizedUsage` (`inputTokens`, `outputTokens`, `totalTokens`). Read it directly, or normalize any provider's usage with `normalizeUsage`: + +```ts group=harness-ag-ui-usage +import { normalizeUsage } from '@tanstack/ai-harness/ag-ui' + +normalizeUsage([{ inputTokens: 10, outputTokens: 5 }]) +// { inputTokens: 10, outputTokens: 5, totalTokens: 15 } +normalizeUsage({ promptTokens: 8, completionTokens: 2, totalTokens: 10 }) +// { inputTokens: 8, outputTokens: 2, totalTokens: 10 } +``` + +For a live spend meter, opt into interim ticks. After each run that reports usage, the stream carries a `tanstack.spend` `CUSTOM` event with that run's usage and a running cumulative total: + +```ts group=harness-ag-ui +const meteredHandler = createAgUiHandler({ + host, + harness: assistant, + authorize: () => ({ id: 'user-1' }), + stream: { emitSpendEvents: true }, +}) +``` + +Spend rides `CUSTOM` because the AG-UI spec has no first-class spend event yet — track that convention deliberately. + +## Pin the AG-UI version + +The AG-UI spec is still evolving. This bridge is built and tested against `@ag-ui/core@1.0.0`, `@ag-ui/encoder@1.0.0`, and `@ag-ui/client@1.0.0`. Pin those versions and move them on purpose, not by floating range. diff --git a/docs/harness/cli.md b/docs/harness/cli.md index 9889069159..cc45a4f81b 100644 --- a/docs/harness/cli.md +++ b/docs/harness/cli.md @@ -2,7 +2,7 @@ title: Run a harness in the terminal id: harness-cli order: 3 -description: "Give your harness a terminal UI, a print mode for scripts and CI, NDJSON output, an ACP mode for editors, and an HTTP server." +description: "Run your harness in a terminal with line mode or your own screen, a print mode for scripts and CI, NDJSON output, an ACP mode for editors, and an HTTP server." keywords: - tanstack ai - harness @@ -28,8 +28,6 @@ octane: @tanstack/ai-harness-cli -The interactive UI uses Ink, which needs Node 22 or later. - ## 1. Write the entry file ```ts group=harness-cli @@ -47,13 +45,19 @@ process.exitCode = await runCli(assistant) ## 2. Pick a mode -- No flags: the interactive UI. Type a message and press Enter. While the agent works, Enter steers it and Esc cancels. +To use the harness yourself: + +- No flags: line mode. Type a message and press Enter. With a `ui`, your own screen starts in its place (see [Run your own screen](#run-your-own-screen)). - `-p "prompt"`: run one prompt, print the answer, and exit. - `-p "prompt" --output ndjson`: print every AG-UI event as one JSON line. + +To use the harness from another program: + - `--acp`: serve the harness as an ACP v2 agent over stdio, for editors. Needs `@tanstack/ai-acp`. -- `--serve`: serve the session protocol over HTTP on `127.0.0.1:8787`. Every request needs the bearer token. Pass `--token`, set `HARNESS_TOKEN`, or copy the token the CLI prints. +- `--mcp`: serve the harness as an MCP server over stdio, for Claude Code, Cursor, and other MCP clients. Needs `@tanstack/ai-mcp`. Add `--yes` to approve every tool call. Read [Use a harness from any MCP client](./mcp-server). +- `--serve`: serve the session protocol over HTTP on `127.0.0.1:8787`. Every request needs the bearer token. Pass `--token`, set `HARNESS_TOKEN`, or copy the token the CLI prints. With `@tanstack/ai-mcp`, it also serves MCP at `/mcp`. -When stdin is a pipe, the CLI reads one message or command per line and waits for each turn. +Line mode reads one message or command per line and waits for each turn. It works the same in a terminal and with piped input. In a terminal, it also opens sign-in links in the browser. ## 3. Use it in CI @@ -66,7 +70,7 @@ When stdin is a pipe, the CLI reads one message or command per line and waits fo | 2 | The turn waits for an approval. | | 130 | The turn was cancelled. | -## Commands in the interactive UI +## Commands in line mode For the session: @@ -83,8 +87,34 @@ For the running work: Plugin commands (for example `/model` or `/todos`) show up in `/help`. When a turn stops for an approval or a plugin asks a question, type your answer. For yes-or-no questions, `y` approves and `n` refuses. +## Run your own screen + +Line mode prints plain lines. For a full screen with your own layout, pass `ui` to `runCli`. It works with any TUI library, for example Ink, OpenTUI, or blessed. + +`ui` gets a ready [session view](./custom-ui) and resolves when the user quits. This entry file starts an Ink screen: + +```tsx ignore +import { render } from 'ink' +import { runCli } from '@tanstack/ai-harness-cli' +import { assistant } from './harness' +import { Screen } from './screen' + +process.exitCode = await runCli(assistant, { + ui: async (view) => { + await render().waitUntilExit() + }, +}) +``` + +- `ui` runs only in an interactive terminal. Piped input uses line mode. `-p`, `--acp`, `--mcp`, `--serve`, and `--dashboard` do not use `ui`. +- When `ui` resolves, the CLI disposes the view and `runCli` returns. +- To write `Screen`, read [Build your own UI](./custom-ui). + +For a full Ink screen with approvals, questions, sign-ins, and child agents, copy [`examples/harness-cli/src/tui.tsx`](https://github.com/TanStack/ai/blob/main/examples/harness-cli/src/tui.tsx). + ## What you have now -- One entry file that runs your harness as a terminal app, a script step, an editor agent, or a server. +- One entry file that runs your harness as a terminal app, a script step, an editor agent, an MCP server, or an HTTP server. +- Your own terminal screen on the same session, with any TUI library. Next: keep long turns alive through crashes with [durable sessions](./durable-sessions). diff --git a/docs/harness/code-mode.md b/docs/harness/code-mode.md index c84a9e51fb..a73963b6f5 100644 --- a/docs/harness/code-mode.md +++ b/docs/harness/code-mode.md @@ -1,7 +1,7 @@ --- title: Code mode in a harness id: harness-code-mode -order: 9 +order: 10 description: "Let the harness model write one TypeScript program that calls many tools, and run it in an isolate. Any TanStack AI isolate driver plugs in." keywords: - tanstack ai diff --git a/docs/harness/coding-agent.md b/docs/harness/coding-agent.md index 8a45100343..d1dfe8e355 100644 --- a/docs/harness/coding-agent.md +++ b/docs/harness/coding-agent.md @@ -68,7 +68,8 @@ Run the file with `npx tsx coder.ts`. Ask for a change. The agent reads files fr | `projectInstructions({ root })` | Adds `AGENTS.md` and `CLAUDE.md` to the system prompt. | | `fileCommands({ dir })` | Each `.md` file becomes a slash command. `$ARGUMENTS` is replaced by what you type after it. | | `compact({ adapter })` | `/compact` replaces a long conversation with a summary. | -| `usage()` | `/usage` shows the tokens of the session. | +| `usage()` | `/usage` shows the tokens of the session: the lead turn and every agent. | +| `goal({ judge })` | `/goal ` keeps the agent working until a judge model says that the goal is met. See [Work until a goal is met](./goal). | ## Add your own rules @@ -90,6 +91,10 @@ A trailing `*` matches every tool that starts with the text. The last matching r The workspace tools run on your machine with your permissions. Run code you do not trust in a sandbox. +## Hand work to Claude Code or Codex + +Your agent can also give tasks to coding agents you already use. [Delegate to coding agents](./coding-agents) shows how. + ## What you have now - A terminal coding agent with file tools, approvals, modes, a todo list, and a model picker. diff --git a/docs/harness/coding-agents.md b/docs/harness/coding-agents.md new file mode 100644 index 0000000000..94e65175c5 --- /dev/null +++ b/docs/harness/coding-agents.md @@ -0,0 +1,122 @@ +--- +title: Delegate to coding agents +id: harness-coding-agents +order: 11 +description: "Let a harness hand coding work to Claude Code, Codex, Grok Build, or any ACP agent. Each one works in a sandbox and keeps its own session." +keywords: + - tanstack ai + - harness + - claude code + - codex + - grok build + - sandbox + - subagents +--- + +Your lead agent plans the work, but you want Claude Code or Codex to do the edits, because they are good at code and you already use them. `codingAgents` gives the lead model one tool per coding agent. Each agent works in a sandbox, keeps its own session between calls, and its tool calls show up in your UI. + +## 1. Add the plugin + + + +react: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +vue: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +solid: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +svelte: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +preact: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +angular: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +octane: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +vanilla: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex + + + +```ts group=harness-coding-agents +import { defineHarness } from '@tanstack/ai-harness' +import { permissions } from '@tanstack/ai-harness/plugins' +import { claudeCodeText } from '@tanstack/ai-claude-code' +import { codexText } from '@tanstack/ai-codex' +import { openaiText } from '@tanstack/ai-openai' +import { defineSandbox, defineWorkspace, localSource } from '@tanstack/ai-sandbox' +import { codingAgents } from '@tanstack/ai-sandbox/harness' +import { localProcessSandbox } from '@tanstack/ai-sandbox-local-process' + +const repo = '/path/to/your/repo' + +export const lead = defineHarness({ + name: 'acme/lead', + adapter: openaiText('gpt-5.6'), + plugins: () => [ + permissions(), + codingAgents({ + sandbox: defineSandbox({ + id: 'repo', + provider: localProcessSandbox({ dir: repo }), + workspace: defineWorkspace({ source: localSource(repo) }), + }), + agents: { + claude_code: { + adapter: claudeCodeText('claude-opus-4-8', { permissionMode: 'acceptEdits' }), + description: 'Larger changes, refactors, and reviews', + }, + codex: { + adapter: codexText('gpt-5.3-codex', { sandboxMode: 'workspace-write' }), + description: 'Quick fixes and tests', + }, + }, + }), + ], +}) +``` + +The lead model now has a `claude_code` tool and a `codex` tool. Each tool takes one `task`: the whole job, in words. The agent sees only that text and the files in its sandbox. + +## 2. Ask for work + +1. Start the harness, for example with the CLI. +2. Ask the lead: `have claude_code add a test for the date parser, then have codex fix what fails`. +3. Watch the child work. The CLI shows each agent's tool calls, then one line with the start of its answer: + +```text +[agent claude_code started] +[claude_code: tool Write] +[agent claude_code finished: Added parse-date.test.ts with three cases.] +``` + +The dashboard shows the same work in a block under the lead's message. An ACP editor shows the child's tool calls next to the lead's own. + +## Sessions + +Each agent keeps its own session per harness thread. The next task for `claude_code` resumes the same Claude Code session, so it still knows the files it read. The session ids live in the plugin state, so they also survive a restart of the host. + +To start over, run `/fresh claude_code`, or `/fresh` for every agent. + +## Workspaces + +`workspace` chooses where the agents work: + +| Value | Where each agent works | At the same time | +| --- | --- | --- | +| `'shared'` (default) | One sandbox per harness thread | One agent at a time | +| `'per-agent'` | A sandbox per agent | Yes | + +Use `'shared'` when the agents build on each other's changes. Use `'per-agent'` when they work on separate copies and you merge the results. + +## Plan mode + +When the `permissions()` plugin is in `plan` mode, the agents start read-only: + +- Claude Code gets `permissionMode: 'plan'`. +- Codex gets `sandboxMode: 'read-only'`. +- Any other agent gets its `planModelOptions`, for example `{ permissionMode: 'default' }` for an ACP agent that asks before each edit. + +## Other agents and sandboxes + +- `adapter` takes any coding-agent adapter: `grokBuildText` from `@tanstack/ai-grok-build`, or `acpCompatibleText` from `@tanstack/ai-acp` for any ACP agent. +- `sandbox` takes any sandbox provider. [Sandbox providers](../sandbox/providers) lists them. +- `modelOptions` on an agent is added to every call, for example a fixed `permissionMode`. + +## What you have now + +- A lead model that hands tasks to Claude Code and Codex. +- One sandbox per thread, and one saved session per agent. +- Child tool calls in the CLI, the dashboard, and ACP editors. diff --git a/docs/harness/custom-ui.md b/docs/harness/custom-ui.md new file mode 100644 index 0000000000..aeedae39de --- /dev/null +++ b/docs/harness/custom-ui.md @@ -0,0 +1,318 @@ +--- +title: Build your own UI +id: harness-custom-ui +order: 13 +description: "Show a harness session in your own terminal screen or web app. One live store holds the messages, tool calls, approvals, and plugin state, and works with any UI library." +keywords: + - tanstack ai + - harness + - session view + - tanstack store + - ink + - custom ui +--- + +You want your own screen for your harness: your own layout in a terminal, or a page in your web app. `session.events()` gives you raw AG-UI events. To show them, you must join text deltas, follow each tool call, and keep a list of open approvals. That is a reducer, and you must keep it correct. + +`createSessionView` does that work. It keeps one live state in a TanStack Store, and it gives you actions and typed events. You show the state with any UI library. + +At the end of this page you have a terminal screen in about 20 lines, and the same view in a web app. + +## Install + +Install the harness and the TanStack Store adapter for your UI library. With no UI library, the harness is enough. + + + +react: @tanstack/ai-harness @tanstack/react-store +vue: @tanstack/ai-harness @tanstack/vue-store +solid: @tanstack/ai-harness @tanstack/solid-store +svelte: @tanstack/ai-harness @tanstack/svelte-store +preact: @tanstack/ai-harness @tanstack/preact-store +angular: @tanstack/ai-harness @tanstack/angular-store +vanilla: @tanstack/ai-harness +octane: @tanstack/ai-harness + + + +The terminal screen on this page uses Ink, which needs Node 22 or later. Install Ink and React too: + + + +react: ink react + + + +## 1. Make a view + +Open a session, then give it to `createSessionView`: + +```ts group=harness-custom-ui +import { createHarnessHost, defineHarness } from '@tanstack/ai-harness' +import { createSessionView } from '@tanstack/ai-harness/view' +import { memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' + +const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), +}) + +const host = createHarnessHost({ persistence: memoryPersistence() }) +const session = await host.open(assistant, { threadId: 'thread-1' }) + +const view = createSessionView(session) +await view.ready +``` + +`view.ready` resolves when the view has the saved messages, the status, and the commands and settings of the session. If you open a thread again, its history is in the view immediately. + +## 2. Show the state + +`view.store` is a TanStack Store. `useSelector` reads one part of the state, and your component updates only when that part changes. This Ink screen shows the messages and the status: + +```tsx ignore +import { Box, Text, render } from 'ink' +import { useSelector } from '@tanstack/react-store' +import type { SessionView } from '@tanstack/ai-harness/view' + +function Screen({ view }: { view: SessionView }) { + const messages = useSelector(view.store, (state) => state.messages) + const status = useSelector(view.store, (state) => state.status) + return ( + + {messages.map((message) => + message.role === 'assistant' ? ( + + {message.parts.map((part) => (part.type === 'text' ? part.text : '')).join('')} + + ) : ( + + {message.role === 'user' ? `> ${message.text}` : message.text} + + ), + )} + {status} + + ) +} + +render() +await view.send('Summarize the README.') +``` + +Put the code of steps 1 and 2 in one `.tsx` file, then run it. The screen shows your prompt, then the answer while it streams, then the status `idle`. Press Ctrl+C to quit. + +A full Ink screen, with approvals, questions, sign-ins, and child agents, is in [`examples/harness-cli/src/tui.tsx`](https://github.com/TanStack/ai/blob/main/examples/harness-cli/src/tui.tsx). It runs from [`runCli({ ui })`](./cli#run-your-own-screen). + +Each TanStack Store adapter reads the store the same way. Vue, Solid, Svelte, and Preact have `useSelector`. Angular has `injectSelector`. With no UI library, subscribe to the store: + +```ts group=harness-custom-ui +const subscription = view.store.subscribe((state) => { + console.log(`${state.status}: ${state.messages.length} messages`) +}) +``` + +Call `subscription.unsubscribe()` to stop. + +### What the state holds + +The conversation: + +- `messages`: the lines to show. Each message has a `role`: `user`, `assistant`, or `notice`. +- `status`: `idle`, `running`, or `requires_action`. +- `connection`: `open`, `reconnecting`, or `closed`. + +What waits for the user: + +- `approvals`: tool calls that wait for a yes or a no. +- `questions`: questions from a command or a plugin. +- `signIns`: connectors that need a sign-in, with a `url` and a `userCode` when the connector gives them. + +The session: + +- `threadId`: the id of the conversation. +- `agents`: the background agents that run now. +- `queuedTurns`: the number of messages that wait for their turn. +- `commands`, `config`, and `tools`: what the session has, for a help screen or a settings panel. +- `plugins`: the saved state of each plugin, by plugin name. + +An assistant message has `parts`. Each part is one of these: + +- `text` or `reasoning`: the text, which grows while it streams. +- `tool-call`: a tool call with its `name`, `args`, and `status` (`running`, `done`, `failed`, or `needs-approval`). +- `agent`: a child agent, with its own `parts`. + +A notice has a `kind`: + +- `info`: the session continues a turn that a crash stopped. +- `error`: a turn or an action failed. +- `rejected`: the session did not accept a message. +- `command`: the text result of a command. +- `ui`: a line that you added with `view.notice(text)`. + +## 3. Act on the session + +Start and stop work: + +- `view.send(text)`: sends a prompt. While a turn runs, the text steers that turn. Text that starts with `/` runs a command, for example `/goal all tests pass`. +- `view.command(name, input)`: runs a command, for example from a button. +- `view.setConfig(key, value)`: changes a setting, for example `view.setConfig('mode', 'plan')`. +- `view.cancel()`: cancels the running turn. +- `view.dispose()`: stops the view when your screen closes. After that, each action throws an error. + +Answer what waits: + +- `approval.approve()` and `approval.reject()`: answer one approval. `view.approve(id)` and `view.reject(id)` do the same by id. +- `view.approveAll()` and `view.rejectAll()`: answer all open approvals. +- `question.answer(value)`: answers a question. + +If a turn waits for more than one approval, the turn continues after you answer all of them. + +## 4. Listen for events + +Some items need attention when they arrive. `view.on` calls your handler for each one. When a tool call waits for approval, this handler rings the terminal bell and adds a line to the screen: + +```ts group=harness-custom-ui +view.on('approval', (approval) => { + process.stdout.write('\u0007') + view.notice(`${approval.tool} waits for your answer.`) +}) +``` + +Items that wait for the user: + +- `'approval'`: an approval. Call `approve()` or `reject()` on it. +- `'question'`: a question with a `message`, a `schema`, and `answer(value)`. +- `'signIn'`: a connector that needs a sign-in, with `connector`, `url`, and `userCode`. + +Progress: + +- `'toolCall'`: `{ id, name }`, when a tool call starts. +- `'agent'`: `{ id, name, status }`, when a child agent starts, finishes, or fails. +- `'error'`: the error text. The view also adds it to `messages` as a notice. +- `'turnEnd'`: `{ operationId }`, when a chat turn ends. + +The view calls your handler after the item is in the state, so the handler can read `view.store.get()` and find the item. `view.on` returns a function that removes the handler. + +## Plugin state and plugin events + +Plugins show their work through state and events. `state.plugins` holds the saved state of each plugin, by plugin name. A plugin can export a selector that checks the shape of its state for you. + +With the [goal plugin](./goal) in your harness, `selectGoal` reads the goal from the state, or gives `null`. `GoalMet` is the typed event of the plugin: + +```ts group=harness-custom-ui +import { GoalMet, selectGoal } from '@tanstack/ai-harness/plugins' + +const goal = selectGoal(view.store.get()) +if (goal) view.notice(`Goal: ${goal.text} (${goal.status}, round ${goal.round})`) + +view.on(GoalMet, (met) => view.notice(`Goal met: ${met.reason}`)) +``` + +In a component, give the selector to `useSelector`: `useSelector(view.store, selectGoal)`. `view.on` takes any plugin event from `createPluginEvent`, and the handler gets its typed value. `@tanstack/ai-harness/plugins` runs only in Node, so use it with a local session. In a browser, read `state.plugins['tanstack/goal']` and check its shape before you use it. + +## A UI in the browser + +Your harness runs on a server, and your UI is a web page. Give `createSessionView` a `HarnessClient` in place of a session. The state, the actions, and the events stay the same. + +### On the server + +Serve the session with `createHarnessHandler`, and mount the handler on `/api/harness`: + +```ts group=harness-custom-ui-server +import { createHarnessHandler, createHarnessHost, defineHarness } from '@tanstack/ai-harness' +import { memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' + +export const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), +}) + +const host = createHarnessHost({ persistence: memoryPersistence() }) + +export const handler = createHarnessHandler({ + host, + harness: assistant, + authorize: (request) => + request.headers.get('authorization') === `Bearer ${process.env.HARNESS_TOKEN}` + ? { id: 'user-1' } + : null, + canAccess: (principal, threadId) => threadId.startsWith(principal.id), +}) +``` + +The handler also serves `GET transcript` and `GET describe`. The view reads them when it starts. For the other routes, see [Connect clients](./connect). + +### In the browser + +Make the view one time, in code that runs in the browser. Then read it in your components: + +```tsx ignore +import { useState } from 'react' +import { useSelector } from '@tanstack/react-store' +import { createHarnessClient } from '@tanstack/ai-harness/client' +import { createSessionView } from '@tanstack/ai-harness/view' + +const view = createSessionView( + createHarnessClient({ + url: '/api/harness', + threadId: 'user-1-thread', + headers: { authorization: 'Bearer my-token' }, + }), +) + +export function Chat() { + const messages = useSelector(view.store, (state) => state.messages) + const approvals = useSelector(view.store, (state) => state.approvals) + const connection = useSelector(view.store, (state) => state.connection) + const [draft, setDraft] = useState('') + + return ( +
+

{connection === 'reconnecting' ? 'Reconnecting...' : ''}

+ {messages.map((message) => ( +

+ {message.role === 'assistant' + ? message.parts.map((part) => (part.type === 'text' ? part.text : '')).join('') + : message.text} +

+ ))} + {approvals.map((approval) => ( +

+ Run {approval.tool}?{' '} + + +

+ ))} +
{ + event.preventDefault() + void view.send(draft) + setDraft('') + }} + > + setDraft(event.target.value)} /> + +
+
+ ) +} +``` + +`connection` tells the user about the network: + +- `open`: events arrive. +- `reconnecting`: the connection dropped. The client connects again and continues after the last event. +- `closed`: the view stopped reading events, for example after `view.dispose()`. + +## What you have now + +- A terminal screen in about 20 lines, with no event loop and no reducer. +- The same view in a web app, on a harness that runs on a server. +- Actions, typed events, and plugin state that work with any UI library. diff --git a/docs/harness/dashboard.md b/docs/harness/dashboard.md index afbd511cd7..f30c6f197c 100644 --- a/docs/harness/dashboard.md +++ b/docs/harness/dashboard.md @@ -1,7 +1,7 @@ --- title: Self-host the dashboard id: harness-dashboard -order: 12 +order: 16 description: "Watch and steer your harness sessions from a browser or a phone. Agents dial out to your dashboard server, so they need no open port." keywords: - tanstack ai diff --git a/docs/harness/deploy.md b/docs/harness/deploy.md index ea6cf77423..387381c432 100644 --- a/docs/harness/deploy.md +++ b/docs/harness/deploy.md @@ -1,7 +1,7 @@ --- title: Deploy a harness id: harness-deploy -order: 11 +order: 15 description: "Run a harness in your server, as a worker process, on another machine, or as a single executable." keywords: - tanstack ai diff --git a/docs/harness/goal.md b/docs/harness/goal.md new file mode 100644 index 0000000000..bdfb182d35 --- /dev/null +++ b/docs/harness/goal.md @@ -0,0 +1,100 @@ +--- +title: Work until a goal is met +id: harness-goal +order: 12 +description: "Give the harness a goal, for example all tests pass. After each turn, a judge model checks the goal, and the harness keeps working until the goal is met." +keywords: + - tanstack ai + - harness + - goal + - judge + - agent loop +--- + +You ask your agent to make all tests pass. It fixes one test, then it stops and waits for you. You type "keep going" again and again. + +The `goal()` plugin keeps the harness working. You give the goal one time. After each turn, a judge model reads the goal and the end of the conversation. If the goal is not met, the plugin starts the next turn. + +## 1. Add the plugin + +```ts group=harness-goal +import { defineHarness } from '@tanstack/ai-harness' +import { goal, permissions, workspaceTools } from '@tanstack/ai-harness/plugins' +import { openaiText } from '@tanstack/ai-openai' + +const root = process.cwd() + +export const coder = defineHarness({ + name: 'acme/coder', + adapter: openaiText('gpt-5.6'), + plugins: () => [ + permissions(), + workspaceTools({ root }), + goal({ judge: openaiText('gpt-5.6-luna') }), + ], +}) +``` + +`judge` is the model that decides if the goal is met. It can be a smaller model than the main model, or the same model. + +## 2. Give it a goal + +1. Start the harness, for example with the [CLI](./cli). +2. Type `/goal all tests pass`. The first turn starts. +3. Watch the turns. Each new turn tells the model the reason of the last check: + +```text +Keep working on the goal: all tests pass. Last check: The date parser test still fails. +``` + +4. Type `/goal` to see the result: + +```text +Goal: all tests pass +Status: met, round 3 of 20. +Last check: All 12 tests pass. +``` + +## How it stops + +The plugin starts no new turn in these cases: + +- The judge says that the goal is met. The status is `met`. +- The plugin started `maxRounds` turns (20 by default). The status is `stopped`, and the last reason stays. +- A turn waits for an approval, or a turn fails. The status is `paused`. +- You send your own message. The status is `paused`, and your message runs as a normal turn. + +To change the limit, set `maxRounds`, for example `goal({ judge, maxRounds: 50 })`. + +## Commands + +| Command | What it does | +| --- | --- | +| `/goal ` | Sets the goal and starts the first turn. | +| `/goal` | Shows the goal, the status, the round, and the last reason. | +| `/goal stop` | Ends the goal. The running turn finishes, and no new turn starts. | +| `/goal resume` | Continues a paused or stopped goal with a new round count. The goal is plugin state, so this also works after a restart. | + +## Do something when the goal is met + +The plugin sends the `GoalMet` event with the goal and the last reason. Listen for it in your own plugin: + +```ts group=harness-goal +import { definePlugin } from '@tanstack/ai-harness' +import { GoalMet } from '@tanstack/ai-harness/plugins' + +export const notify = definePlugin({ + name: 'acme/notify', + setup: (ctx) => { + ctx.on(GoalMet, (met) => console.log(`Done: ${met.goal}. ${met.reason}`)) + }, +}) +``` + +Clients get the same event as a `harness.plugin.event` with the name `tanstack/goal:met`. + +## What you have now + +- A harness that keeps working until the judge says that the goal is met. +- A round limit, so the loop cannot run without end. +- `/goal`, `/goal stop`, and `/goal resume` to see and control the goal. diff --git a/docs/harness/mcp-server.md b/docs/harness/mcp-server.md new file mode 100644 index 0000000000..ff0af7099e --- /dev/null +++ b/docs/harness/mcp-server.md @@ -0,0 +1,224 @@ +--- +title: Use a harness from any MCP client +id: harness-mcp-server +order: 9 +description: "Serve your harness as an MCP server. Chat with it, answer its approvals, and run its agents and commands from Claude Code, Claude Desktop, Cursor, or another agent." +keywords: + - tanstack ai + - harness + - mcp + - mcp server + - claude code + - cursor +--- + +You built a harness, but you do your work in Claude Code, Cursor, or another agent. You want to give work to your harness from there. These clients speak MCP. With `--mcp`, the CLI serves your harness as an MCP server. The client can then chat with the harness, answer its approvals, and run its agents and commands. + +## 1. Install the MCP package + + + +react: @tanstack/ai-harness-cli @tanstack/ai-mcp +vue: @tanstack/ai-harness-cli @tanstack/ai-mcp +solid: @tanstack/ai-harness-cli @tanstack/ai-mcp +svelte: @tanstack/ai-harness-cli @tanstack/ai-mcp +preact: @tanstack/ai-harness-cli @tanstack/ai-mcp +angular: @tanstack/ai-harness-cli @tanstack/ai-mcp +octane: @tanstack/ai-harness-cli @tanstack/ai-mcp +vanilla: @tanstack/ai-harness-cli @tanstack/ai-mcp + + + +## 2. Add the harness to your client + +Start from the entry file in [Run a harness in the terminal](./cli). With `--mcp`, it reads MCP messages on stdin and writes the answers on stdout. + +For Claude Code, run this command one time: + +```bash +claude mcp add --env OPENAI_API_KEY=your-key assistant -- npx tsx /path/to/cli.ts --mcp +``` + +For Claude Desktop, Cursor, and other clients with a JSON file, add the same command under `mcpServers`: + +```json +{ + "mcpServers": { + "assistant": { + "command": "npx", + "args": ["tsx", "/path/to/cli.ts", "--mcp"], + "env": { "OPENAI_API_KEY": "your-key" } + } + } +} +``` + +- Claude Desktop keeps this in `claude_desktop_config.json`. +- Cursor keeps this in `.cursor/mcp.json`. + +The client starts the process when it needs the harness. The process stops when the client closes stdin. In this mode, stdout carries only MCP messages. Write your own logs with `console.error`. + +## 3. Use the tools + +Ask the client to use the harness, for example: "Ask assistant to summarize my open tickets." The client calls these tools. + +To talk to the harness: + +- `chat`: send a message and wait for the answer. While a turn runs, the message waits in the queue. +- `steer`: add a message to the running turn. +- `cancel`: cancel the running turn. +- `status`: show the status, the interrupts and questions that wait, the background agents, and the queued turns. + +To answer the harness: + +- `approve` and `reject`: answer every approval that waits. +- `resolve`: answer every interrupt that waits, in one call. +- `answer`: answer a question from a command or a plugin. + +To run what the harness has: + +- `agent_`: one tool for each agent in `expose.agents`. It takes the agent input and returns the agent result. +- `command_`: one tool for each plugin command. The command `connect:notion` becomes the tool `command_connect_notion`. + +Two names can give the same tool name, for example `connect:notion` and `connect_notion`. Then the later one in name order gets a number: `command_connect_notion_2`. The description of each of these tools names its command or agent. + +Every tool takes an optional `threadId`. Each thread is its own conversation. Without a `threadId`, a tool uses the `--thread` conversation (default `main`). + +`chat` answers with JSON: + +```json +{ + "status": "completed", + "text": "You have three open tickets.", + "interrupts": [], + "questions": [] +} +``` + +## Interrupts and questions + +A turn can stop for an interrupt. The turn then waits for your answer. Each interrupt has a `kind`: + +- `approval`: a tool call that needs your yes or no. +- `client-tool`: a tool that the client runs. The turn waits for the tool output. +- `generic`: a question from a middleware. The answer must match its `responseSchema`. + +For an approval, the client asks you in its own window (MCP elicitation). Answer `yes` to run the tool. Any other answer rejects it. + +In these cases, `chat` returns the interrupts with `status: "interrupted"`: + +- The client cannot ask you. +- You close the question without an answer. +- An interrupt of another kind also waits. + +```json +{ + "status": "interrupted", + "text": "Here is the plan.", + "interrupts": [ + { + "id": "interrupt-1", + "kind": "generic", + "message": "Review the plan", + "responseSchema": { + "type": "object", + "properties": { "note": { "type": "string" } }, + "required": ["note"] + } + } + ], + "questions": [] +} +``` + +An approval and a client tool also have `tool` and `args`, the tool call that waits. + +Then the client calls `approve`, `reject`, or `resolve`. That call waits for the turn to continue and returns the same JSON as `chat`. `approve` and `reject` answer approvals only. When an interrupt of another kind also waits, they refuse, and the client uses `resolve`. + +`resolve` takes one decision for each interrupt that waits. This example answers a turn with one approval and one generic interrupt: + +```json +{ + "decisions": [ + { "interruptId": "approval-1", "approved": true }, + { "interruptId": "interrupt-1", "payload": { "note": "Ship it." } } + ] +} +``` + +- For an approval, set `approved`. `true` runs the tool, and `false` rejects it. +- For a client tool, set `payload` to the tool output. +- For a generic interrupt, set `payload` to a value that matches its `responseSchema`. + +`resolve` refuses a list that does not answer every interrupt. It also refuses a decision without the field that its kind needs. + +To approve every tool call without a question, start the CLI with `--yes`. Use it only with tools that you trust. `--yes` answers approvals only, so other interrupts still come back in the result. + +A command or a plugin can also ask you a question. The result then has `status: "waiting"` and the question in `questions`. The client calls `answer` with the question `id` as `questionId` and your `value`. `answer` waits for the work that asked, then returns its result. + +## Serve MCP over HTTP + +If `@tanstack/ai-mcp` is installed, `--serve` also serves the MCP server at `/mcp`. The route uses the same bearer token as the other routes: + +```bash +HARNESS_TOKEN=your-token npx tsx cli.ts --serve +claude mcp add --transport http assistant http://127.0.0.1:8787/mcp --header "Authorization: Bearer your-token" +``` + +A request without the token gets a 401. `--yes` works on this route too. + +## Serve it from your own server + +To put the MCP server in your own app, use `createHarnessMcpServer`. It opens sessions on your host and returns a server with a `fetch` handler: + +```ts group=harness-mcp-server +import { createHarnessHost, defineHarness } from '@tanstack/ai-harness' +import { createHarnessMcpServer } from '@tanstack/ai-mcp/harness' +import { memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' + +const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), +}) + +const host = createHarnessHost({ persistence: memoryPersistence() }) + +export const server = await createHarnessMcpServer({ + host, + harness: assistant, + threadId: 'main', + approvals: 'ask', +}) + +export async function handleMcp(request: Request) { + const token = request.headers.get('authorization') + if (token !== `Bearer ${process.env.MCP_TOKEN}`) { + return new Response('Unauthorized', { status: 401 }) + } + return server.fetch(request) +} +``` + +Mount `handleMcp` on a route, for example `/mcp`. The options are: + +- `threadId`: the conversation of a tool call without a `threadId`. Default `main`. +- `approvals`: `'ask'` (default) asks in the client, or returns the approvals when the client cannot ask. `'auto'` approves every tool call. Both answer approvals only. +- `name` and `version`: the MCP server name and version. Default: the harness name and `1.0.0`. + +For a process that a client starts, serve the same server on stdio: + +```ts group=harness-mcp-server +import { serveMCPStdio } from '@tanstack/ai-mcp/server/stdio' + +serveMCPStdio(server) +``` + +## What you have now + +- Your harness in Claude Code, Claude Desktop, Cursor, or any other MCP client. +- Approvals that the client asks you about, or answers with `approve` and `reject`. +- Every kind of interrupt, answered with one `resolve` call. +- The agents and commands of your harness as MCP tools. + +Next: put the harness on a server with [Deploy a harness](./deploy). diff --git a/docs/harness/mcp.md b/docs/harness/mcp.md index 52dff3f012..6e62f78d0b 100644 --- a/docs/harness/mcp.md +++ b/docs/harness/mcp.md @@ -15,6 +15,8 @@ keywords: Your agent needs to read the user's Linear issues and Notion pages. Both services run MCP servers that sign in with OAuth, so you do not register an app or copy an API key. `mcpConnector` adds `/connect linear`, keeps the token in your credential store, and gives the model the server's tools after sign-in. +This page connects your harness to MCP servers. To use your harness from an MCP client such as Claude Code, read [Use a harness from any MCP client](./mcp-server). + ## 1. Add the connectors diff --git a/docs/harness/plugins.md b/docs/harness/plugins.md index 4e01346a91..9efa2dd835 100644 --- a/docs/harness/plugins.md +++ b/docs/harness/plugins.md @@ -30,9 +30,10 @@ Add it with `plugins: () => [today]` in `defineHarness`. `setup` runs once per s - `tools`: tools for the model, made with `toolDefinition`. - `prompts`: text for the system prompt. A function runs for each turn, so it can show current state. -- `middleware`: chat middleware, the same type as `chat({ middleware })`. +- `middleware`: chat middleware for the lead turn, the same type as `chat({ middleware })`. `agentMiddleware` is for agent runs, see [Middleware in every agent](#middleware-in-every-agent). - `generationMiddleware`: middleware for the activities agents call. - `agents`: agents added to `session.agents`. +- `subagents`: agents the model can call as tools. They are also added to `session.agents`. [Delegate to coding agents](./coding-agents) uses them. - `commands`: user actions, see below. - `config`: session settings, see below. - `contribute`: items for another plugin's extension point. @@ -98,6 +99,42 @@ export const counter = definePlugin({ When two writers race, `update` runs your function again with fresh state. +## Middleware in every agent + +You want to track token cost, or apply a policy, for every model call. But `middleware` runs only in the lead turn. The agents that the lead model calls, the agents you start in the background, and their own children each run a separate chat. Put the same middleware in `agentMiddleware` too. Then it runs in each of those chats. + +```ts group=harness-plugins +import type { ChatMiddleware } from '@tanstack/ai' + +export const usageByAgent = definePlugin({ + name: 'acme/usage-by-agent', + setup: () => { + const tokens = new Map() + const tracker: ChatMiddleware = { + name: 'acme/usage-by-agent', + onUsage: (ctx, usage) => { + const agent = ctx.subagentName ?? 'lead' + tokens.set(agent, (tokens.get(agent) ?? 0) + usage.totalTokens) + }, + } + return { + middleware: [tracker], + agentMiddleware: [tracker], + commands: { + tokens: defineCommand({ + description: 'Show the tokens of each agent', + run: () => Object.fromEntries(tokens), + }), + }, + } + }, +}) +``` + +- `ctx.subagentName` is the name of the agent. In the lead turn, it is `undefined`. +- `ctx.subagentRunId` is different for each run of an agent. For a nested child, `ctx.parentSubagentRunId` is the id of the child that started it. +- Each model call goes to `onUsage` one time, in the run that made the call. The lead turn does not count the usage of a child again. + ## Let plugins work together Three ways, from simple to loose: @@ -151,6 +188,7 @@ Set `lifetime: 'run'` to set a plugin up again for each turn. ## What you have now - A plugin that adds tools, prompts, commands, settings, and state to any harness. +- Middleware that sees every model call, in the lead turn and in every agent. - Plugins that share services, lists, and events without knowing each other. Next: see the [first-party plugins](./coding-agent) that turn a harness into a coding agent. diff --git a/docs/harness/subagents.md b/docs/harness/subagents.md index 78de061c97..0e33a72367 100644 --- a/docs/harness/subagents.md +++ b/docs/harness/subagents.md @@ -1,7 +1,7 @@ --- title: Run agents from a harness id: harness-subagents -order: 10 +order: 14 description: "Start typed agents from commands and plugins, run them in groups, call a whole harness as a child, and keep the tree within limits." keywords: - tanstack ai @@ -62,6 +62,8 @@ export const review = definePlugin({ - `ctx.agents.start(agent, input, { wake: true })` runs it in the background and starts a turn when it is done. - `ctx.agents.group(options, body)` runs several. With `onFailure: 'cancel-siblings'`, one failure cancels the others. With `'collect'`, use `group.runSettled` to get every result or error. Every child settles before `group` returns. +To track usage or apply a policy in each of these runs, see [Middleware in every agent](./plugins#middleware-in-every-agent). + ## Call a harness as a child `harnessAgent` turns a harness into an agent. Put it in `subagents.agents`, and the main model calls it as a tool. The child harness keeps its own tools, plugins, and history. diff --git a/examples/agent-dashboard/.gitignore b/examples/agent-dashboard/.gitignore new file mode 100644 index 0000000000..771f390df3 --- /dev/null +++ b/examples/agent-dashboard/.gitignore @@ -0,0 +1,15 @@ +node_modules +.DS_Store +dist +dist-ssr +*.local +.env +.env.local +.nitro +.tanstack +.output +.vinxi +test-results +playwright-report +*.log +.data diff --git a/examples/agent-dashboard/README.md b/examples/agent-dashboard/README.md new file mode 100644 index 0000000000..5142a1278c --- /dev/null +++ b/examples/agent-dashboard/README.md @@ -0,0 +1,152 @@ +# Agent Dashboard + +Mission control for TanStack AI agents — a TanStack Start app that watches live +agent sessions, approves tool calls mid-run, and tracks spend, all as a live +projection of the [AG-UI](../../docs/harness/ag-ui.md) event stream. + +It embeds a deterministic **support-triage** agent, so it runs with **no API +key**: the agent looks up a ticket (an auto tool), drafts a customer reply (an +approval-gated tool that pauses the run), and sends it once you approve. (One +team is the exception — the **Reddit pod** uses a real service and a real LLM; +see below.) + +```bash +pnpm --filter agent-dashboard dev # http://localhost:3002 +``` + +## What it exercises + +- **TanStack Start** — app shell + routing: hosts → sessions → session detail + (`src/routes`), and server API routes that host the agent (`src/routes/api.*`). +- **`@tanstack/ai-harness/ag-ui`** — the session view consumes the AG-UI SSE + stream with a bare `@ag-ui/client` `HttpAgent` (`src/lib/session-controller.ts`); + no bespoke protocol code. +- **TanStack DB** — run state (messages, tool calls, approvals, spend) lives in + `localOnly` collections written from the stream and read with `useLiveQuery` + (`src/db/collections.ts`). The UI is a projection of the stream, not a poller. +- **TanStack Query** — server state: host/session/run lists and agent config. +- **Composing teams (product UI)** — the home page is an **agents table**; each + row's **Add to team** starts a new team with that agent or drops it into an + existing one. On a team, **+ Add agent** adds any available agent, and each + roster agent has a **🔧 tools** button to run one of its public tools with + JSON parameters. Agents carry **default subscriptions** by harness (what they + react to), so a hand-composed team behaves like a seeded one. +- **TanStack DevTools** — the demo-only scaffolding (the seeded-team launchers, + the triage demo, automations, pod memory) lives in a custom **Demo Controls** + panel, kept out of the product UI so it's clear what's scaffolding vs. the real + experience. The panel renders from the devtools root (outside the route tree) + and drives the app purely by reading the same live TanStack DB state the UI + does — so it doubles as a state-management stress test. + +## Reddit pod (real service, real AI) + +The **+ React-news demo** team is the first pod wired to a _real_ external +service and a _real_ LLM — the graduation from the scripted demo agents: + +- **`reddit/fetcher`** — a procedural agent (no LLM) carrying one real tool, + `reddit.search_react_news`, which reads Reddit's public **RSS (Atom)** feed + (read-only, no auth, no key). Run it from the roster's **🔧 tools** button, a + 30-min schedule, or the Demo Controls panel. +- **`sentiment/react`** — a **real LLM** agent (Anthropic). Its harness default + subscription is the fetcher's tool _result_ + (`{ event: 'tool_result', tool: 'reddit.search_react_news', action: 'trigger' }`), + so whether you spin up the seeded demo or compose the team by hand, a news + batch triggers it automatically — nobody runs it — and it posts a sentiment + digest, persisting standout signals to pod memory. + +The loop is **timer → tool result → subscription → LLM digest**. The trigger +path spends zero tokens; the only cost is the digest itself. Chat messages don't +match the `tool_result` subscription, so the loop doesn't feed itself (a Phase 4 +scoping case study — verified by letting it run several cycles). + +### The one manual step: an API key + +`sentiment/react` needs a real provider. Set `ANTHROPIC_API_KEY` before starting +the dev server to get a live digest: + +```bash +ANTHROPIC_API_KEY=sk-ant-... pnpm --filter agent-dashboard dev +``` + +Without a key the agent posts a "set the key" message instead of a digest — the +rest of the dashboard still runs key-free. The e2e suite uses a deterministic +double and a recorded Reddit fixture, so it needs neither a key nor the network. + +> **Why RSS, not `.json`:** Reddit's public JSON (`/r/x/new.json`) returns `403` +> for many datacenter/VPN egress IPs regardless of `User-Agent`, while the Atom +> feed (`/r/x/new.rss`) is served — so the tool reads RSS. RSS carries title, +> link, author, timestamp and body (enough for a digest) but not score/comment +> counts. Reddit still rate-limits bursts (a rapid retry can `429`); the 30-min +> schedule stays well clear. The e2e suite uses the recorded fixture, so it +> needs neither network nor key. + +## Control plane + +- **History** (`/history`) — past runs backed by `HarnessPersistence` + (`runs.listByThread`). Opening a session replays it from stored events + (`/api/replay`), so a session you didn't run in this tab rehydrates from the + persisted feed. +- **Spend** (`/spend`) — per-session token rollups (a live query over the + `spend` collection) with per-session budgets and over-budget alerts. +- **Config** (`/config`) — a form generated from the agent's typed + `ConfigOption` schemas; edits write through the harness protocol + (`op: 'config'`). Endpoints are versioned with `HARNESS_PROTOCOL_VERSION`. + +## The approval queue + +When a run pauses on an approval, an approval card appears. It resolves the +interrupt through **both** paths the interrupt supports: + +- **Approve** → the AG-UI resume flow (`runAgent({ resume })` over `/api/agent`), + so the continuation streams back into the view. +- **Deny** → the harness-native control endpoint (`/api/harness/control`). +- **Edit** → approve with edited tool arguments. + +## Meta-chat (the demo) + +`/chat` is the dashboard's **own** agent (`dashboard/meta`), a tool-using chat +over live dashboard state. It runs on the same host as every other agent, so its +runs and tool calls show up in History and the trace view — the dashboard +dogfooding itself. Its tools: `list_agents`, `list_sessions`, `query_runs`, +`get_agent_config`, `set_agent_config`, `summarize_session`. + +Demo script: + +1. Open `/chat`. +2. Ask **"List the agents on this host"** — it calls `list_agents` (visible + inline) and answers with the agents registered on the host. +3. Run a triage session (open the **Demo Controls** devtools panel → **+ New + team** → **Start triage demo**). +4. Back in `/chat`, ask **"How many runs so far?"** and **"Summarize the latest + session"** — it queries live run history and the session snapshot. +5. Open **History** — the meta-chat's own runs are listed alongside the agents', + and each replays. + +## Server wiring + +- `POST /api/agent` — the AG-UI run stream (`createAgUiHandler`, spend ticks on). +- `GET|POST /api/harness/*` — the harness control tier (`createHarnessHandler`): + `snapshot`, `events`, `control`. +- `GET /api/hosts`, `GET /api/sessions` — host and session lists. +- `GET|POST /api/config` — read/write agent config (versioned). +- `GET /api/runs` — run history; `GET /api/replay` — a session's stored events. +- `POST /api/meta` — the meta-chat's AG-UI run stream. + +## State + +Server state is durable, so you can close the tab, restart the server, and find +each team where you left it. It lives in one JSON file (`.data/state.json`, +gitignored; override with `DASHBOARD_STATE_FILE`), written through on every +mutation and replayed on boot (`src/server/store.ts`). Persisted: chat run state +(messages, runs, interrupts, agent config) plus the dashboard's side tables +(threads, pod memory, schedules, webhooks). Not persisted: the in-flight +injection queue and the dev offline toggle. It's a single-process file store — a +multi-node dashboard would swap in a real database behind the same seam. + +## Test + +```bash +pnpm --filter agent-dashboard test:e2e # Playwright: stream + approve mid-run +``` + +Generated with Claude Code. diff --git a/examples/agent-dashboard/e2e/control-plane.spec.ts b/examples/agent-dashboard/e2e/control-plane.spec.ts new file mode 100644 index 0000000000..b5a58031fd --- /dev/null +++ b/examples/agent-dashboard/e2e/control-plane.spec.ts @@ -0,0 +1,56 @@ +import { expect, test } from '@playwright/test' + +const runInput = (threadId: string) => ({ + data: { + threadId, + runId: `${threadId}-r1`, + messages: [{ id: 'u1', role: 'user', content: 'handle it' }], + tools: [], + context: [], + }, +}) + +test('config form reads and writes ConfigOption schemas', async ({ page }) => { + await page.goto('/config') + const tone = page.getByLabel('tone') + await expect(tone).toBeVisible() + await tone.selectOption('formal') + + // Written through the harness protocol and persisted. + await page.waitForTimeout(500) + const res = await page.request.get('/api/config?threadId=settings') + const body = await res.json() + const toneValue = body.options.find((o: { key: string }) => o.key === 'tone') + .value + expect(toneValue).toBe('formal') +}) + +test('spend dashboard shows live token usage after a run', async ({ page }) => { + const threadId = `spend-${Date.now()}` + await page.goto(`/sessions/${threadId}`) + await page.getByRole('button', { name: 'Start triage demo' }).click() + await expect(page.getByText('Approval required', { exact: true })).toBeVisible() + + // Client-side nav (no reload) so the in-memory TanStack DB projection survives. + await page.getByRole('link', { name: 'Spend' }).click() + await expect(page.getByText(/tokens across/)).toBeVisible() + await expect(page.getByText(threadId).first()).toBeVisible() +}) + +test('run history lists a run and replays it into the session view', async ({ + page, +}) => { + const threadId = `replay-${Date.now()}` + // Run server-side so the browser has no local state for this thread. + await page.request.post('/api/agent', runInput(threadId)) + + await page.goto('/history') + await expect(page.getByText(threadId).first()).toBeVisible() + + // Open the session fresh — it rehydrates from stored events (replay). + await page.goto(`/sessions/${threadId}`) + await expect(page.getByText("I'll pull up that ticket first.")).toBeVisible() + await expect(page.getByText('lookup_ticket').first()).toBeVisible() + // The pending approval is restored from the live snapshot. + await expect(page.getByText('Approval required', { exact: true })).toBeVisible() +}) diff --git a/examples/agent-dashboard/e2e/dashboard.spec.ts b/examples/agent-dashboard/e2e/dashboard.spec.ts new file mode 100644 index 0000000000..fb309066b0 --- /dev/null +++ b/examples/agent-dashboard/e2e/dashboard.spec.ts @@ -0,0 +1,32 @@ +import { expect, test } from '@playwright/test' + +// The dashboard embeds a deterministic support-triage agent, so this runs with +// no API key: send a prompt, watch the AG-UI stream, approve a tool call +// mid-run, and see the run resume — the north-star approval demo. +test('streams a session and approves a tool call mid-run', async ({ page }) => { + const threadId = `e2e-${Date.now()}` + await page.goto(`/sessions/${threadId}`) + + const startButton = page.getByRole('button', { name: 'Start triage demo' }) + await expect(startButton).toBeVisible() + await startButton.click() + + // The agent looks up the ticket (auto tool) then drafts a reply for approval. + await expect(page.getByText('lookup_ticket').first()).toBeVisible() + await expect( + page.getByText("Here's a draft reply for your approval."), + ).toBeVisible() + + // The approval card appears (the run paused on the interrupt). + await expect(page.getByText('Approval required', { exact: true })).toBeVisible() + await expect(page.getByText('send_reply').first()).toBeVisible() + + // Spend meter is live (tokens accrued from the stream). + await expect(page.getByText(/[1-9][0-9,]* tokens/)).toBeVisible() + + // Approve via the AG-UI resume flow; the run continues and finishes. + await page.getByRole('button', { name: 'Approve', exact: true }).click() + + await expect(page.getByText(/Sent ✅/)).toBeVisible() + await expect(page.getByText('Approval required', { exact: true })).toHaveCount(0) +}) diff --git a/examples/agent-dashboard/e2e/devtools.ts b/examples/agent-dashboard/e2e/devtools.ts new file mode 100644 index 0000000000..6a58024033 --- /dev/null +++ b/examples/agent-dashboard/e2e/devtools.ts @@ -0,0 +1,34 @@ +import type { Page } from '@playwright/test' + +// The demo controls now live in a custom TanStack DevTools panel (see +// src/components/demo-controls.tsx). Under E2E the panel opens on load +// (VITE_E2E, see playwright.config.ts), so demo controls are reachable by +// default. The open panel is a fixed overlay along the bottom, so it can cover +// bottom-of-page product controls (the message box, roster "run" buttons) — +// close it before clicking those, reopen before the next demo control. + +const trigger = (page: Page) => + page.getByRole('button', { name: 'Open TanStack Devtools' }) + +/** Open the demo-controls panel if it's currently closed. */ +export async function openDemo(page: Page) { + // When the panel is open the trigger is hidden; if it's visible, we're closed. + if ( + await trigger(page) + .isVisible() + .catch(() => false) + ) { + await trigger(page).click() + } +} + +/** Close the panel so bottom-of-page product controls are clickable. */ +export async function closeDemo(page: Page) { + if ( + await trigger(page) + .isVisible() + .catch(() => false) + ) + return // already closed + await page.locator('button.close').first().click() +} diff --git a/examples/agent-dashboard/e2e/injection.spec.ts b/examples/agent-dashboard/e2e/injection.spec.ts new file mode 100644 index 0000000000..65a59c217b --- /dev/null +++ b/examples/agent-dashboard/e2e/injection.spec.ts @@ -0,0 +1,83 @@ +import { expect, test } from '@playwright/test' +import { openDemo } from './devtools' + +// Phase 2: the dashboard invokes work deterministically — run-now, timers, and +// webhooks — with the structured result streaming into the channel. These share +// server-side state (the scheduler + the offline flag), so run them serially. +test.describe.configure({ mode: 'serial' }) + +async function newTeam(page: import('@playwright/test').Page) { + await page.goto('/') + // The team-creation demos now live in the Demo Controls devtools panel. + await openDemo(page) + await page.getByRole('button', { name: '+ New team' }).click() + await expect(page).toHaveURL(/\/teams\//) + // The automations panel is present once the channel view mounts. + await expect(page.getByText('Automations')).toBeVisible() +} + +test('run-now executes a public tool; private tools are never listed', async ({ + page, +}) => { + await newTeam(page) + + // Only the public tool (fetch_stats) is offered — the reply tools stay private. + await expect( + page.getByRole('button', { name: 'run fetch_stats' }), + ).toBeVisible() + await expect( + page.getByRole('button', { name: /run send_reply/ }), + ).toHaveCount(0) + await expect( + page.getByRole('button', { name: /run lookup_ticket/ }), + ).toHaveCount(0) + + // Run it now — the result streams back as a structured, injected card. + await page.getByRole('button', { name: 'run fetch_stats' }).click() + await expect(page.getByText('⏵ manual')).toBeVisible() + await expect(page.getByText('fetch_stats').first()).toBeVisible() +}) + +test('a scheduled timer fires a public tool into the channel', async ({ + page, +}) => { + await newTeam(page) + + // Fire every second so the test doesn't wait on a slow interval. + await page.getByLabel('schedule interval seconds').fill('1') + await page.getByRole('button', { name: '+ Add schedule' }).click() + // The timer (dashboard-owned clock) fires within a couple of seconds. + await expect(page.getByText('⏵ timer').first()).toBeVisible({ + timeout: 10000, + }) + + // Pause it so it doesn't keep firing after the test. + await page.getByRole('button', { name: 'pause' }).first().click() +}) + +test('a webhook produces the same injected result as the other triggers', async ({ + page, +}) => { + await newTeam(page) + + await page.getByRole('button', { name: 'Send test webhook' }).click() + await expect(page.getByText('⏵ webhook')).toBeVisible() +}) + +test('an injected job queues while the host is offline and flushes on reconnect', async ({ + page, +}) => { + await newTeam(page) + + await page.getByRole('button', { name: 'Simulate host offline' }).click() + await expect(page.getByText(/host offline/)).toBeVisible() + + // Injecting while offline queues instead of running — no card appears yet. + await page.getByRole('button', { name: 'run fetch_stats' }).click() + await expect(page.getByText(/1 queued/)).toBeVisible() + await expect(page.getByText('⏵ manual')).toHaveCount(0) + + // Reconnect → the queue flushes → the result appears. + await page.getByRole('button', { name: 'Bring host online' }).click() + await expect(page.getByText('⏵ manual')).toBeVisible({ timeout: 8000 }) +}) diff --git a/examples/agent-dashboard/e2e/meta-chat.spec.ts b/examples/agent-dashboard/e2e/meta-chat.spec.ts new file mode 100644 index 0000000000..a4ebfcfde4 --- /dev/null +++ b/examples/agent-dashboard/e2e/meta-chat.spec.ts @@ -0,0 +1,24 @@ +import { expect, test } from '@playwright/test' + +// The meta-chat is the dashboard's own tool-using agent. A quick prompt makes +// it call a tool over live state and summarize the result, with the tool call +// visible inline (dogfooding the trace view). +test('meta-chat calls a tool over live state and summarizes', async ({ + page, +}) => { + await page.goto('/chat') + + await page + .getByRole('button', { name: 'List the agents on this host' }) + .click() + + // The tool call is visible in the trace… + await expect(page.getByText('list_agents').first()).toBeVisible() + // …and the agent summarizes the real result: triage, meta, the two PR-watcher + // demo agents, and the two Reddit-pod agents are registered on this host. + await expect(page.getByText(/host runs 6 agent/)).toBeVisible() + + // Its run shows up in history like any other agent. + await page.getByRole('link', { name: 'History' }).click() + await expect(page.getByText(/meta-/).first()).toBeVisible() +}) diff --git a/examples/agent-dashboard/e2e/persistence.spec.ts b/examples/agent-dashboard/e2e/persistence.spec.ts new file mode 100644 index 0000000000..f24422eece --- /dev/null +++ b/examples/agent-dashboard/e2e/persistence.spec.ts @@ -0,0 +1,36 @@ +import { expect, test } from '@playwright/test' +import { openDemo } from './devtools' + +// Server-side persistence: a team, its runs, and its resolved approvals survive a +// full page reload (the roster rehydrates from the server; the session replays +// from stored events). Regression guard for two bugs: +// 1. the team vanished on reload (roster was client-only), and +// 2. an already-approved tool call resurfaced as "Approval required" because the +// tail replayed the historical interrupt as if it were live. +test('a team and its approved run survive a reload', async ({ page }) => { + await page.goto('/') + + // Create a triage team and run it to the approval (demos live in the panel). + await openDemo(page) + await page.getByRole('button', { name: '+ New team' }).click() + await expect(page).toHaveURL(/\/teams\//) + await page.getByRole('button', { name: 'Start triage demo' }).click() + await expect( + page.getByText('Approval required', { exact: true }).first(), + ).toBeVisible() + + // Approve via the AG-UI resume flow; the run continues and finishes. + await page.getByRole('button', { name: 'Approve', exact: true }).click() + await expect(page.getByText(/Sent ✅/)).toBeVisible() + await expect( + page.getByText('Approval required', { exact: true }), + ).toHaveCount(0) + + // Reload: the team must reappear (roster is persisted) and the finished run must + // replay — WITHOUT the resolved approval coming back. + await page.reload() + await expect(page.getByText(/Sent ✅/)).toBeVisible() + await expect( + page.getByText('Approval required', { exact: true }), + ).toHaveCount(0) +}) diff --git a/examples/agent-dashboard/e2e/pr-watcher.spec.ts b/examples/agent-dashboard/e2e/pr-watcher.spec.ts new file mode 100644 index 0000000000..49d1940e82 --- /dev/null +++ b/examples/agent-dashboard/e2e/pr-watcher.spec.ts @@ -0,0 +1,118 @@ +import { expect, test } from '@playwright/test' +import { closeDemo, openDemo } from './devtools' +import type { Page } from '@playwright/test' + +// Phase 3: the PR-watcher "the pod learns" loop. A webhook opens a per-PR channel, +// a subscribed security agent reviews it, the human corrects it, the correction is +// persisted to pod memory, and the next PR is handled better. These share +// server-side state (the seeded PR sequence, pod memory), so run serially. +test.describe.configure({ mode: 'serial' }) + +async function newPrWatcherTeam(page: Page) { + await page.goto('/') + // The seeded demos live in the Demo Controls devtools panel now. + await openDemo(page) + await page.getByRole('button', { name: '+ PR-watcher demo' }).click() + await expect(page).toHaveURL(/\/teams\//) + await expect(page.getByText('Members · 2')).toBeVisible() + await expect(page.getByText('Automations')).toBeVisible() +} + +test('the PR-watcher loop: review, correct, remember, handle the next PR better', async ({ + page, +}) => { + await newPrWatcherTeam(page) + + // 1–2. Webhook in → the watcher opens a per-PR channel and posts a summary. + await page.getByRole('button', { name: 'Send PR webhook' }).click() + await expect(page.getByText(/Channel #pr-\d+ created/)).toBeVisible() + // The new channel shows up in the sidebar; open it. + const firstChannel = page.getByRole('button', { name: /# pr-\d+/ }).first() + await expect(firstChannel).toBeVisible() + const firstChannelName = ((await firstChannel.textContent()) ?? '').trim() + await firstChannel.click() + + // 3. The watcher's summary and the security agent's finding are both in-channel, + // and the finding flags the public-internet exposure (no memory yet). + await expect(page.getByText(/Requesting a security review/)).toBeVisible() + await expect(page.getByText(/Security finding/)).toBeVisible() + // Match the finding message specifically (the PR risk also mentions "public + // internet" but lives in the collapsed tool-result JSON tree). + await expect( + page.getByText(/exposes an endpoint to the public internet/), + ).toBeVisible() + + // 4. The human corrects it. The agent persists the standing instruction. + // Close the demo panel so it doesn't cover the message box / Send button. + await closeDemo(page) + await page + .getByPlaceholder('Send a message…') + .fill("not a security problem — we're intranet-only here") + await page.getByRole('button', { name: 'Send', exact: true }).click() + + // The write is a visible tool call; its args render as a collapsed JSON tree, + // so the key is present in the DOM but not visible until expanded. + await expect(page.getByText('pod.memory_write').first()).toBeVisible() + await expect(page.locator('body')).toContainText('intranet-policy') + + // 5. The next PR: a second webhook opens a new channel; this time the review is + // clean and the run carries the "memory attached" badge. + await page.getByRole('button', { name: /# main/ }).click() + // Reopen the demo panel to reach the webhook button again. + await openDemo(page) + await page.getByRole('button', { name: 'Send PR webhook' }).click() + + // A second, different PR channel appears; open it. + const secondChannel = page + .getByRole('button', { name: /# pr-\d+/ }) + .filter({ hasNotText: firstChannelName.replace(/^#\s*/, '') }) + await expect(secondChannel.first()).toBeVisible({ timeout: 15000 }) + await secondChannel.first().click() + + // The pod learned: no intranet finding this time, and memory was attached. + await expect(page.getByText(/intranet-only per policy/)).toBeVisible() + await expect(page.getByText(/memory entr(y|ies) attached/)).toBeVisible() + await expect(page.getByText(/Security finding/)).toHaveCount(0) +}) + +test('a DM with an agent is created and routes messages to it', async ({ + page, +}) => { + await newPrWatcherTeam(page) + + // Create a DM with the security member from the roster (a page control, so + // close the demo panel that would otherwise cover it). + await closeDemo(page) + await page.getByRole('button', { name: /New DM with security/ }).click() + + // A dm channel appears in the sidebar; open it. + const dm = page.getByRole('button', { name: /@ dm-/ }) + await expect(dm).toBeVisible({ timeout: 15000 }) + await dm.click() + + // The DM seats the security agent as its member, so a message reaches it and + // it replies (regression: DMs used to be created with no members → no reply). + await page.getByPlaceholder('Send a message…').fill('please review this') + await page.getByRole('button', { name: 'Send', exact: true }).click() + await expect(page.getByText(/Security finding/).first()).toBeVisible({ + timeout: 15000, + }) +}) + +test('the memory panel adds and removes entries', async ({ page }) => { + await newPrWatcherTeam(page) + // Memory now lives on the team page (one panel per agent); close the demo + // overlay and scope to the watcher's panel. + await closeDemo(page) + const mem = page.getByRole('group', { name: 'memory pr-watcher' }) + await expect(mem.getByRole('heading', { name: /Memory ·/ })).toBeVisible() + await mem.getByLabel('memory key').fill('watchlist') + await mem.getByLabel('memory value').fill('repo:tanstack/ai') + await mem.getByRole('button', { name: '+ Add entry' }).click() + + await expect(mem.getByText('watchlist')).toBeVisible() + await expect(mem.getByText('repo:tanstack/ai')).toBeVisible() + + await mem.getByRole('button', { name: 'delete memory watchlist' }).click() + await expect(mem.getByText('repo:tanstack/ai')).toHaveCount(0) +}) diff --git a/examples/agent-dashboard/e2e/react-news.spec.ts b/examples/agent-dashboard/e2e/react-news.spec.ts new file mode 100644 index 0000000000..6c0db5ee1a --- /dev/null +++ b/examples/agent-dashboard/e2e/react-news.spec.ts @@ -0,0 +1,86 @@ +import { expect, test } from '@playwright/test' +import { closeDemo, openDemo } from './devtools' + +// The Reddit pod — the first real-service + real-AI team. A procedural fetcher +// tool pulls React news (served from a recorded fixture under VITE_E2E, so this +// runs offline and deterministically), and a sentiment agent — subscribed to +// the tool's *result* — posts a digest with nobody clicking run on it. +// +// This also exercises the product controls: the demo team is seeded from the +// Demo Controls devtools panel, and the fetch is driven from the roster's +// per-agent "run tool" dialog. It asserts the LOOP structurally (fetch → result +// card → sentiment message), not specific live headlines (see design doc §10). +test('the Reddit pod loop: a news batch triggers an unprompted sentiment digest', async ({ + page, +}) => { + await page.goto('/') + // The seeded demos live in the Demo Controls devtools panel now. + await openDemo(page) + await page.getByRole('button', { name: '+ React-news demo' }).click() + await expect(page).toHaveURL(/\/teams\//) + + // Two members: the fetcher and the sentiment agent. + await expect(page.getByText('Members · 2')).toBeVisible() + + // Close the panel and drive the fetch from the roster's run-tool dialog (the + // product control): open it on the fetcher, then Run its default tool. + await closeDemo(page) + await page.getByRole('button', { name: 'Run a tool on fetcher' }).click() + await expect(page.getByRole('heading', { name: /Run a tool/ })).toBeVisible() + await page.getByRole('button', { name: 'Run', exact: true }).click() + // Dismiss the dialog so it doesn't cover the channel. + await page.getByRole('button', { name: 'Close' }).click() + + // 1. The batch lands as a result card with (fixtured) real headlines. + // The result JSON renders as a collapsed tree, so the headline is present in + // the DOM but not visible until expanded — assert content, not visibility. + await expect(page.locator('body')).toContainText( + /React Compiler is now stable/, + { timeout: 15000 }, + ) + + // 2. The sentiment agent posts a digest — unprompted, driven by the + // tool_result subscription. + await expect(page.getByText(/Sentiment digest/).first()).toBeVisible({ + timeout: 15000, + }) +}) + +// Regression: a team COMPOSED BY HAND (agents table → New team, then + Add +// agent) must react the same as the seeded demo — the sentiment agent carries +// its tool_result subscription by harness default, so nobody has to wire it. +test('a hand-composed Reddit team auto-reacts to a fetch', async ({ page }) => { + await page.goto('/') + // Close the demo panel so it doesn't overlay the lower agents-table rows. + await closeDemo(page) + + // New team from the agents table with reddit/fetcher. + const fetcherRow = page.getByRole('row').filter({ hasText: 'reddit/fetcher' }) + await fetcherRow.getByRole('button', { name: /Add to team/ }).click() + await page.getByRole('button', { name: /New team with this agent/ }).click() + await expect(page).toHaveURL(/\/teams\//) + + // Add the sentiment agent from the team-page product control. + await page.getByRole('button', { name: /Add agent/ }).click() + await page + .getByRole('button', { name: 'sentiment/react', exact: true }) + .click() + await expect(page.getByText('Members · 2')).toBeVisible() + + // Run the fetch from the roster — the sentiment agent reacts with no manual + // subscription wiring. + await page.getByRole('button', { name: 'Run a tool on fetcher' }).click() + await expect(page.getByRole('heading', { name: /Run a tool/ })).toBeVisible() + await page.getByRole('button', { name: 'Run', exact: true }).click() + await page.getByRole('button', { name: 'Close' }).click() + + // The result JSON renders as a collapsed tree, so the headline is present in + // the DOM but not visible until expanded — assert content, not visibility. + await expect(page.locator('body')).toContainText( + /React Compiler is now stable/, + { timeout: 15000 }, + ) + await expect(page.getByText(/Sentiment digest/).first()).toBeVisible({ + timeout: 15000, + }) +}) diff --git a/examples/agent-dashboard/e2e/team.spec.ts b/examples/agent-dashboard/e2e/team.spec.ts new file mode 100644 index 0000000000..a0a79e3982 --- /dev/null +++ b/examples/agent-dashboard/e2e/team.spec.ts @@ -0,0 +1,48 @@ +import { expect, test } from '@playwright/test' +import { closeDemo, openDemo } from './devtools' + +// The teams reframe: one agent looks like a plain chat; a second member reveals +// the team (roster appears), and both members' AG-UI streams merge into ONE +// channel view. Exercises the product controls: the home agents table's +// "Add to team" and the team page's "+ Add agent" picker. +test('a second agent reveals the team and both members share one channel', async ({ + page, +}) => { + await page.goto('/') + + // Start a team from the agents table: support/triage → New team. + const triageRow = page.getByRole('row').filter({ hasText: 'support/triage' }) + await triageRow.getByRole('button', { name: /Add to team/ }).click() + await page.getByRole('button', { name: /New team with this agent/ }).click() + await expect(page).toHaveURL(/\/teams\//) + + // One member: it looks like a normal chat — no team roster. + await expect(page.getByText(/Members ·/)).toHaveCount(0) + + // Run the first agent from the Demo Controls panel: it looks up the ticket and + // pauses for approval. + await openDemo(page) + await page.getByRole('button', { name: 'Start triage demo' }).click() + await expect(page.getByText('lookup_ticket').first()).toBeVisible() + await expect( + page.getByText('Approval required', { exact: true }).first(), + ).toBeVisible() + + // Add a second agent from the team page's product control → the roster appears. + await closeDemo(page) + await page.getByRole('button', { name: /Add agent/ }).click() + await page + .getByRole('button', { name: 'support/triage', exact: true }) + .click() + await expect(page.getByText('Members · 2')).toBeVisible() + + // Run the second member from the roster; its stream joins the SAME channel. + await page.getByRole('button', { name: 'Run triage 2' }).click() + + // Both members' runs are visible in one timeline — two lookup_ticket cards and + // two approvals. If the two threads' row ids collided, we'd see only one each. + await expect(page.getByText('lookup_ticket')).toHaveCount(2) + await expect( + page.getByText('Approval required', { exact: true }), + ).toHaveCount(2) +}) diff --git a/examples/agent-dashboard/package.json b/examples/agent-dashboard/package.json new file mode 100644 index 0000000000..7fe362728b --- /dev/null +++ b/examples/agent-dashboard/package.json @@ -0,0 +1,48 @@ +{ + "name": "agent-dashboard", + "private": true, + "type": "module", + "scripts": { + "dev": "vite dev --port 3002", + "build": "vite build", + "serve": "vite preview", + "start": "node .output/server/index.mjs", + "test": "vitest run", + "test:lib": "vitest run", + "test:e2e": "playwright test", + "test:types": "tsc --noEmit" + }, + "dependencies": { + "@ag-ui/client": "1.0.0", + "@tailwindcss/vite": "^4.1.18", + "@tanstack/ai": "workspace:*", + "@tanstack/ai-anthropic": "workspace:*", + "@tanstack/ai-harness": "workspace:*", + "@tanstack/ai-persistence": "workspace:*", + "@tanstack/react-db": "^0.1.55", + "@tanstack/react-devtools": "^0.9.10", + "@tanstack/react-query": "^5.90.12", + "@tanstack/react-router": "^1.158.4", + "@tanstack/react-router-devtools": "^1.158.4", + "@tanstack/react-start": "^1.159.0", + "@tanstack/router-plugin": "^1.158.4", + "nitro": "3.0.260610-beta", + "react": "^19.2.3", + "react-dom": "^19.2.3", + "react-markdown": "^10.1.0", + "remark-gfm": "^4.0.1", + "tailwindcss": "^4.1.18", + "zod": "^4.2.0" + }, + "devDependencies": { + "@playwright/test": "^1.57.0", + "@tanstack/devtools-vite": "^0.5.3", + "@types/node": "^24.10.1", + "@types/react": "^19.2.7", + "@types/react-dom": "^19.2.3", + "@vitejs/plugin-react": "^5.2.0", + "typescript": "5.9.3", + "vite": "^8.2.1", + "vitest": "^4.1.10" + } +} diff --git a/examples/agent-dashboard/playwright.config.ts b/examples/agent-dashboard/playwright.config.ts new file mode 100644 index 0000000000..60dfaf7af8 --- /dev/null +++ b/examples/agent-dashboard/playwright.config.ts @@ -0,0 +1,29 @@ +import { defineConfig, devices } from '@playwright/test' + +export default defineConfig({ + testDir: './e2e', + fullyParallel: false, + forbidOnly: !!process.env.CI, + retries: 0, + workers: 1, + reporter: [['list']], + timeout: 30_000, + expect: { timeout: 15_000 }, + use: { + baseURL: 'http://localhost:3002', + screenshot: 'only-on-failure', + trace: 'on-first-retry', + }, + projects: [{ name: 'chromium', use: { ...devices['Desktop Chrome'] } }], + webServer: { + // Isolate persisted state to a throwaway file, wiped on start, so runs never + // inherit a previous run's runs/memory/automations (see src/server/store.ts). + // VITE_E2E opens the devtools panel on load so specs can reach the demo + // controls that now live in the "Demo Controls" panel. + command: + 'rm -f .data/e2e-state.json && VITE_E2E=1 DASHBOARD_STATE_FILE=.data/e2e-state.json pnpm run dev', + url: 'http://localhost:3002', + reuseExistingServer: !process.env.CI, + timeout: 120_000, + }, +}) diff --git a/examples/agent-dashboard/src/components/automations-panel.tsx b/examples/agent-dashboard/src/components/automations-panel.tsx new file mode 100644 index 0000000000..2c1645eeeb --- /dev/null +++ b/examples/agent-dashboard/src/components/automations-panel.tsx @@ -0,0 +1,281 @@ +/** + * The channel's automations: the tool registry (public tools only, with run-now), + * the schedule table (the dashboard's clock), and a webhook tester. All three + * produce the same thing — an injected tool call whose structured result streams + * into the channel via the feed tail. + */ +import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query' +import { useState } from 'react' +import { runInjection } from '@/lib/session-controller' +import type { MembershipRow } from '@/db/collections' + +interface ToolInfo { + name: string + description: string +} +interface ScheduleRow { + id: string + tool: string + everySeconds?: number + cron?: string + enabled: boolean + nextFire?: number +} + +export function AutomationsPanel({ + channelId, + primary, +}: { + channelId: string + primary: MembershipRow +}) { + const qc = useQueryClient() + const invalidate = () => { + void qc.invalidateQueries({ queryKey: ['schedules', channelId] }) + void qc.invalidateQueries({ queryKey: ['offline'] }) + } + + const tools = useQuery<{ tools: Array }>({ + // Pass the harness so the registry is resolved directly — otherwise the + // thread may not yet be noted with its harness and we'd get triage's tools. + queryKey: ['tools', primary.threadId, primary.harness], + queryFn: () => + fetch( + `/api/tools?threadId=${primary.threadId}&harness=${encodeURIComponent(primary.harness)}`, + ).then((r) => r.json()), + }) + const schedules = useQuery<{ schedules: Array }>({ + queryKey: ['schedules', channelId], + queryFn: () => + fetch(`/api/schedules?channelId=${channelId}`).then((r) => r.json()), + refetchInterval: 2000, + }) + const offline = useQuery<{ offline: boolean; queued: number }>({ + queryKey: ['offline'], + queryFn: () => fetch('/api/dev/offline').then((r) => r.json()), + refetchInterval: 1500, + }) + + const [everySeconds, setEverySeconds] = useState(10) + const [tool, setTool] = useState('') + const toolNames = tools.data?.tools ?? [] + const selectedTool = tool || toolNames[0]?.name || '' + + const addSchedule = useMutation({ + mutationFn: () => + fetch('/api/schedules', { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ + threadId: primary.threadId, + channelId, + tool: selectedTool, + everySeconds, + }), + }).then((r) => r.json()), + onSuccess: invalidate, + }) + const toggleSchedule = useMutation({ + mutationFn: (row: ScheduleRow) => + fetch('/api/schedules', { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ id: row.id, enabled: !row.enabled }), + }).then((r) => r.json()), + onSuccess: invalidate, + }) + const deleteSchedule = useMutation({ + mutationFn: (id: string) => + fetch(`/api/schedules?id=${id}`, { method: 'DELETE' }).then((r) => + r.json(), + ), + onSuccess: invalidate, + }) + const setOffline = useMutation({ + mutationFn: (value: boolean) => + fetch('/api/dev/offline', { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ offline: value }), + }).then((r) => r.json()), + onSuccess: invalidate, + }) + const sendWebhook = useMutation({ + mutationFn: async () => { + const created = await fetch('/api/webhooks', { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ + threadId: primary.threadId, + channelId, + tool: selectedTool, + argMapping: { queue: 'queue' }, + }), + }).then((r) => r.json()) + const token = created.webhook.token as string + return fetch(`/api/webhooks/${token}`, { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ queue: 'from-webhook' }), + }).then((r) => r.json()) + }, + }) + // The PR-watcher demo: a prompt-mode webhook drives the watcher's full run + // (check_pr → channel_create → message_post). Each send is the next seeded PR. + const sendPrWebhook = useMutation({ + mutationFn: async () => { + const created = await fetch('/api/webhooks', { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ + threadId: primary.threadId, + channelId, + harness: primary.harness, + mode: 'prompt', + message: + 'A PR webhook arrived. Check for a new PR and open a review channel.', + }), + }).then((r) => r.json()) + const token = created.webhook.token as string + return fetch(`/api/webhooks/${token}`, { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ event: 'pull_request' }), + }).then((r) => r.json()) + }, + }) + + return ( +
+
+

Automations

+ + deterministic tool runs — no tokens + + {offline.data?.offline && ( + + host offline — {offline.data.queued} queued + + )} +
+ + {/* Tool registry + run-now (public tools only) */} +
+
+ Public tools +
+
+ {toolNames.length === 0 && ( + none + )} + {toolNames.map((t) => ( + + ))} +
+
+ + {/* Schedule table */} +
+
+ Schedules +
+
+ + + +
+
    + {(schedules.data?.schedules ?? []).map((row) => ( +
  • + {row.tool} + + {row.everySeconds + ? `every ${row.everySeconds}s` + : (row.cron ?? '')} + + + +
  • + ))} +
+
+ + {/* Webhook + offline simulation */} +
+ + {primary.harness === 'ops/pr-watcher' && ( + + )} + +
+
+ ) +} diff --git a/examples/agent-dashboard/src/components/channel-view.tsx b/examples/agent-dashboard/src/components/channel-view.tsx new file mode 100644 index 0000000000..29fa4555bc --- /dev/null +++ b/examples/agent-dashboard/src/components/channel-view.tsx @@ -0,0 +1,622 @@ +/** + * The shared channel view. It unions every member thread of a channel into one + * timeline (a live query over the collections keyed by `channelId`). With one + * member it looks exactly like a single-agent chat; when a second member joins, + * team chrome (the member list, per-agent attribution) appears — the "second + * agent reveals the team" moment. + * + * A channel's members are its team roster for the `main` channel, or the opt-in + * `channelMembers` for a `dynamic`/`dm` channel. Attribution is resolved against + * the whole team roster either way. + */ +import { eq, useLiveQuery } from '@tanstack/react-db' +import { useQuery } from '@tanstack/react-query' +import { useEffect, useState } from 'react' +import { + approvals, + channelMembers, + channels, + memberships, + messages, + runMeta, + sessions, + spend, + toolCalls, + uiState, + upsert, +} from '@/db/collections' +import { + addAgentToChannel, + channelSendPrompt, + createDm, + defaultSubscriptions, + openChannelMember, + resolveApproval, +} from '@/lib/session-controller' +import { MemberList } from '@/components/member-list' +import { MemoryPanel } from '@/components/memory-panel' +import { JsonTree, tryParse } from '@/components/json-tree' +import { Markdown } from '@/components/markdown' +import type { + ApprovalRow, + ChannelMemberRow, + ChannelRow, + MembershipRow, + MessageRow, + RunMetaRow, + SessionRow, + ToolCallRow, + UiStateRow, +} from '@/db/collections' + +// `pod.channel_create` / `pod.message_post` are realized as a channel and a +// message respectively, so their raw tool cards are hidden (they'd duplicate). +// `pod.memory_write` stays visible (the write is auditable mechanics). +const HIDDEN_TOOL_CARDS = new Set([ + 'pod.channel_create', + 'pod.message_post', + 'pod.memory_read', +]) + +export function ChannelView({ + channelId, + teamId, +}: { + channelId: string + teamId?: string +}) { + const [input, setInput] = useState('') + + const { data: chanRows = [] } = useLiveQuery( + (q) => q.from({ c: channels }).where(({ c }) => eq(c.id, channelId)), + [channelId], + ) + const channel = (chanRows as Array)[0] + const resolvedTeamId = teamId ?? channel?.teamId + + // The whole team roster (for attribution + main-channel membership). + const { data: roster = [] } = useLiveQuery( + (q) => + q + .from({ m: memberships }) + .where(({ m }) => eq(m.teamId, resolvedTeamId ?? '')), + [resolvedTeamId], + ) + const rosterRows = roster as Array + // Opt-in members for a non-main channel. + const { data: chanMembers = [] } = useLiveQuery( + (q) => + q + .from({ cm: channelMembers }) + .where(({ cm }) => eq(cm.channelId, channelId)), + [channelId], + ) + const chanMemberRows = chanMembers as Array + + const isMain = channel?.kind !== 'dynamic' && channel?.kind !== 'dm' + const memberRows: Array = isMain + ? rosterRows + : chanMemberRows + .map((cm) => rosterRows.find((m) => m.agentId === cm.agentId)) + .filter((m): m is MembershipRow => Boolean(m)) + + // Open a live tail for every team member so background runs (a watcher firing, + // a subscribed agent reviewing) project even when we're not looking at them. + const rosterKey = rosterRows.map((m) => m.id).join(',') + useEffect(() => { + for (const m of rosterRows) { + openChannelMember({ + channelId: m.channelId, + agentId: m.agentId, + threadId: m.threadId, + teamId: m.teamId, + harness: m.harness, + role: m.role, + displayName: m.displayName, + }) + } + // eslint-disable-next-line react-hooks/exhaustive-deps + }, [rosterKey]) + + // Publish the active channel so the demo-controls devtools panel (rendered + // out of the route tree) knows which channel to drive. Clear on unmount + // unless another channel already took over. + useEffect(() => { + upsert(uiState, { id: 'active' }, (d) => { + d.channelId = channelId + d.teamId = resolvedTeamId + }) + return () => { + if (uiState.get('active')?.channelId === channelId) { + upsert(uiState, { id: 'active' }, (d) => { + d.channelId = undefined + d.teamId = undefined + }) + } + } + }, [channelId, resolvedTeamId]) + + const { data: msgs = [] } = useLiveQuery( + (q) => q.from({ m: messages }).where(({ m }) => eq(m.channelId, channelId)), + [channelId], + ) + const { data: tools = [] } = useLiveQuery( + (q) => + q.from({ t: toolCalls }).where(({ t }) => eq(t.channelId, channelId)), + [channelId], + ) + const { data: apprs = [] } = useLiveQuery( + (q) => + q.from({ a: approvals }).where(({ a }) => eq(a.channelId, channelId)), + [channelId], + ) + const { data: spendRows = [] } = useLiveQuery( + (q) => q.from({ s: spend }).where(({ s }) => eq(s.channelId, channelId)), + [channelId], + ) + const { data: sess = [] } = useLiveQuery( + (q) => q.from({ s: sessions }).where(({ s }) => eq(s.channelId, channelId)), + [channelId], + ) + const { data: runMetaRows = [] } = useLiveQuery((q) => q.from({ r: runMeta })) + + const isTeam = memberRows.length > 1 + const nameByAgent = new Map(rosterRows.map((m) => [m.agentId, m.displayName])) + const statusByThread: Record = {} + for (const s of sess as Array) + statusByThread[s.threadId] = s.status + + const sessionRows = sess as Array + const status = sessionRows.some((s) => s.status === 'requires_action') + ? 'requires_action' + : sessionRows.some((s) => s.status === 'running') + ? 'running' + : 'idle' + const tokens = (spendRows as Array<{ totalTokens: number }>).reduce( + (sum, s) => sum + (s.totalTokens ?? 0), + 0, + ) + const pending = (apprs as Array).filter( + (a) => a.status === 'pending', + ) + + const timeline = [ + ...(msgs as Array).map((m) => ({ + kind: 'message' as const, + at: m.createdAt, + m, + })), + ...(tools as Array) + .filter((t) => !HIDDEN_TOOL_CARDS.has(t.name)) + .map((t) => ({ kind: 'tool' as const, at: t.createdAt, t })), + ].sort((a, b) => a.at - b.at) + + // Human input targets the primary agent member (broadcast is a later phase). + const primary = + memberRows.find((m) => m.role === 'agent') ?? memberRows[0] ?? undefined + + // How much pod memory the platform attached to this channel's agents' last run. + const attachedById = new Map( + (runMetaRows as Array).map((r) => [r.threadId, r.attached]), + ) + const attached = primary ? (attachedById.get(primary.threadId) ?? 0) : 0 + + const send = async (text: string) => { + const t = text.trim() + if (!t || !primary) return + setInput('') + await channelSendPrompt(primary, t, channelId) + } + + // A generic nudge to run a specific member. (The triage-specific demo prompt + // lives in the Demo Controls panel, not here — this must work for any agent.) + const runMember = (member: MembershipRow) => + channelSendPrompt(member, 'Please proceed.', channelId) + + const createDmWith = (member: MembershipRow) => { + if (!primary || member.agentId === primary.agentId) return + void createDm( + { + channelId: primary.channelId, + agentId: primary.agentId, + threadId: primary.threadId, + teamId: primary.teamId, + harness: primary.harness, + role: primary.role, + displayName: primary.displayName, + }, + member.agentId, + member.displayName, + ) + } + + // Toggle a member's subscription on/off. "On" restores the harness's default + // triggers (what it reacts to) rather than a hardcoded channel_created one. + const toggleSubscription = (member: MembershipRow) => { + const on = (member.subscriptions ?? []).length > 0 + memberships.update(member.id, (draft) => { + draft.subscriptions = on + ? [] + : (defaultSubscriptions(member.harness) ?? []) + }) + } + + const title = isMain + ? isTeam + ? 'main' + : (primary?.displayName ?? channel?.name ?? channelId) + : `#${channel?.name ?? channelId}` + + return ( +
+
+ + ← teams + +

{title}

+ {channel?.topic && ( + {channel.topic} + )} + + {status.replace('_', ' ')} + + {attached > 0 && ( + + 🧠 {attached} memory {attached === 1 ? 'entry' : 'entries'} attached + + )} + + {tokens.toLocaleString()} tokens + + {isMain && } +
+ + {pending.map((approval) => ( + ).find( + (t) => t.id === approval.toolCallId, + )} + /> + ))} + +
+ {isTeam && ( + + )} +
+ {timeline.length === 0 && ( +

+ No activity yet. Send a message below, or drive it from the Demo + controls devtools panel. +

+ )} + {timeline.map((entry) => + entry.kind === 'message' ? ( + entry.m.role === 'system' ? ( + + ) : ( + + ) + ) : ( + + ), + )} +
+
+ +
+ setInput(e.target.value)} + onKeyDown={(e) => e.key === 'Enter' && send(input)} + placeholder="Send a message…" + className="min-w-40 flex-1 rounded-md border border-white/15 bg-transparent px-3 py-2 text-sm outline-none focus:border-white/30" + /> + +
+ + {isTeam && isMain && ( +
+

+ Memory +

+
+ {memberRows + .filter((m) => m.role === 'agent') + .map((m) => ( + + ))} +
+
+ )} +
+ ) +} + +/** Add any available agent to this team (product control, main channel only). */ +function AddAgentControl({ channelId }: { channelId: string }) { + const [open, setOpen] = useState(false) + const hosts = useQuery< + Array<{ agents: Array<{ name: string; description: string }> }> + >({ + queryKey: ['hosts'], + queryFn: () => fetch('/api/hosts').then((r) => r.json()), + }) + const agents = (hosts.data ?? []).flatMap((h) => h.agents) + + return ( +
+ + {open && ( +
+ {agents.map((a) => ( + + ))} +
+ )} +
+ ) +} + +function AuthorTag({ author }: { author?: string }) { + if (!author) return null + return ( + + {author} + + ) +} + +function SystemCard({ message }: { message: MessageRow }) { + const icon = message.system?.kind === 'channel_created' ? '📢' : '👋' + return ( +
+ {icon} + {message.text} + {message.system?.topic && ( + — {message.system.topic} + )} +
+ ) +} + +function MessageBubble({ + message, + author, + showAuthor, +}: { + message: MessageRow + author?: string + showAuthor: boolean +}) { + const isUser = message.role === 'user' + return ( +
+
+ {showAuthor && !isUser && } + {message.subagentRunId && ( + + subagent + + )} + {message.text ? ( + isUser ? ( + message.text + ) : ( + {message.text} + ) + ) : ( + … + )} +
+
+ ) +} + +function ToolCard({ + tool, + author, + showAuthor, +}: { + tool: ToolCallRow + author?: string + showAuthor: boolean +}) { + // Injected tool calls (timer/manual/webhook) get a distinct border + badge. + const injected = Boolean(tool.trigger) + return ( +
+
+ {showAuthor && } + {injected && ( + + ⏵ {tool.trigger} + + )} + ⚙ {tool.name} + + {tool.status} + +
+ {tool.args && + (() => { + const parsed = tryParse(tool.args) + return parsed !== undefined ? ( +
+ +
+ ) : ( +
{tool.args}
+ ) + })()} + {tool.result && + (() => { + const parsed = tryParse(tool.result) + return parsed !== undefined ? ( +
+ +
+ ) : ( +
→ {tool.result}
+ ) + })()} + {tool.truncated && ( +
(result truncated)
+ )} +
+ ) +} + +function ApprovalCard({ + approval, + tool, +}: { + approval: ApprovalRow + tool?: ToolCallRow +}) { + const [editing, setEditing] = useState(false) + const [draft, setDraft] = useState(tool?.args ?? '{}') + const [busy, setBusy] = useState(false) + + const act = async (decision: 'approve' | 'deny', edited?: boolean) => { + setBusy(true) + let editedArgs: Record | undefined + if (edited) { + try { + editedArgs = JSON.parse(draft) + } catch { + setBusy(false) + return + } + } + await resolveApproval(approval.threadId, approval.id, decision, editedArgs) + setBusy(false) + } + + return ( +
+
+ 🔔 + Approval required + {tool && ( + + {tool.name} + + )} +
+

{approval.message}

+ {editing ? ( +