Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
84 changes: 30 additions & 54 deletions .agents/skills/kane-cli/SKILL.md

Large diffs are not rendered by default.

6 changes: 3 additions & 3 deletions .agents/skills/kane-cli/references/assurance-parsing.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,8 +41,8 @@ A stream that ends **without** `done` means the process crashed — outcome unkn
| `held` / `update_held` *(0.7.1+)* | items held for the user's review instead of committed: `source_id` + `count` + `reason` / `count` + `targets[]` | surface the count and that review happens at resume/`context review` |
| `commit` | what landed: counts + `minted[]` (`cid` + `logical_id`); extract adds `proposal_id` | translate ("5 use-cases extracted"); `logical_id` slugs are how you reference nodes later |
| `receipt` | per-phase commit receipt (design; extract also emits one at its commits): `commit_n`, `phase`, `committed[]`, `reused`, `rejected[]`, `warnings[]`, `next`, and (design only) `parity` | surface non-empty `rejected[]` and `warnings[]` in plain language; meaningful reuse is worth one line |
| `variables_declared` *(0.8.12+)* | design: stubs written by a phase commit — `file` (pool file) + `variables[]` (`name`, `description`, `secret`); only names that existed in no variable file | surface as a to-do ("N variables need values — in <file>") |
| `variables_summary` *(0.8.12+)* | design, end of run: every declared name still needing a value — same shape | repeat the to-do in the closing summary; fills gate authoring |
| `variables_declared` *(0.8.12+)* | design: the stubs the tests phase wrote, sent just before that phase's `commit` event — `file` (pool file) + `variables[]` (`name`, `description`, `secret`); only names that existed in no variable file | surface as a to-do ("N variables need values — in <file>"), then fill them before any run of these tests (SKILL.md §3) |
| `variables_summary` *(0.8.12+)* | design, after `session_complete` on a clean completion only (a paused or refused run never sends it): every name this run declared, same shape, no re-check of the pool | repeat the to-do in the closing summary; on a pause build it from `variables_declared` alone; unfilled names are typed as written when the test is authored |
| `message_sent` | `--message` delivered: `sid`, `chars` | confirmation only |
| `panel_resolved` *(0.7.1+)* | a `--answer` flag landed on a pending question: `id`, `by`, `via` | confirmation only |
| `ask_deferred` *(0.7.1+)* | `--with-source` set the pending batch aside: `source_id`, `cid`, `questions` (count) | tell the user the questions were deferred while the agent reads the new source |
Expand Down Expand Up @@ -99,4 +99,4 @@ One payload event carrying the full `--json` document — `coverage` for `cover`
| `3` | **paused and resumable** — not a failure; run the pause loop. Includes crash-pauses (0.7.1+). On the sync verbs (0.8.14+) `done{status: "paused"}` means decisions are waiting: after a rebase walk the rebase is open (answer with `--answer`); after `kane-cli context sync doctor --abort` it is closed and the unanswered decisions stayed in the backup — the `sync_rebase_done` before it says which (`status` `paused` or `aborted`); `paused` with no `sync_rebase_done` means doctor set aside a rebase whose saved state it could not read — `sync_doctor.detail` says the next `kane-cli context sync` reapplies the saved backup. `done{status: "refused", exit_code: 3}` means a pull or a rebase is needed first — `references/context-sync.md` §4 |
| `130` | force-interrupted — for extract and design, resumable only if a `session_paused` event arrived; on the sync verbs the rebase state is on disk: `kane-cli context sync doctor --mode agent` shows it, `kane-cli context sync --mode agent` resumes it |

Reminder: this exit-3 meaning is **specific to these assurance commands**. `run` / `testmd` / `testrun` / `generate` keep their own meanings (3 = timeout/cancelled).
Reminder: this exit-3 meaning is **specific to these assurance commands**. `run` / `testmd` / `testrun` keep their own meanings (3 = timeout/cancelled).
13 changes: 6 additions & 7 deletions .agents/skills/kane-cli/references/assurance.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,9 +2,8 @@

# Assurance — Agent Surface

When the user has **requirements** — a PRD, a spec, acceptance notes — and wants tests designed from them, wants to know what's covered, or wants the suite kept current, use the **assurance commands** (`kane-cli context`, `design`, `cover`, `maintain`). Do not hand-write the tests, and do not reach for `generate`:
When the user has **requirements** — a PRD, a spec, acceptance notes — and wants tests designed from them, wants to know what's covered, or wants the suite kept current, use the **assurance commands** (`kane-cli context`, `design`, `cover`, `maintain`). Do not hand-write the tests:

- `kane-cli generate` = quick scenarios/cases from a one-line description. No requirement linkage.
- **Assurance** = tests derived from the actual documents, every claim cited, every test permanently tagged with the acceptance criteria it verifies, coverage measured against requirements. Use it whenever the user cares about "what exactly is covered, and how do we know?"

Everything here works over a local store (`.context/` in the project directory) that the commands create and manage themselves.
Expand All @@ -19,7 +18,7 @@ kane-cli context review --verdicts <file> --json # 2. CHECKPOINT: us
kane-cli design tests --use-case <uc-ref> --mode agent --max 8 # 3. design ACs, scenarios, tests
kane-cli context review --verdicts <file> --json # 4. CHECKPOINT: user approves the design
kane-cli testmd run .testmuai/tests/<t>_test.md --agent # 5. author each kept test once (real browser)
kane-cli testrun run --match 't-' # 6. batch replays from then on
kane-cli testrun run --match 't-' < /dev/null # 6. batch replays from then on
kane-cli cover gaps # 7. designed % × proven % + per-use-case debt
kane-cli maintain reconcile --from <new.md> --source-id <id> --mode agent # when a source changes (§11)
```
Expand All @@ -35,12 +34,12 @@ Extract, design, and reconcile call the KaneAI service and consume credits; ever

## 2. The pause loop — exit 3 is a pause, NOT a failure

**Scope: this rule applies to `context extract`/`context ingest`, `design tests`, and `maintain reconcile` (§11).** (For `run`/`testmd`/`testrun`/`generate`, exit 3 still means timeout/cancelled.)
**Scope: this rule applies to `context extract`/`context ingest`, `design tests`, and `maintain reconcile` (§11).** (For `run`/`testmd`/`testrun`, exit 3 still means timeout/cancelled.)

These commands take **`--mode agent`** — not `--agent`; they reject that flag, and a bare non-TTY invocation exits `2` asking for an explicit mode. In `--mode agent`, **every question pauses the run** (0.8.8+ — earlier CLIs auto-answered low/medium-risk questions with their recommended defaults and paused only on high risk):

- The run exits `3`, emits `session_paused` with the session id, the questions in full (text, options, the recommended one, risk, rationale), and the verbatim resume command.
- **Never drop a pause** (same rule as generate clarifications). Answer it: if your own context clearly resolves the question, answer it yourself; otherwise surface the question — with its options and recommendation — to your user and get their answer.
- **Never drop a pause.** Answer it: if your own context clearly resolves the question, answer it yourself; otherwise surface the question — with its options and recommendation — to your user and get their answer.
- **0.7.1+ sessions are durable from the first turn**: a crash that left a checkpoint exits `3` with a `session_paused` carrying `crashed: true` (no `pending_questions`) and the resume command — exit 3 always means "resumable". A crash before anything durable was saved still exits `1`; check `context sessions --json` before retrying anything paid.

Three ways to resume:
Expand Down Expand Up @@ -119,13 +118,13 @@ kane-cli design tests --use-case <uc-ref> --mode agent --max 8
- There is **no `--because` flag on `design tests`** — interactively the session collects the redesign reason itself; headless `--force` proceeds with an auto-stamped reason.
- `--phase <grounding|acs|scenarios|wiring|tests>` (0.7.1+) re-enters a design at a phase, re-seeded from the committed earlier phases; missing predecessors exit 2 with the commands to run first in `next`.
- Output: acceptance criteria, scenarios, exactly one test per scenario — written as runnable files under `.testmuai/tests/*_test.md`, each assert step tagged with the criteria it verifies. Plus **gaps** (recorded, ranked missing pieces) and **warnings** (e.g. a test claiming more criteria than its check asserts). Citations are verified against the pinned source text before commit (0.7.1+) — a `CITE_UNVERIFIED` error means a citation could not be verified even after repair.
- **Variables (0.8.12+).** Design reads every `*.json` in the user's variable directories (names + has-value only, never values) and reuses a name that exists before inventing one. Each name it invents is declared: an empty stub lands in `.testmuai/variables/assurance.json`, never a value, never over an existing key. The stream tells you which: `variables_declared` at the commit that wrote them, `variables_summary` at the end (schema in `references/assurance-parsing.md`). **Surface these as a to-do** — "2 variables need values: login_email, login_password — in .testmuai/variables/assurance.json" — because a designed test refuses to author until they are filled (SKILL.md §3, Unresolved variables). Fill them yourself only if the user gave you the values.
- **Variables (0.8.12+).** Design reads every `*.json` in the user's variable directories (names + has-value only, never values) and reuses a name that exists before inventing one. Each name it invents is declared: an empty stub lands in `.testmuai/variables/assurance.json`, never a value, never over an existing key. The stream tells you which: `variables_declared` in the tests phase, just before its `commit` event, and `variables_summary` after `session_complete` on a clean completion (a paused run sends only the first; schema in `references/assurance-parsing.md`). **Surface these as a to-do** — "2 variables need values: login_email, login_password — in .testmuai/variables/assurance.json" — because a designed test authored before they are filled types the placeholders as written (SKILL.md §3, Unresolved variables). Before any run of these tests, fill them: SKILL.md §3 **Fill the variables before any run**, asking with each row's `description`.
- **Present tests, gaps, warnings, AND variables needing values** — first-class output, not noise. Then go to the review checkpoint (§4) before any authoring.
- `kane-cli design explain <t-ref>` replays *why* a test exists (technique, boundary values, criteria) with zero AI cost — use it when the user asks "why this test?".

## 6. The authoring bridge — from designed files to batch runs

A freshly designed test has never been executed. On 0.8.4+ hand the set straight to `kane-cli testrun run`: unauthored members classify as `author`, the run authors them in a real browser, and afterwards the authored and replayed evidence consolidates into one published execution — best-effort: when consolidation cannot complete, evidence stays split rather than lost. `--from-context` (0.8.4+) selects members by assurance test ids and follows edit supersessions. Designed tests may carry `{{variables}}` for values the requirements never pinned (a store URL, a product name) — supply them per `references/testmd.md`. `kane-cli testmd run <file> --agent` remains the single-test authoring path, and on pre-0.8.4 CLIs it is REQUIRED first — `testrun` there refuses never-authored members (`missing_meta`). Evidence packs seal per `references/evidence.md`.
A freshly designed test has never been executed. On 0.8.4+ hand the set straight to `kane-cli testrun run`: unauthored members classify as `author`, the run authors them in a real browser, and afterwards the authored and replayed evidence consolidates into one published execution — best-effort: when consolidation cannot complete, evidence stays split rather than lost. `--from-context` (0.8.4+) selects members by assurance test ids and follows edit supersessions. Designed tests may carry `{{variables}}` for values the requirements never pinned (a store URL, a product name) — fill them first (SKILL.md §3), or the run types the placeholder as written. `kane-cli testmd run <file> --agent` remains the single-test authoring path, and on pre-0.8.4 CLIs it is REQUIRED first — `testrun` there refuses never-authored members (`missing_meta`). Evidence packs seal per `references/evidence.md`.

### 6.1 A designed test is the design — do not edit it by hand

Expand Down
3 changes: 2 additions & 1 deletion .agents/skills/kane-cli/references/cards.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@ Every result is an emoji table. A one-line "Test passed" instead of the card is
- **`🟡 Didn't start` is not `🔴 Failed`.** When nothing ran, say what to fix.
- **Secret-looking values never go in chat.** For a missing value whose name contains `password`, `secret`, `token` or `key`, add an empty entry to the variables file for the person to fill. Ask in chat only for plain values (a URL, a user name).
- If the run's output carried an update notice, add one quiet last line under the card: `kane-cli <version> is available.`
- **Variables with no value go on the card.** When the run's output carried the `unresolved_variables` warning, add a `⚠️ **Variables**` row before ➡️ Next naming each one.

## 2. Run, passed

Expand Down Expand Up @@ -72,7 +73,7 @@ Exit code `1`, or `status: "failed"`. Show the failing step's screenshot under t

## 4. Didn't start

Exit code `2`: nothing ran and no credits were used. Causes include missing variable values, no start URL, sign-in or setup errors, a test file that does not parse, an invalid suite plan, a cloud grid refusal.
Exit code `2`: nothing ran and no credits were used. Causes include no start URL, sign-in or setup errors, a test file that does not parse, an invalid suite plan, a cloud grid refusal.

```markdown
| | |
Expand Down
2 changes: 1 addition & 1 deletion .agents/skills/kane-cli/references/debug.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,7 +66,7 @@ For `run` objectives and plain `_test.md` files. For a designed test, read the s
| 🎯 Agent clicks wrong element | Ambiguous UI, multiple similar elements | Be more specific: "click the **blue** 'Submit' button in the **checkout form**" |
| 👁️ Agent says done but didn't finish | Objective too vague | Add explicit assertions: "assert the confirmation page shows order number" |
| 💀 Exit code 2, no steps | Auth, TMS credential exchange, or Chrome failure | Check `kane-cli whoami`, verify Chrome is available |
| ❓ Exit code 2 with "did you mean …" | Bare-objective shortcut — agent ran `kane-cli "<objective>"` without the `run` subcommand | Re-invoke as `kane-cli run "<objective>" --agent` (same rule for `testmd run` / `generate`) |
| ❓ Exit code 2 with "did you mean …" | Bare-objective shortcut — agent ran `kane-cli "<objective>"` without the `run` subcommand | Re-invoke as `kane-cli run "<objective>" --agent` (same rule for `testmd run`) |
| 📤 Upload silently fails after configuring a project/folder by hand | Saved ID is invalid (typo, deleted, no access) | No action needed — the next run detects the 4xx and auto-defaults a working project/folder. To rebind manually: `kane-cli config project` (TTY picker) or `kane-cli projects list` → `kane-cli config project <id>` (see `references/test-manager.md`) |
| ⏱️ Exit code 3 | Timeout or cancelled | Increase `--timeout` or `--max-steps`, or split into smaller objectives |
| 🚫 "CDP endpoint not reachable" | Chrome not running | Let kane-cli manage Chrome (remove `--cdp-endpoint`) |
Expand Down
Loading
Loading