diff --git a/.agents/skills/kane-cli/SKILL.md b/.agents/skills/kane-cli/SKILL.md index 3663572..81c6e03 100644 --- a/.agents/skills/kane-cli/SKILL.md +++ b/.agents/skills/kane-cli/SKILL.md @@ -1,18 +1,18 @@ --- name: kane-cli -description: Browser automation + AI test authoring via kane-cli - run browser objectives, generate & refine test scenarios/cases from a description, design requirement-linked test suites from a PRD/spec (assurance), parse NDJSON output, inspect logs, save runnable _test.md. Use for any task requiring a real browser (navigate, click, fill forms, test web UI, take screenshots), or to author test cases, quick cases from a description via kane-cli generate; a designed, coverage-accounted suite from requirement documents via the assurance commands. Never write test cases by hand. Also runs mobile app tests - native Android app on a virtual emulator or iOS app on a simulator via --target emulator|simulator (desktop browser stays the default target): locally on macOS Apple Silicon, or on the LambdaTest cloud grid from any machine via kane-cli testrun run --remote. Also shares the assurance store (.context/) with a team through a location (a GitHub repository, an S3-compatible bucket, or a folder) via kane-cli context sync, and resolves the decisions a sync rebase asks. +description: Browser automation + AI test authoring via kane-cli - run browser objectives, design requirement-linked test suites from a PRD/spec or from a description (assurance), parse NDJSON output, inspect logs, save runnable _test.md. Use for any task requiring a real browser (navigate, click, fill forms, test web UI, take screenshots), or to author test cases: a designed, coverage-accounted suite from requirement documents, or from a description the user gives you, via the assurance commands. Never write test cases by hand. Also runs mobile app tests - native Android app on a virtual emulator or iOS app on a simulator via --target emulator|simulator (desktop browser stays the default target): locally on macOS Apple Silicon, or on the LambdaTest cloud grid from any machine via kane-cli testrun run --remote. Also shares the assurance store (.context/) with a team through a location (a GitHub repository, an S3-compatible bucket, or a folder) via kane-cli context sync, and resolves the decisions a sync rebase asks. --- # Kane CLI — Browser Automation Skill -Use `kane-cli` for **any task that requires a real browser**: navigating websites, clicking elements, filling forms, searching, testing web UI, taking screenshots, or verifying deployments. Do NOT use Playwright, Puppeteer, or Selenium directly. Use `--agent` for `run`, `testmd run`, and `generate`. `testrun run` has no `--agent`: it emits NDJSON when **stdin** is not a TTY (use `< /dev/null` for terminal automation). Assurance conversational commands use `--mode agent`. +Use `kane-cli` for **any task that requires a real browser**: navigating websites, clicking elements, filling forms, searching, testing web UI, taking screenshots, or verifying deployments. Do NOT use Playwright, Puppeteer, or Selenium directly. Use `--agent` for `run` and `testmd run`. `testrun run` has no `--agent`: it emits NDJSON only when **stdin** is not a TTY, so every `testrun run` line you write ends in `< /dev/null` (bash and zsh: macOS, Linux, Git Bash; in cmd.exe write `< NUL`; from PowerShell run it through cmd: `cmd /c "kane-cli testrun run … < NUL"`). Check the first stdout line: on 0.8.17+ it is `{"type":"stream_start"…}`. A prose plan there instead means the CLI is in terminal mode: it shows the human view and, after a real run, opens an evidence table that waits until `q` or Esc, so the process never exits on its own; stop it and rerun with the redirect. An empty stdout with a message on stderr is a usage error: read it. Assurance conversational commands use `--mode agent`. -**Authoring test cases or scenarios?** Never write them by hand — kane-cli has two authoring pipelines, and the routing matters: +**Authoring test cases or scenarios?** Never write them by hand: every test case comes from the **assurance** commands — Read `references/assurance.md` first. -- The user describes what to test in a sentence or two, or wants quick scenario/case ideas → `kane-cli generate` (§6). -- The user has **requirement documents** (a PRD, a spec, acceptance notes) and wants a designed suite, requirement-linked coverage, or "what exactly is covered?" answers → the **assurance** commands — Read `references/assurance.md` first. +- The user has **requirement documents** (a PRD, a spec, acceptance notes) → ingest them, then design. +- The user only **describes** what to test, in chat → write their description, in their words, to a requirements file, ingest that file, then design (§6). The description is the requirement. -Don't draft test cases in chat or scratch files: both pipelines produce structured, refinable, runnable `_test.md` output. +Don't draft test cases in chat or scratch files: design produces structured, refinable, runnable `_test.md` output. --- @@ -50,7 +50,7 @@ On Windows PowerShell: `$env:KANE_CLI_USER_AGENT=''; kane-cli run **Keeping runs.** When the person's saved purpose is `suite` or `ask`, add `--name ` to every one-off `run`. A named run is recorded as a `_test.md` while it runs, so keeping it afterwards costs nothing, and a run launched without a name cannot be kept without running again. With `one-off`, leave the flag out. Details: `references/first-run.md` §4. -Bash blocks until kane-cli exits, then hands you the complete stdout. Parse it, summarize what happened, and present the result card. Wait for process completion on `testmd run` and `generate` too, but parse their own completion events: `test_md_done` and `generate_done`, respectively. An intermediate `run_end` does not finish a saved test. +Bash blocks until kane-cli exits, then hands you the complete stdout. Parse it, summarize what happened, and present the result card. Wait for process completion on `testmd run` too, but parse its own completion event, `test_md_done`. An intermediate `run_end` does not finish a saved test. Set a generous timeout (up to 600000ms) since browser runs can take a while. @@ -76,7 +76,8 @@ Progress events have `step`/`status`/`remark` fields and **no `type` field**. |------|-------------|-----| | **Failures** | Any step with `status: "failed"` | `Step failed: ` | | **Flow changes** | `bifurcation`, `child_agent_start`, `child_agent_end` | Plain-language one-liner (e.g. "The agent split the objective into 2 sub-tasks") | -| **Errors** | `error` typed events | `Error: `. The exception is `code: "unresolved_variables"`, which is a pre-run refusal, not a failure: see §3 **Unresolved variables** | +| **Errors** | `error` typed events | `Error: ` | +| **Unresolved variables** | `warning` with `code: "unresolved_variables"` | Before any progress: name the variables with no value and where a value goes (§3 **Unresolved variables**). The run continued | | **Overall progress** | All passing steps | One summary line: ` steps completed: <2–4 key actions from remarks>` | #### What to skip @@ -159,9 +160,9 @@ When the user's request involves a browser — or writing test cases: **What does the user want?** - A single one-shot browser task → build a `kane-cli run --agent` command (§3 + §4) - A test they want to save / re-run / commit → Read `references/testmd.md` first, then use `kane-cli testmd` -- Run a suite of saved tests (several `_test.md` at once) → Read `references/testrun.md` first, then use `kane-cli testrun run` -- Need test cases or scenarios from a short description — because the user asked, or because the task needs them (no browser) → **don't hand-write them**; Read `references/generate.md` first, then use `kane-cli generate` (§6) -- Has requirement documents (PRD/spec) and wants a designed suite, coverage accounting, or suite upkeep → Read `references/assurance.md` first — the assurance commands (`context`/`design`/`cover`/`maintain reconcile`, kane-cli 0.6.1+; several features need newer releases, up to 0.8.14+ — the reference marks each), NOT `generate` +- Run a suite of saved tests (several `_test.md` at once) → Read `references/testrun.md` first, then use `kane-cli testrun run … < /dev/null` +- Need test cases from a description the user gives in chat, with no document → **don't hand-write them**; write the description to a requirements file and design from it (§6) +- Has requirement documents (PRD/spec) and wants a designed suite, coverage accounting, or suite upkeep → Read `references/assurance.md` first — the assurance commands (`context`/`design`/`cover`/`maintain reconcile`, kane-cli 0.6.1+; several features need newer releases, up to 0.8.14+ — the reference marks each) - A **designed** test (its `_test.md` carries an `assurance:` block) failed a run → Read `references/assurance.md` §6.1 **before touching the file** — an edit is adopted as the next design version on the next run; classify the failure first (app bug → report it; requirement changed → reconcile; wording → redesign through the CLI; capability missing → stop), and edit a step of an existing designed test by hand only after reading `references/objectives-cookbook.md` - Share the context store with a team, join a teammate's, keep two stores level, or resolve a sync conflict → Read `references/context-sync.md` first — `kane-cli context sync`, `kane-cli context push`, `kane-cli context pull` and `kane-cli context clone` (kane-cli 0.8.14+); the store is shared through a location, never by copying or git-merging `.context/` - Multiple independent browser tasks → Read `references/parallel.md` first @@ -175,7 +176,7 @@ When the user's request involves a browser — or writing test cases: - The person wants to watch runs live, or asks about the status line → Read `references/live-strip.md` (Claude Code only) - You need the full NDJSON event schema (rare — §5's summary covers 90% of cases) → Read `references/parsing.md` - Compare / evaluate / justify kane-cli against another tool or approach (cost, tokens, effort, ROI) → Read `references/fair-evaluation.md` first — comparisons are only honest like-for-like across the test lifecycle -- **Mobile**: drive a native app on a virtual Android emulator or iOS simulator instead of the browser → Read `references/mobile.md` first. Desktop (the browser) stays the **default** target; mobile is opt-in via `--target emulator|simulator` and always drives an app you provide (`--app `), never a URL. **Local** mobile runs (`run`, `testmd run`, `testrun run`) need macOS Apple Silicon. **From any other machine** (Linux, Windows, Intel Mac, a Mac without Xcode/Android Studio), run saved mobile `_test.md` files on the cloud grid with `kane-cli testrun run --remote --device-name "" --os-version ` — the grid boots the emulator/simulator on a HyperExecute macOS host (the account needs a HyperExecute plan with macOS runners). Never tell a non-Mac user mobile is impossible: point them at `--remote`. +- **Mobile**: drive a native app on a virtual Android emulator or iOS simulator instead of the browser → Read `references/mobile.md` first. Desktop (the browser) stays the **default** target; mobile is opt-in via `--target emulator|simulator` and always drives an app you provide (`--app `), never a URL. **Local** mobile runs (`run`, `testmd run`, `testrun run`) need macOS Apple Silicon. **From any other machine** (Linux, Windows, Intel Mac, a Mac without Xcode/Android Studio), run saved mobile `_test.md` files on the cloud grid with `kane-cli testrun run --remote --device-name "" --os-version < /dev/null` — the grid boots the emulator/simulator on a HyperExecute macOS host (the account needs a HyperExecute plan with macOS runners). Never tell a non-Mac user mobile is impossible: point them at `--remote`. **Every run, always:** follow §1 above. @@ -187,7 +188,7 @@ When the user's request involves a browser — or writing test cases: kane-cli run "" --agent [options] ``` -> The `run` subcommand is **mandatory**. `kane-cli ""` (no `run`) does **not** work — unknown first tokens exit `2` with a "did you mean" suggestion. Same rule applies to `kane-cli testmd run …` and `kane-cli generate …`. +> The `run` subcommand is **mandatory**. `kane-cli ""` (no `run`) does **not** work — unknown first tokens exit `2` with a "did you mean" suggestion. Same rule applies to `kane-cli testmd run …`. `--agent` is mandatory — it switches stdout to NDJSON. Most-used flags: @@ -209,7 +210,14 @@ Other flags (`--global-context`, `--local-context`, `--cdp-endpoint`, `--allow-m **Exit codes:** `0` passed · `1` failed · `2` auth/infra error · `3` timeout/cancelled. -**Unresolved variables (0.8.12+):** every `{{name}}` in the objective must have a value before the run starts. If one does not, the run **refuses before anything launches** — exit `2`, and with `--agent` a single `{"type":"error","code":"unresolved_variables", ...}` event carrying `variables[]` (`name`, `reason`: `value_missing` = the key exists in `file` with no value · `not_declared` = the key is in no file, add it to `suggested_file`; `used_by[]`). **This is terminal — do not retry the same command.** Either ask the user for the values, or write them yourself (`--variables '{"name":{"value":"…"}}'`, or `{"name":{"value":""}}` stubs into `suggested_file` for the user to fill), then run again. There is no bypass flag. Never checked: `{{smart.*}}`/`{{environment.*}}`/`{{secrets.*}}`/`{{totp.*}}`, and names an earlier step stores (`store … as 'x'`). Numbers in a variable file count as values (loaded as strings); booleans do not. +**Unresolved variables (0.8.15+):** kane-cli checks every `{{name}}` in the objective before the run starts. A name with no value is a warning: the run goes ahead and types the name as written, unless a step sets it first. With `--agent` the warning is one `{"type":"warning","code":"unresolved_variables", ...}` event before the first progress frame. It carries `suggested_file` and `variables[]`: `name`; `reason` (`value_missing` = the key exists in `file` with no value · `not_declared` = the key is in no file, add it to `suggested_file`); `used_by[]`. Never checked: an explicit `{{global.*}}` (it resolves from Test Manager at run time), `{{smart.*}}`, `{{environment.*}}`, `{{secrets.*}}`, `{{totp.*}}`, and names an earlier step stores (`store … as 'x'`). Numbers in a variable file count as values (loaded as strings); booleans do not. + +**Fill the variables before any run.** Do this before `run`, `testmd run` and `testrun run`, and always before `--remote`, which books a grid job. A missing value is asked for before the run; the first-run rule that nothing is asked before the first result does not cover it. + +1. Collect the names with no value: the `warning`, a dry-run plan's `unresolved[]` rows, or a design run's `variables_declared` and `variables_summary` rows. Design rows list what design declared, not every missing value, and a dry run's `valid: true` says nothing about values: the dry run's `warning` is the check. +2. A value you already have, because the user said it or the requirement document states it, you write into that key in the file the event names (for `run`, `--variables '{"name":{"value":"…"}}'` also works), and you tell the user what you filled and where. +3. The rest you ask for once, in one message, using each variable's `description` when design gave one. Plain values (a URL, an email, a user name) the user gives you here or adds to the file, their choice. Secrets, which are any row with `secret: true`, any description that names a credential, and any name containing `password`, `secret`, `token` or `key`, the user fills in the file and tells you when done; you never ask for the value in chat, never echo it, and report names and file paths only. +4. Then a fresh `--dry-run` of the exact selection: a valid plan with no `warning` is the check. A member whose rows are all filled may run while the others wait. Never fill a placeholder just to silence the check; a throwaway value is right only when the description asks for one, such as a deliberately wrong password. A frontmatter declaration with an empty value also silences the check, so look at the values a step relies on, not only at the warning. A name still empty is typed into the page as written; if a run then failed at that step, say so. ### Examples @@ -298,7 +306,7 @@ Action → extraction → assertion in one objective: Stdout is NDJSON, one event per line. On kane-cli 0.8.17+ every line also carries `v` (contract version, `1`) and `ts` (when it was emitted), and the first line is `{"type":"stream_start","cli_version":…,"surface":"run"|"testmd"|"testrun"}`. Ignore fields and event types you do not know: new ones can appear in any release. There are two shapes: - **Progress events** (most events) have `step` (1-based), `status` (`running` at start, `done`/`failed` at completion), `remark` — and **no `type` field**. -- **Typed events** have a `type` field: `project_folder_auto_defaulted` (run-startup gate, fires before any progress when no project/folder is configured), `bifurcation`, `child_agent_start`, `child_agent_end`, `ask_user`, `error` (an `error` with `code: "unresolved_variables"` is a pre-run refusal and the **only** line — no `run_end` follows; handle per §3), and finally `run_end`. +- **Typed events** have a `type` field: `project_folder_auto_defaulted` (run-startup gate, fires before any progress when no project/folder is configured), `bifurcation`, `child_agent_start`, `child_agent_end`, `ask_user`, `error`, `warning` (pre-run `code: "unresolved_variables"` — the run continues; handle per §3), and finally `run_end`. Parsing strategy: @@ -313,51 +321,21 @@ For one-shot `run`, build post-run logic on `run_end` and process exit. Saved te For full event schemas (`bifurcation` flow fields, `child_agent_*`, `ask_user` semantics, `cancel`/`user_response` outbound events, complete `run_end` field list), Read `references/parsing.md`. -`kane-cli generate` (§6) emits a **different** stream — every line is typed `generate_*` (no untyped progress lines), terminated by `generate_done`. Its schema is in `references/generate-parsing.md`. - The assurance conversational commands (`context ingest`/`context extract`, `design tests`, `maintain reconcile`, `cover`) do NOT take `--agent` — they take **`--mode agent`** and speak their own typed stream ending in `done` (open vocabulary — tolerate unknown event types; on 0.7.2+ the stream is strict — every stdout line parses, stderr silent — while on 0.7.1 a merged `context ingest` prints a few prose receipt lines BEFORE the stream — skip to the first `{` line, harmless on 0.7.2+ — and a landing-phase ingest failure ends with prose + exit 1/2 and no stream at all: a refusal, not a crash; on 0.7.2+ those failures ride the stream as `error` + `done`); **for those commands only, exit `3` means paused-and-resumable, not timeout** — schema in `references/assurance-parsing.md`, behavior in `references/assurance.md`. The context sync verbs (`kane-cli context sync`, `kane-cli context push`, `kane-cli context pull`, `kane-cli context clone`; kane-cli 0.8.14+) take `--mode agent` the same way and speak a `sync_*` family ending in `done`; there too exit `3` means a decision or a pull is needed, not a failure — `references/context-sync.md`. `kane-cli context sync setup` refuses without a terminal (`TTY_REQUIRED`): agents use `kane-cli context sync add` and `kane-cli context clone`. `kane-cli testrun run` also emits its own typed stream (`testrun_plan` … terminal `testrun_done`) — schema in `references/testrun.md`. `kane-cli testmd run` may additionally emit `test_md_evidence_ingest` (replay evidence published) and `test_md_bundle_sync` (test bundle synced) — informational; describe in plain language, never surface raw names. The post-run evidence hint (`` evidence: view locally with `kane-cli evidence serve ` ``) is a **stderr** text line, not a stdout event — don't try to parse it from the NDJSON stream; see `references/evidence.md` for how to act on it. --- -## 6. Generate test cases (authoring — no browser) - -`kane-cli generate` authors **Test Scenarios → Test Cases** from a plain-language description. It does **not** drive a browser. **Use it whenever a task needs quick test cases or scenarios from a description — don't hand-author them in chat or a file.** (Requirement documents + coverage accounting → assurance instead: `references/assurance.md`.) Reach for it to: turn a feature / requirement description into a test suite; expand or refine coverage (more edge cases, negative paths, a narrower focus); or save the Functional cases as runnable `_test.md` and hand them to `kane-cli testmd run`. Full details + event schema: **Read `references/generate.md`**. - -Three explicit modes, each runs **one turn then exits**: - -| Mode | Command | -|---|---| -| **New** | `kane-cli generate "" --agent` | -| **Refine** | `kane-cli generate "" --refine --req --agent` | -| **Save** | `kane-cli generate --save --req --agent` → writes runnable `_test.md` | - -**Launch + present** — same as §1: use `Bash` (not Monitor), emit "Generating test cases…" before launch, then parse the output when it returns. Generate is a **quick single turn** — it exits on its own at `generate_done`. - -**After Bash returns**, parse the NDJSON and present only what matters: - -| Show | Event | How | -|------|-------|-----| -| **The deliverable** | `generate_snapshot` | Present scenarios + cases (see below) | -| **Clarifications** | `generate_clarification` | Surface the question — it needs an answer | -| **Save results** | `generate_save_result` | List files written | -| **Errors** | `error` | Surface the message | -| **Skip everything else** | `generate_thinking`, `generate_progress`, `generate_chat`, `generate_start` | Noise — don't narrate | - -At `generate_done`, **present the result adaptively**: -- **≤ ~30 cases** → a nested tree: each scenario, then its cases tagged Positive / Negative / Edge. -- **more than that** → a summary line + a bulleted scenario list (title + case count); expand a scenario's cases only when asked. - -Then offer the next commands from the terminal line's Refine / Save hints (they carry the request id) — don't hand-build them. - -**Clarification → refine (do not skip):** if the turn ends with a clarification, that's **exit 0 — not an error**. Act on it: answer it yourself, or ask your own user, then **re-invoke** `kane-cli generate "" --refine --req --agent`. Never drop a clarification. +## 6. Test cases from a description (no requirement document) -**Attach files:** `--files a,b,c` adds local files (docs / images / PDF / CSV — up to 10, ≤ 50 MB each) as generation context on a **new** or **`--refine`** turn (not `--save`); each emits a `generate_upload` line before `generate_start`. Details in `references/generate.md`. +When the user describes what to test in chat and has no document, use the description as the requirement. -**Save is Functional-only:** `--save` writes only **Functional** cases to `_test.md` (under `/.testmuai/tests` by default). Non-functional cases (Security, Performance, …) are generated and shown but not saved. Run saved files with **`kane-cli testmd run`** (`references/testmd.md`) — that's the generate → testmd pipeline. +1. Write the user's words to `requirements/.md` in the project, as given, one heading per feature, nothing invented. Tell the user the file exists and that the tests will cite it. +2. `kane-cli context ingest requirements/.md --mode agent`, review the extracted use-cases at the checkpoint, then `kane-cli design tests …`, exactly as `references/assurance.md` describes. +3. Present the designed tests, gaps, warnings and the variables that need values (§3 **Fill the variables before any run**). The designed tests get their own review checkpoint before any run (`references/assurance.md`). Then hand the set to `kane-cli testrun run … < /dev/null` (`references/assurance.md`, the authoring bridge). -Internal event/field names (`generate_snapshot`, `request_id`, …) are for parsing only — never show them to the user (§5 rule). Wire schema: `references/generate-parsing.md`. +A draft in chat cannot be run, and a hand-written `_test.md` carries no requirement link and no coverage accounting; the deliverable is the designed test. --- @@ -368,7 +346,6 @@ Internal event/field names (`generate_snapshot`, `request_id`, …) are for pars | User wants to save/persist/re-run a test | `references/testmd.md` | | Run a suite of saved `_test.md` tests as one batch | `references/testrun.md` | | Run a suite on the cloud grid (`--remote`), incl. mobile suites from any machine | `references/testrun.md` §Remote + `references/mobile.md` §Remote | -| You need quick test cases or scenarios from a description | `references/generate.md` | | User has requirement docs (PRD/spec) → designed suite, coverage, or suite upkeep | `references/assurance.md` | | Need the assurance NDJSON event schema (`--mode agent`) | `references/assurance-parsing.md` | | Share the context store with a team, join one, keep stores level, or resolve a sync conflict (kane-cli 0.8.14+) | `references/context-sync.md` | @@ -376,7 +353,6 @@ Internal event/field names (`generate_snapshot`, `request_id`, …) are for pars | View, share, validate, or merge evidence packs | `references/evidence.md` | | Multiple independent browser tasks | `references/parallel.md` | | Need full NDJSON event schema (`run`) | `references/parsing.md` | -| Need the `generate` NDJSON event schema | `references/generate-parsing.md` | | Browse / create projects or folders, or parse the auto-default event | `references/test-manager.md` | | Start of every session: preflight, the ready card, sign-in | `references/ready-check.md` | | A person's first session: run first, the tour, three choices | `references/first-run.md` | @@ -388,7 +364,7 @@ Internal event/field names (`generate_snapshot`, `request_id`, …) are for pars ## Command-specific completion -The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). `generate` emits `generate_done`. Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. +The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. Progress is for live display: count only `done`/`failed` completions, retaining child and execution context when step indices repeat. diff --git a/.agents/skills/kane-cli/references/assurance-parsing.md b/.agents/skills/kane-cli/references/assurance-parsing.md index a69d44e..9562a82 100644 --- a/.agents/skills/kane-cli/references/assurance-parsing.md +++ b/.agents/skills/kane-cli/references/assurance-parsing.md @@ -41,8 +41,8 @@ A stream that ends **without** `done` means the process crashed — outcome unkn | `held` / `update_held` *(0.7.1+)* | items held for the user's review instead of committed: `source_id` + `count` + `reason` / `count` + `targets[]` | surface the count and that review happens at resume/`context review` | | `commit` | what landed: counts + `minted[]` (`cid` + `logical_id`); extract adds `proposal_id` | translate ("5 use-cases extracted"); `logical_id` slugs are how you reference nodes later | | `receipt` | per-phase commit receipt (design; extract also emits one at its commits): `commit_n`, `phase`, `committed[]`, `reused`, `rejected[]`, `warnings[]`, `next`, and (design only) `parity` | surface non-empty `rejected[]` and `warnings[]` in plain language; meaningful reuse is worth one line | -| `variables_declared` *(0.8.12+)* | design: stubs written by a phase commit — `file` (pool file) + `variables[]` (`name`, `description`, `secret`); only names that existed in no variable file | surface as a to-do ("N variables need values — in ") | -| `variables_summary` *(0.8.12+)* | design, end of run: every declared name still needing a value — same shape | repeat the to-do in the closing summary; fills gate authoring | +| `variables_declared` *(0.8.12+)* | design: the stubs the tests phase wrote, sent just before that phase's `commit` event — `file` (pool file) + `variables[]` (`name`, `description`, `secret`); only names that existed in no variable file | surface as a to-do ("N variables need values — in "), then fill them before any run of these tests (SKILL.md §3) | +| `variables_summary` *(0.8.12+)* | design, after `session_complete` on a clean completion only (a paused or refused run never sends it): every name this run declared, same shape, no re-check of the pool | repeat the to-do in the closing summary; on a pause build it from `variables_declared` alone; unfilled names are typed as written when the test is authored | | `message_sent` | `--message` delivered: `sid`, `chars` | confirmation only | | `panel_resolved` *(0.7.1+)* | a `--answer` flag landed on a pending question: `id`, `by`, `via` | confirmation only | | `ask_deferred` *(0.7.1+)* | `--with-source` set the pending batch aside: `source_id`, `cid`, `questions` (count) | tell the user the questions were deferred while the agent reads the new source | @@ -99,4 +99,4 @@ One payload event carrying the full `--json` document — `coverage` for `cover` | `3` | **paused and resumable** — not a failure; run the pause loop. Includes crash-pauses (0.7.1+). On the sync verbs (0.8.14+) `done{status: "paused"}` means decisions are waiting: after a rebase walk the rebase is open (answer with `--answer`); after `kane-cli context sync doctor --abort` it is closed and the unanswered decisions stayed in the backup — the `sync_rebase_done` before it says which (`status` `paused` or `aborted`); `paused` with no `sync_rebase_done` means doctor set aside a rebase whose saved state it could not read — `sync_doctor.detail` says the next `kane-cli context sync` reapplies the saved backup. `done{status: "refused", exit_code: 3}` means a pull or a rebase is needed first — `references/context-sync.md` §4 | | `130` | force-interrupted — for extract and design, resumable only if a `session_paused` event arrived; on the sync verbs the rebase state is on disk: `kane-cli context sync doctor --mode agent` shows it, `kane-cli context sync --mode agent` resumes it | -Reminder: this exit-3 meaning is **specific to these assurance commands**. `run` / `testmd` / `testrun` / `generate` keep their own meanings (3 = timeout/cancelled). +Reminder: this exit-3 meaning is **specific to these assurance commands**. `run` / `testmd` / `testrun` keep their own meanings (3 = timeout/cancelled). diff --git a/.agents/skills/kane-cli/references/assurance.md b/.agents/skills/kane-cli/references/assurance.md index cab6a28..05b188a 100644 --- a/.agents/skills/kane-cli/references/assurance.md +++ b/.agents/skills/kane-cli/references/assurance.md @@ -2,9 +2,8 @@ # Assurance — Agent Surface -When the user has **requirements** — a PRD, a spec, acceptance notes — and wants tests designed from them, wants to know what's covered, or wants the suite kept current, use the **assurance commands** (`kane-cli context`, `design`, `cover`, `maintain`). Do not hand-write the tests, and do not reach for `generate`: +When the user has **requirements** — a PRD, a spec, acceptance notes — and wants tests designed from them, wants to know what's covered, or wants the suite kept current, use the **assurance commands** (`kane-cli context`, `design`, `cover`, `maintain`). Do not hand-write the tests: -- `kane-cli generate` = quick scenarios/cases from a one-line description. No requirement linkage. - **Assurance** = tests derived from the actual documents, every claim cited, every test permanently tagged with the acceptance criteria it verifies, coverage measured against requirements. Use it whenever the user cares about "what exactly is covered, and how do we know?" Everything here works over a local store (`.context/` in the project directory) that the commands create and manage themselves. @@ -19,7 +18,7 @@ kane-cli context review --verdicts --json # 2. CHECKPOINT: us kane-cli design tests --use-case --mode agent --max 8 # 3. design ACs, scenarios, tests kane-cli context review --verdicts --json # 4. CHECKPOINT: user approves the design kane-cli testmd run .testmuai/tests/_test.md --agent # 5. author each kept test once (real browser) -kane-cli testrun run --match 't-' # 6. batch replays from then on +kane-cli testrun run --match 't-' < /dev/null # 6. batch replays from then on kane-cli cover gaps # 7. designed % × proven % + per-use-case debt kane-cli maintain reconcile --from --source-id --mode agent # when a source changes (§11) ``` @@ -35,12 +34,12 @@ Extract, design, and reconcile call the KaneAI service and consume credits; ever ## 2. The pause loop — exit 3 is a pause, NOT a failure -**Scope: this rule applies to `context extract`/`context ingest`, `design tests`, and `maintain reconcile` (§11).** (For `run`/`testmd`/`testrun`/`generate`, exit 3 still means timeout/cancelled.) +**Scope: this rule applies to `context extract`/`context ingest`, `design tests`, and `maintain reconcile` (§11).** (For `run`/`testmd`/`testrun`, exit 3 still means timeout/cancelled.) These commands take **`--mode agent`** — not `--agent`; they reject that flag, and a bare non-TTY invocation exits `2` asking for an explicit mode. In `--mode agent`, **every question pauses the run** (0.8.8+ — earlier CLIs auto-answered low/medium-risk questions with their recommended defaults and paused only on high risk): - The run exits `3`, emits `session_paused` with the session id, the questions in full (text, options, the recommended one, risk, rationale), and the verbatim resume command. -- **Never drop a pause** (same rule as generate clarifications). Answer it: if your own context clearly resolves the question, answer it yourself; otherwise surface the question — with its options and recommendation — to your user and get their answer. +- **Never drop a pause.** Answer it: if your own context clearly resolves the question, answer it yourself; otherwise surface the question — with its options and recommendation — to your user and get their answer. - **0.7.1+ sessions are durable from the first turn**: a crash that left a checkpoint exits `3` with a `session_paused` carrying `crashed: true` (no `pending_questions`) and the resume command — exit 3 always means "resumable". A crash before anything durable was saved still exits `1`; check `context sessions --json` before retrying anything paid. Three ways to resume: @@ -119,13 +118,13 @@ kane-cli design tests --use-case --mode agent --max 8 - There is **no `--because` flag on `design tests`** — interactively the session collects the redesign reason itself; headless `--force` proceeds with an auto-stamped reason. - `--phase ` (0.7.1+) re-enters a design at a phase, re-seeded from the committed earlier phases; missing predecessors exit 2 with the commands to run first in `next`. - Output: acceptance criteria, scenarios, exactly one test per scenario — written as runnable files under `.testmuai/tests/*_test.md`, each assert step tagged with the criteria it verifies. Plus **gaps** (recorded, ranked missing pieces) and **warnings** (e.g. a test claiming more criteria than its check asserts). Citations are verified against the pinned source text before commit (0.7.1+) — a `CITE_UNVERIFIED` error means a citation could not be verified even after repair. -- **Variables (0.8.12+).** Design reads every `*.json` in the user's variable directories (names + has-value only, never values) and reuses a name that exists before inventing one. Each name it invents is declared: an empty stub lands in `.testmuai/variables/assurance.json`, never a value, never over an existing key. The stream tells you which: `variables_declared` at the commit that wrote them, `variables_summary` at the end (schema in `references/assurance-parsing.md`). **Surface these as a to-do** — "2 variables need values: login_email, login_password — in .testmuai/variables/assurance.json" — because a designed test refuses to author until they are filled (SKILL.md §3, Unresolved variables). Fill them yourself only if the user gave you the values. +- **Variables (0.8.12+).** Design reads every `*.json` in the user's variable directories (names + has-value only, never values) and reuses a name that exists before inventing one. Each name it invents is declared: an empty stub lands in `.testmuai/variables/assurance.json`, never a value, never over an existing key. The stream tells you which: `variables_declared` in the tests phase, just before its `commit` event, and `variables_summary` after `session_complete` on a clean completion (a paused run sends only the first; schema in `references/assurance-parsing.md`). **Surface these as a to-do** — "2 variables need values: login_email, login_password — in .testmuai/variables/assurance.json" — because a designed test authored before they are filled types the placeholders as written (SKILL.md §3, Unresolved variables). Before any run of these tests, fill them: SKILL.md §3 **Fill the variables before any run**, asking with each row's `description`. - **Present tests, gaps, warnings, AND variables needing values** — first-class output, not noise. Then go to the review checkpoint (§4) before any authoring. - `kane-cli design explain ` replays *why* a test exists (technique, boundary values, criteria) with zero AI cost — use it when the user asks "why this test?". ## 6. The authoring bridge — from designed files to batch runs -A freshly designed test has never been executed. On 0.8.4+ hand the set straight to `kane-cli testrun run`: unauthored members classify as `author`, the run authors them in a real browser, and afterwards the authored and replayed evidence consolidates into one published execution — best-effort: when consolidation cannot complete, evidence stays split rather than lost. `--from-context` (0.8.4+) selects members by assurance test ids and follows edit supersessions. Designed tests may carry `{{variables}}` for values the requirements never pinned (a store URL, a product name) — supply them per `references/testmd.md`. `kane-cli testmd run --agent` remains the single-test authoring path, and on pre-0.8.4 CLIs it is REQUIRED first — `testrun` there refuses never-authored members (`missing_meta`). Evidence packs seal per `references/evidence.md`. +A freshly designed test has never been executed. On 0.8.4+ hand the set straight to `kane-cli testrun run`: unauthored members classify as `author`, the run authors them in a real browser, and afterwards the authored and replayed evidence consolidates into one published execution — best-effort: when consolidation cannot complete, evidence stays split rather than lost. `--from-context` (0.8.4+) selects members by assurance test ids and follows edit supersessions. Designed tests may carry `{{variables}}` for values the requirements never pinned (a store URL, a product name) — fill them first (SKILL.md §3), or the run types the placeholder as written. `kane-cli testmd run --agent` remains the single-test authoring path, and on pre-0.8.4 CLIs it is REQUIRED first — `testrun` there refuses never-authored members (`missing_meta`). Evidence packs seal per `references/evidence.md`. ### 6.1 A designed test is the design — do not edit it by hand diff --git a/.agents/skills/kane-cli/references/cards.md b/.agents/skills/kane-cli/references/cards.md index 9904459..b035268 100644 --- a/.agents/skills/kane-cli/references/cards.md +++ b/.agents/skills/kane-cli/references/cards.md @@ -16,6 +16,7 @@ Every result is an emoji table. A one-line "Test passed" instead of the card is - **`🟡 Didn't start` is not `🔴 Failed`.** When nothing ran, say what to fix. - **Secret-looking values never go in chat.** For a missing value whose name contains `password`, `secret`, `token` or `key`, add an empty entry to the variables file for the person to fill. Ask in chat only for plain values (a URL, a user name). - If the run's output carried an update notice, add one quiet last line under the card: `kane-cli is available.` +- **Variables with no value go on the card.** When the run's output carried the `unresolved_variables` warning, add a `⚠️ **Variables**` row before ➡️ Next naming each one. ## 2. Run, passed @@ -72,7 +73,7 @@ Exit code `1`, or `status: "failed"`. Show the failing step's screenshot under t ## 4. Didn't start -Exit code `2`: nothing ran and no credits were used. Causes include missing variable values, no start URL, sign-in or setup errors, a test file that does not parse, an invalid suite plan, a cloud grid refusal. +Exit code `2`: nothing ran and no credits were used. Causes include no start URL, sign-in or setup errors, a test file that does not parse, an invalid suite plan, a cloud grid refusal. ```markdown | | | diff --git a/.agents/skills/kane-cli/references/debug.md b/.agents/skills/kane-cli/references/debug.md index 7d8faf0..372d188 100644 --- a/.agents/skills/kane-cli/references/debug.md +++ b/.agents/skills/kane-cli/references/debug.md @@ -66,7 +66,7 @@ For `run` objectives and plain `_test.md` files. For a designed test, read the s | 🎯 Agent clicks wrong element | Ambiguous UI, multiple similar elements | Be more specific: "click the **blue** 'Submit' button in the **checkout form**" | | 👁️ Agent says done but didn't finish | Objective too vague | Add explicit assertions: "assert the confirmation page shows order number" | | 💀 Exit code 2, no steps | Auth, TMS credential exchange, or Chrome failure | Check `kane-cli whoami`, verify Chrome is available | -| ❓ Exit code 2 with "did you mean …" | Bare-objective shortcut — agent ran `kane-cli ""` without the `run` subcommand | Re-invoke as `kane-cli run "" --agent` (same rule for `testmd run` / `generate`) | +| ❓ Exit code 2 with "did you mean …" | Bare-objective shortcut — agent ran `kane-cli ""` without the `run` subcommand | Re-invoke as `kane-cli run "" --agent` (same rule for `testmd run`) | | 📤 Upload silently fails after configuring a project/folder by hand | Saved ID is invalid (typo, deleted, no access) | No action needed — the next run detects the 4xx and auto-defaults a working project/folder. To rebind manually: `kane-cli config project` (TTY picker) or `kane-cli projects list` → `kane-cli config project ` (see `references/test-manager.md`) | | ⏱️ Exit code 3 | Timeout or cancelled | Increase `--timeout` or `--max-steps`, or split into smaller objectives | | 🚫 "CDP endpoint not reachable" | Chrome not running | Let kane-cli manage Chrome (remove `--cdp-endpoint`) | diff --git a/.agents/skills/kane-cli/references/generate-parsing.md b/.agents/skills/kane-cli/references/generate-parsing.md deleted file mode 100644 index d767ebe..0000000 --- a/.agents/skills/kane-cli/references/generate-parsing.md +++ /dev/null @@ -1,65 +0,0 @@ - - -# Reading `generate --agent` Output - -> **Internal reference only.** The event types and field names below are for you to parse programmatically. **Never expose them to the user** — present plain-language scenarios/cases per `references/generate.md`, never `generate_snapshot`, `request_id`, `generate_done`, or raw JSON. - -With `--agent`, `kane-cli generate` writes **one JSON object per line** to **stdout**. Unlike `run` (which has untyped `step`/`status` progress lines), **every generate line is typed** — it always has a `type` field. That makes parsing simpler: - -``` -for each line of NDJSON: - parse JSON, switch on obj.type - if obj.type === "generate_done" → terminal event, stop parsing - if obj.type === "generate_snapshot" → the deliverable (full scenarios + cases) - if obj.type === "generate_clarification" → turn ended awaiting an answer (still exit 0) - else → progress / informational -``` - -Build post-turn logic on **`generate_done`** (terminal, stable schema) and read the result from **`generate_snapshot`** (emitted once, at turn end). - -## Event types - -Generate-specific (typed `generate_*`): - -| `type` | Key fields | Meaning → what to do | -|---|---|---| -| `generate_upload` | `file`, `index`, `total`, `status` | Only when `--files` is used. Emitted once per attached file **before `generate_start`**, while the file uploads; `status` goes `"uploading"` → `"done"` (or `"failed"`). Progress only — narrate "attaching …" or ignore. A failed upload surfaces as an `error` + non-zero exit. | -| `generate_start` | `request_id`, `objective_chars`, `scenario_limit`, `per_scenario_limit`, `is_refine` | Turn began. **Capture `request_id`** — it's the handle for every later `--refine` / `--save`. `is_refine` distinguishes a new request from a continuation. | -| `generate_thinking` | `took_ms` | Liveness only. Narrate "thinking…" or ignore. | -| `generate_progress` | `pct` | Milestone (25 / 50 / 75 / 100). Optional progress display — not a completion signal. | -| `generate_snapshot` | `scenario_count`, `case_count`, `scenarios[]` | **The deliverable** — full scenarios, each with its cases. Each case carries `title`, `polarity` (`"p"`/`"n"`/`"e"` = Positive / Negative / Edge), `category` (`"Functional"`, `"Security"`, …), `priority`. Present it per `generate.md`. Emitted exactly once. | -| `generate_clarification` | `text` | The generator needs an answer; the turn ended (exit 0) awaiting it. **Not an error** — answer via `--refine --req` (see `generate.md`). | -| `generate_chat` | `text` | The model's prose reply (e.g. what a refine changed). Show as info. | -| `generate_save_result` | `suite_dir`, `saved`, `fell_back`, `warning?` | `--save` wrote files. `saved` = files written; `fell_back` = cases written as prose because they couldn't be expanded; `warning` e.g. `"no functional test cases"`. | -| `generate_done` | `request_id`, `status`, `scenario_count`, `case_count`, `refine_hint`, `save_hint`, `suite_dir?` | **Terminal line.** Branch on `status`; use `refine_hint` / `save_hint` **verbatim** as the next command. `suite_dir` present only after a `--save`. | - -Shared with `run` (NOT generate-specific — don't treat as part of the generate contract): - -| `type` | Fields | Meaning | -|---|---|---| -| `project_folder_auto_defaulted` | resolved project + folder (id, name) | Run-startup gate auto-resolved a project/folder when none was configured (or the cached one was stale/invalid). Fires before `generate_start`. Translate to plain language (see `references/test-manager.md`). | -| `error` | `message` | A failure occurred; pair with the exit code to decide retry vs abort. | -| `update_available` | `current`, `latest`, `severity` | A newer kane-cli exists. Informational; emitted before `generate_start`. | - -## Terminal `generate_done` - -Example after a **`--save`** run (note `suite_dir` — a new/refine turn omits it): - -```json -{ - "type": "generate_done", - "request_id": "23271", - "status": "completed", - "scenario_count": 3, - "case_count": 11, - "refine_hint": "kane-cli generate \"\" --refine --req 23271", - "save_hint": "kane-cli generate --save --req 23271", - "suite_dir": ".testmuai/tests/checkout-23271" -} -``` - -`status`: `completed` (exit 0 — including a turn that ended with a clarification) · `failed` or `ended` (exit 1) · `stopped` (exit 3). `suite_dir` is present only after a `--save`. The hints carry no `--agent` flag, but it is auto-on for non-TTY callers (agents/pipes), so re-invoking a hint verbatim still yields NDJSON. - -## Exit codes - -`0` completed · `1` failed · `2` error (auth / setup / transport, or an invalid flag combination) · `3` stopped / cancelled · `130` interrupted (Ctrl-C). Interactive `--refine`/`--save` continuity is by re-invocation with `--req` — there is no stdin relay. diff --git a/.agents/skills/kane-cli/references/generate.md b/.agents/skills/kane-cli/references/generate.md deleted file mode 100644 index d947c1e..0000000 --- a/.agents/skills/kane-cli/references/generate.md +++ /dev/null @@ -1,139 +0,0 @@ - - -# Generating Test Cases with `kane-cli generate` - -`kane-cli generate` turns a plain-language description of *what to test* into structured **Test Scenarios** (logical groupings) each containing **Test Cases** (typed Positive / Negative / Edge). It calls the AI Test Case Generator — **no browser is launched**. The result is a tree of scenarios + cases you present to the user, refine conversationally, and optionally save as runnable `_test.md` files. - -**Use this whenever a task needs test cases or scenarios written — don't hand-author them in chat or a scratch file.** Reach for it to: turn a requirement / feature description into a test suite; expand or refine coverage (more edge cases, negative paths, a narrower or broader focus); or save the Functional cases as runnable `_test.md`. It is **not** for driving a browser — that's `kane-cli run` (§3 of SKILL.md). - -> For the full web-product picture (dashboards, issue-link inputs), see the public docs: . The CLI takes a **text objective**, optionally with local files attached via **`--files`** (see "Attaching files" below). - -## The three modes — one turn per invocation, then exit - -There is **no interactive session**. Each invocation runs exactly one generation turn and exits. Continuity across turns is carried by a **request id** (`--req `) that the previous turn's terminal line hands back. - -| Mode | Command | Notes | -|---|---|---| -| **New** | `kane-cli generate "" --agent` | Starts a fresh request. Capture the request id from the terminal line. | -| **Refine** | `kane-cli generate "" --refine --req --agent` | Adjusts an existing request. `--refine` **and** `--req` required; needs a change description. | -| **Save** | `kane-cli generate --save --req [--out ] --agent` | Writes the request's Functional cases to `_test.md`. No new turn, takes no objective. | - -`--refine` and `--save` always run headless (even from a terminal). - -### Flags - -| Flag | Purpose | -|---|---| -| `--agent` | Typed NDJSON on stdout (auto-on when stdin is not a TTY) | -| `--req ` | The request id to `--refine` or `--save` | -| `--out ` | Save target — **only** with `--save`; default `/.testmuai/tests` | -| `--name ` | Names the run and the saved suite folder | -| `--scenario-limit ` / `--per-scenario-limit ` | Cap scenarios / cases-per-scenario | -| `--memory` | Use the memory layer — reuse relevant existing cases, reduce duplicates | -| `--files ` | Comma-separated local files to attach as context (new / refine only — see "Attaching files") | -| `--project ` / `--folder ` | Test Manager project / folder | -| `--username` / `--access-key` | Auth (same as `run`) | - -If neither `--project`/`--folder` nor a saved project/folder is set when generation starts, kane-cli auto-resolves one headlessly and emits a `project_folder_auto_defaulted` event before `generate_start`. Translate it to a one-line note for the user — full handling lives in `references/test-manager.md`. - -## Attaching files - -Pass local files as extra context with `--files ` on a **new** generation or a **`--refine`** (not `--save`). The generator reads them and reflects them in the scenarios + cases — attach a spec, a screenshot of the UI, a PDF / Word doc, or a CSV of inputs. - -```bash -kane-cli generate "test the login flow described in the attached spec" --files ./login-spec.pdf,./wireframe.png --agent -``` - -- **Supported types** — documents (`.txt .json .xml .csv .pdf .docx .xlsx`), images (`.jpg .jpeg .png .gif .bmp .webp`), audio (`.mp3 .wav .m4a`), video (`.mp4 .mov .webm .mpeg .mpga`). -- **Limits** — up to **10 files**, each **≤ 50 MB**. -- **Validated as a set, up front** — if any path is missing, an unsupported type, too large, or over the count, the whole command is rejected (exit `2`) **before anything is sent** and the offending paths are listed; fix and re-run. Files outside the current directory are allowed but flagged with a warning on stderr. -- **`new` / `--refine` only** — combining `--files` with `--save` exits `2`. - -Under `--agent`, each file emits a `generate_upload` line (`status` `uploading` → `done`) **before** `generate_start` — see `references/generate-parsing.md`. *(Interactively in the TUI, type `@` in the generate prompt to attach a file inline.)* - -## Presenting a result (adaptive) - -The terminal data carries the full scenarios + cases. **Present it based on size**: - -- **≤ ~30 cases → a nested tree** (scenario, then each case with its type tag): - ``` - ✓ Generated 3 scenarios · 11 cases (request 23271) - - ▸ Login - - Valid credentials [Positive] - - Wrong password [Negative] - - Empty fields [Edge] - ▸ Checkout - - Guest checkout [Positive] - - Expired card [Negative] - ... - ``` -- **more than ~30 cases → a summary + scenario list** (cases on request): - ``` - ✓ Generated 6 scenarios · 84 cases (request 23271) - - • Login (12 cases) - • Checkout (20 cases) - • Cart management (14 cases) - ... - ``` - -Always end with the **next-step commands the terminal line provides** (Refine / Save) — they already carry the request id, so don't hand-build them. - -## Clarifications — act on them, never drop them - -If a turn ends with a **clarification question**, that is **success (exit 0)**, not an error — the generator needs an answer before it can continue. You must act on it: - -1. Read the question. -2. **Decide** — answer it yourself from context, **or** surface it to your user and get an answer. -3. **Re-invoke** with the answer as a refine: - ```bash - kane-cli generate "" --refine --req --agent - ``` - -## The refine → save → run loop - -```bash -# 1. New request -kane-cli generate "checkout flow on a shopping site" --agent -# → terminal line carries request id 23271 + Refine/Save hints - -# 2. Refine (repeat as needed) -kane-cli generate "also cover an expired card and an out-of-stock item" --refine --req 23271 --agent - -# 3. Save the Functional cases as runnable _test.md -kane-cli generate --save --req 23271 --agent -# → /.testmuai/tests///_test.md - -# 4. Run / replay them -kane-cli testmd run .testmuai/tests///_test.md --agent -``` - -**Save is Functional-only.** `--save` writes only test cases whose category is **Functional** — those are the ones runnable as `_test.md`. Non-functional cases (Security, Performance, etc.) are generated and shown in the result but are **not** written; saving a request with no Functional cases writes nothing and says so. Saved files are ordinary `_test.md` tests — see `references/testmd.md` for running, editing, and replay. This is the **generate → testmd** pipeline: author cases here, run them there. - -## Exit codes - -| Code | Meaning | -|---|---| -| 0 | Turn completed (including a turn that ended with a clarification) | -| 1 | Generation failed (or ended) | -| 2 | Error — auth / setup / transport, or an invalid flag combination | -| 3 | Generation stopped / cancelled | -| 130 | Interrupted (Ctrl-C) | - -Invalid flag combinations exit `2` with a message on stderr. The full set: - -- `--refine` and `--save` together -- `--refine` without `--req` -- `--refine` without a change description -- `--refine` combined with `--out` (`--out` is save-only) -- `--save` without `--req` -- `--save` with a description (it takes none) -- `--out` without `--save` -- `--files` with `--save` (files attach to a new generation or a refine, not a save) -- `--req` without `--refine` or `--save` -- a new generation with no description - -## Reading the output - -`--agent` emits one typed JSON object per line. For the full event schema and the parse strategy, Read **`references/generate-parsing.md`**. As with all `--agent` output, the field names are for parsing only — **never show them to the user**; present plain-language scenarios and cases (per "Presenting a result" above). diff --git a/.agents/skills/kane-cli/references/mobile.md b/.agents/skills/kane-cli/references/mobile.md index 222a8c4..eca231e 100644 --- a/.agents/skills/kane-cli/references/mobile.md +++ b/.agents/skills/kane-cli/references/mobile.md @@ -9,10 +9,10 @@ Desktop (the browser) is the **default** target and the primary use of kane-cli. | Where the device runs | Host | Commands | |---|---|---| | **Local** (a simulator/emulator on this machine) | **macOS on Apple Silicon (arm64) only** — not Intel Macs, Linux, or Windows | `run --target …`, `testmd run`, `testrun run` | -| **Cloud grid** (a virtual device on a HyperExecute macOS host) | **Any machine** — no Xcode / Android Studio needed. The account needs a HyperExecute plan with macOS runners | `testrun run --remote …` — see §Remote below | +| **Cloud grid** (a virtual device on a HyperExecute macOS host) | **Any machine** — no Xcode / Android Studio needed. The account needs a HyperExecute plan with macOS runners | `testrun run --remote … < /dev/null` — see §Remote below | - **Desktop stays the default.** The `--target` axis is what selects mobile. Leave it off and you get the browser. -- If the user is not on mac-arm64, local mobile is not an option — **offer the grid**: save the objective as a `_test.md` (`target: emulator|simulator` + `app:`) and run it with `kane-cli testrun run --remote --device-name "" --os-version `. Do not tell them mobile is unavailable. +- If the user is not on mac-arm64, local mobile is not an option — **offer the grid**: save the objective as a `_test.md` (`target: emulator|simulator` + `app:`) and run it with `kane-cli testrun run --remote --device-name "" --os-version < /dev/null`. Do not tell them mobile is unavailable. ## The three targets @@ -111,7 +111,7 @@ Everything else about `_test.md` (step bodies, replay/cascade, commands) is unch `kane-cli testrun run` accepts mobile `_test.md` members (0.8.7+): -- **Locally** (mac-arm64 with the setup above): `kane-cli testrun run --device-name "" --os-version ` — the device as `kane-cli devices list --target ` prints it, or the members' own `device_name:`/`os_version:`. +- **Locally** (mac-arm64 with the setup above): `kane-cli testrun run --device-name "" --os-version < /dev/null` — the device as `kane-cli devices list --target ` prints it, or the members' own `device_name:`/`os_version:`. - **On the cloud grid, from any machine**: add `--remote` — next section. Full testrun flags, events, and rollup: `references/testrun.md`. ## Remote: mobile suites on the cloud grid (`testrun run --remote`) @@ -121,9 +121,9 @@ One command turns a mobile suite into a HyperExecute job on a macOS host that bo ```bash kane-cli plugin install remote-execution # once; then `kane-cli plugin doctor remote-execution` kane-cli devices list --target emulator --remote --agent # grid catalog: name + os_versions per row -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 --dry-run # validate + resolve device, no job -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 -kane-cli testrun run tests/ios/ --remote --device-name "iPhone 15" --os-version 17.5 +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 --dry-run < /dev/null # validate + resolve device, no job +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 < /dev/null +kane-cli testrun run --match '^tests/ios/' --remote --device-name "iPhone 15" --os-version 17.5 < /dev/null ``` **Always `--dry-run` first** — it runs the remote preflight and resolves the device against the catalog at no cost. Use `Bash` with a long timeout (up to 600000 ms) for the real run: device setup + app install + members take several minutes. diff --git a/.agents/skills/kane-cli/references/parallel.md b/.agents/skills/kane-cli/references/parallel.md index 5bc51bb..0866044 100644 --- a/.agents/skills/kane-cli/references/parallel.md +++ b/.agents/skills/kane-cli/references/parallel.md @@ -4,7 +4,7 @@ For multiple independent browser tasks, decompose and run in parallel using the Agent tool. -> **Saved tests? Use testrun instead.** If the tasks are committed `_test.md` files, do NOT hand-roll parallelism — `kane-cli testrun run --parallel N` gives you isolated Chromes, a pooled scheduler, one rollup, and one evidence pack. Read `references/testrun.md`. This reference is for **ad-hoc `run` objectives** only. +> **Saved tests? Use testrun instead.** If the tasks are committed `_test.md` files, do NOT hand-roll parallelism — `kane-cli testrun run --parallel N < /dev/null` gives you isolated Chromes, a pooled scheduler, one rollup, and one evidence pack. Read `references/testrun.md`. This reference is for **ad-hoc `run` objectives** only. ## When to Split diff --git a/.agents/skills/kane-cli/references/parsing.md b/.agents/skills/kane-cli/references/parsing.md index 2e6aac9..ec3952e 100644 --- a/.agents/skills/kane-cli/references/parsing.md +++ b/.agents/skills/kane-cli/references/parsing.md @@ -55,24 +55,25 @@ These are **untyped** — they have no `type` field. Do **not** key on `event.ty | Event (`type` field) | Key Fields | Purpose | |-------|-----------|---------| -| `project_folder_auto_defaulted` | resolved project + folder (id, name) | Run-startup gate auto-resolved a project/folder when none was configured (or the cached one was stale/invalid). Fires before any progress event on `run` / `testmd run` / `generate`. Translate to plain language (see `references/test-manager.md`). | +| `project_folder_auto_defaulted` | resolved project + folder (id, name) | Run-startup gate auto-resolved a project/folder when none was configured (or the cached one was stale/invalid). Fires before any progress event on `run` / `testmd run`. Translate to plain language (see `references/test-manager.md`). | | `bifurcation` | `flows[]`, `count` | Agent split objective into sub-flows | | `child_agent_start` | `child_id`, `objective`, `parent_step` | Child agent spawned | | `child_agent_end` | `child_id`, `success`, `steps_taken`, `summary` | Child agent finished | | `ask_user` | `question`, `step_index`, `options?` | Agent needs user input | -| `error` | `message`, `code?` | Error occurred. With `code: "unresolved_variables"` it is the pre-run refusal — the only event of the run, no `run_end` follows; schema below. | +| `error` | `message` | Error occurred | +| `warning` | `code`, `message`, … | *(0.8.15+)* Before any progress event: `code: "unresolved_variables"` — a `{{name}}` had no value; the run continues. Schema below. | | `test_md_evidence_ingest` | `status: "ok"\|"failed"`, `evidence_id`, `stage?` (failure only) | `testmd run` only: a replay's evidence pack published to the dashboard. Informational. | | `test_md_bundle_sync` | `status: "ok"\|"failed"`, `commit_id`, `bytes?` (success) / `stage?` (failure) | `testmd run` / `testmd sync`: test bundle pushed to the cloud after an authored commit. Informational. | | `testrun_*` family | see `references/testrun.md` | Emitted only by `kane-cli testrun run`; terminal event is `testrun_done`, not `run_end`. | **Note:** The `run` stream has no `run_start` event; startup metadata or errors can precede progress. -### `error` with `code: "unresolved_variables"` (0.8.12+) +### `warning` with `code: "unresolved_variables"` (0.8.15+) -Emitted by `run`, `testmd run` and (after `testrun_plan`) `testrun run` when an authored step references a `{{name}}` that has no value. Nothing was dispatched; exit code `2`; stderr is silent. +Emitted by `run`, `testmd run` and (after `testrun_plan`) `testrun run` when an authored step references a `{{name}}` that has no value. The run goes ahead: the name is typed as written unless a step sets it first. The warning does not change the exit code. A run given a Test Manager dataset row can emit a second `warning` with the same code for `${x}` names that are no column of that row, still before any progress. ```json -{"type":"error","code":"unresolved_variables","message":"2 variable(s) have no value — nothing was dispatched", +{"type":"warning","code":"unresolved_variables","message":"2 variable(s) have no value — typed as written unless a step sets them first", "suggested_file":".testmuai/variables/variables.json", "variables":[ {"name":"checkout_url","reason":"not_declared","used_by":[{"file":"objective","step":1}]}, @@ -82,14 +83,12 @@ Emitted by `run`, `testmd run` and (after `testrun_plan`) `testrun run` when an | Field | Meaning | |---|---| -| `variables[].reason` | `value_missing` — the key exists in `variables[].file` with an empty value · `not_declared` — the key is in no variable file | +| `variables[].reason` | `value_missing` — the key exists in `variables[].file` with an empty value · `not_declared` — the key is in no variable file · `not_a_dataset_column` — a `${x}` that is no column of the Test Manager dataset the run was given (`variables[].dataset` names it) | | `variables[].file` | the pool file that holds the empty key (`value_missing` only); `--variables` when it came inline | | `variables[].used_by[]` | `{file, step}` — `file` is `objective` for `kane-cli run`, else the test file (flattened step index) | | `suggested_file` | where to add a `not_declared` key: `.testmuai/variables/assurance.json` inside an assurance store, `variables.json` otherwise | -Terminal: do not re-run the same command. Supply values (`--variables`, or fill the file) and run again. - -**Note:** `ask_user` is auto-disabled when stdin is not a TTY. Since agents typically run kane-cli as a subprocess, ask_user events will not be emitted. Write objectives that don't require interactive input. +Report it with the run's outcome. If a later step failed on a literal placeholder, point at this. ## Parsing Strategy for one-shot `run` @@ -159,6 +158,6 @@ To cancel a run: ## Command-specific completion -The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). `generate` emits `generate_done`. Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. +The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. Progress is for live display: count only `done`/`failed` completions, retaining child and execution context when step indices repeat. diff --git a/.agents/skills/kane-cli/references/setup-and-config.md b/.agents/skills/kane-cli/references/setup-and-config.md index b876526..ace5e83 100644 --- a/.agents/skills/kane-cli/references/setup-and-config.md +++ b/.agents/skills/kane-cli/references/setup-and-config.md @@ -82,7 +82,7 @@ Variables parameterize objectives with reusable values and secrets. Use `{{key}} **Values (0.8.12+):** a string is a value when non-empty; a number is accepted and loaded as its string; a boolean, object or array is not a value. An empty `value` is a declared-but-unfilled key. -**Before a run (0.8.12+):** every `{{name}}` an authored step references must have a value — `run`, `testmd run` and `testrun run` refuse before launching anything otherwise (exit `2`; with `--agent`, one `error` event with `code: "unresolved_variables"` — SKILL.md §3). No bypass flag. +**Before a run (0.8.15+):** `run`, `testmd run` and `testrun run` check every `{{name}}` an authored step references before launching anything. A name with no value is a warning (with `--agent`, one `warning` event with `code: "unresolved_variables"`; SKILL.md §3): the run goes ahead and types the name as written unless a step sets it first. **`assurance.json`:** `kane-cli design tests` writes empty stubs (`{"name":{"value":"","secret":false,"description":"…"}}`) into `{cwd}/.testmuai/variables/assurance.json` for every variable it declares, never a value, and never over a key that already exists in any pool file. Fill those before authoring the designed tests. diff --git a/.agents/skills/kane-cli/references/test-manager.md b/.agents/skills/kane-cli/references/test-manager.md index 91653cf..65e9030 100644 --- a/.agents/skills/kane-cli/references/test-manager.md +++ b/.agents/skills/kane-cli/references/test-manager.md @@ -104,7 +104,7 @@ To use the result for subsequent runs, persist with `kane-cli config project ` | Filter candidates by project-relative path regex | — | +| `--match ` | Filter candidates by project-relative path regex | The path is as the OS writes it: `tests/app/` on macOS and Linux, `tests\app\` on Windows. Quote the regex with double quotes in cmd.exe; single quotes are literal there. | | `--tags ` | ANY-match on frontmatter `tags:` (repeatable or comma-separated, case-insensitive) | — | | `--parallel ` | Worker count; each desktop worker gets an isolated Chrome with a fresh temp profile | `1` | | `--on-failure ` | `continue` (run everything) \| `fail-fast` (stop dispatching new members after a failure) | `continue` | @@ -45,16 +45,16 @@ All members must share one org + project. *(0.8.4+)* Members need **not** be aut | `org_mismatch` | Different organisation than the other tests | "Check `kane-cli testmd status ` — it belongs to another org" | | `project_mismatch` | Different project than the other tests | "Run it separately or per-project" | -If **any** member fails preflight, the plan is invalid: nothing runs, exit `2`. Suggest `--dry-run` to preview the plan cheaply before a big run. *(0.8.12+)* Preflight also checks variables: a member whose authored steps reference a `{{name}}` with no value fails with `unresolved_variables` — `testrun run` has no `--variables` flag, so fill the pool file (`.testmuai/variables/*.json`) or the member's own `variables:` frontmatter. +If **any** member fails preflight, the plan is invalid: nothing runs, exit `2`. Suggest `--dry-run` to preview the plan cheaply before a big run. *(0.8.15+)* Preflight also checks variables. A `{{name}}` with no value is a warning: the member stays in the plan (`valid` does not look at variables), and the run types the name as written unless a step sets it first. The check covers every step of the member, replayed ones included. `testrun run` has no `--variables` flag, so fill the pool file (`.testmuai/variables/*.json`) or the member's own `variables:` frontmatter first. ## Remote: the suite as one HyperExecute job (`--remote`) -`kane-cli testrun run --remote` ships the cwd as the job payload, provisions a grid runtime on a **HyperExecute macOS runner** — Chrome for web members, a **virtual Android emulator or iOS simulator** for mobile members — runs every member there as its own headless `testmd run`, and brings the recordings (`output-/`) and the sealed evidence pack back into the project. It works **from any machine** with nothing local but Node and the plugin (no Chrome needed); the account needs a HyperExecute plan with macOS runners and the plugin (`kane-cli plugin install remote-execution`; check with `kane-cli plugin doctor remote-execution`). Auth is a LambdaTest username + access key — an OAuth profile is exchanged automatically. Not the same as `--ws-endpoint`, which attaches a remote browser to a run that still executes locally. +`kane-cli testrun run --remote < /dev/null` ships the cwd as the job payload, provisions a grid runtime on a **HyperExecute macOS runner** — Chrome for web members, a **virtual Android emulator or iOS simulator** for mobile members — runs every member there as its own headless `testmd run`, and brings the recordings (`output-/`) and the sealed evidence pack back into the project. It works **from any machine** with nothing local but Node and the plugin (no Chrome needed); the account needs a HyperExecute plan with macOS runners and the plugin (`kane-cli plugin install remote-execution`; check with `kane-cli plugin doctor remote-execution`). Auth is a LambdaTest username + access key — an OAuth profile is exchanged automatically. Not the same as `--ws-endpoint`, which attaches a remote browser to a run that still executes locally. ```bash -kane-cli testrun run --tags smoke --remote --dry-run # web suite: validate, dispatch nothing -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 # Android suite on the grid -kane-cli testrun run tests/ios/ --remote --device-name "iPhone 15" --os-version 17.5 # iOS suite on the grid +kane-cli testrun run --tags smoke --remote --dry-run < /dev/null # web suite: validate, dispatch nothing +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 < /dev/null # Android suite on the grid +kane-cli testrun run --match '^tests/ios/' --remote --device-name "iPhone 15" --os-version 17.5 < /dev/null # iOS suite on the grid ``` - **Always `--dry-run` first.** It runs the normal preflight plus the **remote preflight** and resolves the device against the grid catalog (`kane-cli devices list --target emulator|simulator --remote --agent`) without creating a job. @@ -87,7 +87,7 @@ All typed; stdout; one JSON object per line. **Local completion: `testrun_done`. | `type` | Payload | Notes | |---|---|---| -| `testrun_plan` | `members: [{path, test_id?, tags, failure?}]`, `valid`, `parallel`, `parallel_clamped?` | If `valid: false`, treat as immediate failure — report each member's `failure` reason and stop expecting more events. *(0.8.12+)* `failure: "unresolved_variables"` means a member references a `{{name}}` with no value; one `error` event with `code: "unresolved_variables"` follows the plan (schema in `references/parsing.md`) and lists every such name across members — surface it, do not retry. | +| `testrun_plan` | `members: [{path, test_id?, tags, failure?, unresolved?}]`, `valid`, `parallel`, `parallel_clamped?` | If `valid: false`, treat as immediate failure — report each member's `failure` reason; a `warning` may still follow before the process exits. *(0.8.15+)* `unresolved[]` lists the member's `{{name}}`s with no value; `valid` does not look at it. One `warning` event with `code: "unresolved_variables"` follows the plan (schema in `references/parsing.md`) and the run goes ahead. | | `testrun_start` | `execution_id`, `members` (paths), `parallel` | | | `testrun_member_start` | `path`, `test_id?`, *(0.8.17+)* `session_id`, `log_path` | A saved test started. `log_path` is the absolute path of that test's own event log (see **Each test's own log** below). | | `testrun_member_end` | `path`, `test_id?`, `status`, `duration_s`, *(0.8.17+)* `session_id`, `log_path`, `failure?: {message, step_index?}` | `status` ∈ `passed \| failed \| broken \| interrupted`. `failure` is present when the test did not pass: use it for the "where" and "why" of the failed-tests table. | @@ -129,6 +129,7 @@ for each line: if type === "testrun_done" → capture suite outcome; remote runs keep reading if type === "remote_done" → capture remote status, exit and sessions_path if type === "testrun_plan" && !valid → report offenders, expect exit 2 + if type === "warning" && code === "unresolved_variables" → name the variables with no value (references/parsing.md); the run continues; ignore other warning codes if type === "testrun_member_end" → note per-member outcome if type === "testrun_summary" → capture totals for the rollup else → informational; narrate sparingly @@ -164,7 +165,7 @@ Local suites containing any mobile member require `--parallel 1`; larger values Healing is enabled by default (three shrinking replay windows, then re-authoring of authorable steps). `--no-adaptive-heal` disables it. Retired `--retry`/`--retry-count` only print a notice and have no effect. Replay-only recorded steps retain their recordings even during healing. -NDJSON selection uses stdin, not stdout: run `kane-cli testrun run < /dev/null` for automation launched from a terminal. Dry-run validates a plan, not runtime authentication or browser/device readiness. Always observe process exit, including paths without a normal completion event. +NDJSON selection uses stdin, not stdout: every `kane-cli testrun run` line ends in `< /dev/null` (bash and zsh on macOS, Linux and Git Bash; `< NUL` in cmd.exe; from PowerShell run it through cmd: `cmd /c "kane-cli testrun run … < NUL"`). Check the first stdout line: on 0.8.17+ it is `{"type":"stream_start"…}`. A prose plan there instead means the CLI is in terminal mode: it prints the human view and, after a real run, opens an evidence table that waits until `q` or Esc, so the process never exits on its own. Stop it and rerun with the redirect. An empty stdout with a message on stderr is a usage error: read it. Dry-run validates a plan, not runtime authentication or browser/device readiness. Always observe process exit, including paths without a normal completion event. ### Remote behavior still requiring verification diff --git a/.claude/skills/kane-cli/SKILL.md b/.claude/skills/kane-cli/SKILL.md index 3663572..81c6e03 100644 --- a/.claude/skills/kane-cli/SKILL.md +++ b/.claude/skills/kane-cli/SKILL.md @@ -1,18 +1,18 @@ --- name: kane-cli -description: Browser automation + AI test authoring via kane-cli - run browser objectives, generate & refine test scenarios/cases from a description, design requirement-linked test suites from a PRD/spec (assurance), parse NDJSON output, inspect logs, save runnable _test.md. Use for any task requiring a real browser (navigate, click, fill forms, test web UI, take screenshots), or to author test cases, quick cases from a description via kane-cli generate; a designed, coverage-accounted suite from requirement documents via the assurance commands. Never write test cases by hand. Also runs mobile app tests - native Android app on a virtual emulator or iOS app on a simulator via --target emulator|simulator (desktop browser stays the default target): locally on macOS Apple Silicon, or on the LambdaTest cloud grid from any machine via kane-cli testrun run --remote. Also shares the assurance store (.context/) with a team through a location (a GitHub repository, an S3-compatible bucket, or a folder) via kane-cli context sync, and resolves the decisions a sync rebase asks. +description: Browser automation + AI test authoring via kane-cli - run browser objectives, design requirement-linked test suites from a PRD/spec or from a description (assurance), parse NDJSON output, inspect logs, save runnable _test.md. Use for any task requiring a real browser (navigate, click, fill forms, test web UI, take screenshots), or to author test cases: a designed, coverage-accounted suite from requirement documents, or from a description the user gives you, via the assurance commands. Never write test cases by hand. Also runs mobile app tests - native Android app on a virtual emulator or iOS app on a simulator via --target emulator|simulator (desktop browser stays the default target): locally on macOS Apple Silicon, or on the LambdaTest cloud grid from any machine via kane-cli testrun run --remote. Also shares the assurance store (.context/) with a team through a location (a GitHub repository, an S3-compatible bucket, or a folder) via kane-cli context sync, and resolves the decisions a sync rebase asks. --- # Kane CLI — Browser Automation Skill -Use `kane-cli` for **any task that requires a real browser**: navigating websites, clicking elements, filling forms, searching, testing web UI, taking screenshots, or verifying deployments. Do NOT use Playwright, Puppeteer, or Selenium directly. Use `--agent` for `run`, `testmd run`, and `generate`. `testrun run` has no `--agent`: it emits NDJSON when **stdin** is not a TTY (use `< /dev/null` for terminal automation). Assurance conversational commands use `--mode agent`. +Use `kane-cli` for **any task that requires a real browser**: navigating websites, clicking elements, filling forms, searching, testing web UI, taking screenshots, or verifying deployments. Do NOT use Playwright, Puppeteer, or Selenium directly. Use `--agent` for `run` and `testmd run`. `testrun run` has no `--agent`: it emits NDJSON only when **stdin** is not a TTY, so every `testrun run` line you write ends in `< /dev/null` (bash and zsh: macOS, Linux, Git Bash; in cmd.exe write `< NUL`; from PowerShell run it through cmd: `cmd /c "kane-cli testrun run … < NUL"`). Check the first stdout line: on 0.8.17+ it is `{"type":"stream_start"…}`. A prose plan there instead means the CLI is in terminal mode: it shows the human view and, after a real run, opens an evidence table that waits until `q` or Esc, so the process never exits on its own; stop it and rerun with the redirect. An empty stdout with a message on stderr is a usage error: read it. Assurance conversational commands use `--mode agent`. -**Authoring test cases or scenarios?** Never write them by hand — kane-cli has two authoring pipelines, and the routing matters: +**Authoring test cases or scenarios?** Never write them by hand: every test case comes from the **assurance** commands — Read `references/assurance.md` first. -- The user describes what to test in a sentence or two, or wants quick scenario/case ideas → `kane-cli generate` (§6). -- The user has **requirement documents** (a PRD, a spec, acceptance notes) and wants a designed suite, requirement-linked coverage, or "what exactly is covered?" answers → the **assurance** commands — Read `references/assurance.md` first. +- The user has **requirement documents** (a PRD, a spec, acceptance notes) → ingest them, then design. +- The user only **describes** what to test, in chat → write their description, in their words, to a requirements file, ingest that file, then design (§6). The description is the requirement. -Don't draft test cases in chat or scratch files: both pipelines produce structured, refinable, runnable `_test.md` output. +Don't draft test cases in chat or scratch files: design produces structured, refinable, runnable `_test.md` output. --- @@ -50,7 +50,7 @@ On Windows PowerShell: `$env:KANE_CLI_USER_AGENT=''; kane-cli run **Keeping runs.** When the person's saved purpose is `suite` or `ask`, add `--name ` to every one-off `run`. A named run is recorded as a `_test.md` while it runs, so keeping it afterwards costs nothing, and a run launched without a name cannot be kept without running again. With `one-off`, leave the flag out. Details: `references/first-run.md` §4. -Bash blocks until kane-cli exits, then hands you the complete stdout. Parse it, summarize what happened, and present the result card. Wait for process completion on `testmd run` and `generate` too, but parse their own completion events: `test_md_done` and `generate_done`, respectively. An intermediate `run_end` does not finish a saved test. +Bash blocks until kane-cli exits, then hands you the complete stdout. Parse it, summarize what happened, and present the result card. Wait for process completion on `testmd run` too, but parse its own completion event, `test_md_done`. An intermediate `run_end` does not finish a saved test. Set a generous timeout (up to 600000ms) since browser runs can take a while. @@ -76,7 +76,8 @@ Progress events have `step`/`status`/`remark` fields and **no `type` field**. |------|-------------|-----| | **Failures** | Any step with `status: "failed"` | `Step failed: ` | | **Flow changes** | `bifurcation`, `child_agent_start`, `child_agent_end` | Plain-language one-liner (e.g. "The agent split the objective into 2 sub-tasks") | -| **Errors** | `error` typed events | `Error: `. The exception is `code: "unresolved_variables"`, which is a pre-run refusal, not a failure: see §3 **Unresolved variables** | +| **Errors** | `error` typed events | `Error: ` | +| **Unresolved variables** | `warning` with `code: "unresolved_variables"` | Before any progress: name the variables with no value and where a value goes (§3 **Unresolved variables**). The run continued | | **Overall progress** | All passing steps | One summary line: ` steps completed: <2–4 key actions from remarks>` | #### What to skip @@ -159,9 +160,9 @@ When the user's request involves a browser — or writing test cases: **What does the user want?** - A single one-shot browser task → build a `kane-cli run --agent` command (§3 + §4) - A test they want to save / re-run / commit → Read `references/testmd.md` first, then use `kane-cli testmd` -- Run a suite of saved tests (several `_test.md` at once) → Read `references/testrun.md` first, then use `kane-cli testrun run` -- Need test cases or scenarios from a short description — because the user asked, or because the task needs them (no browser) → **don't hand-write them**; Read `references/generate.md` first, then use `kane-cli generate` (§6) -- Has requirement documents (PRD/spec) and wants a designed suite, coverage accounting, or suite upkeep → Read `references/assurance.md` first — the assurance commands (`context`/`design`/`cover`/`maintain reconcile`, kane-cli 0.6.1+; several features need newer releases, up to 0.8.14+ — the reference marks each), NOT `generate` +- Run a suite of saved tests (several `_test.md` at once) → Read `references/testrun.md` first, then use `kane-cli testrun run … < /dev/null` +- Need test cases from a description the user gives in chat, with no document → **don't hand-write them**; write the description to a requirements file and design from it (§6) +- Has requirement documents (PRD/spec) and wants a designed suite, coverage accounting, or suite upkeep → Read `references/assurance.md` first — the assurance commands (`context`/`design`/`cover`/`maintain reconcile`, kane-cli 0.6.1+; several features need newer releases, up to 0.8.14+ — the reference marks each) - A **designed** test (its `_test.md` carries an `assurance:` block) failed a run → Read `references/assurance.md` §6.1 **before touching the file** — an edit is adopted as the next design version on the next run; classify the failure first (app bug → report it; requirement changed → reconcile; wording → redesign through the CLI; capability missing → stop), and edit a step of an existing designed test by hand only after reading `references/objectives-cookbook.md` - Share the context store with a team, join a teammate's, keep two stores level, or resolve a sync conflict → Read `references/context-sync.md` first — `kane-cli context sync`, `kane-cli context push`, `kane-cli context pull` and `kane-cli context clone` (kane-cli 0.8.14+); the store is shared through a location, never by copying or git-merging `.context/` - Multiple independent browser tasks → Read `references/parallel.md` first @@ -175,7 +176,7 @@ When the user's request involves a browser — or writing test cases: - The person wants to watch runs live, or asks about the status line → Read `references/live-strip.md` (Claude Code only) - You need the full NDJSON event schema (rare — §5's summary covers 90% of cases) → Read `references/parsing.md` - Compare / evaluate / justify kane-cli against another tool or approach (cost, tokens, effort, ROI) → Read `references/fair-evaluation.md` first — comparisons are only honest like-for-like across the test lifecycle -- **Mobile**: drive a native app on a virtual Android emulator or iOS simulator instead of the browser → Read `references/mobile.md` first. Desktop (the browser) stays the **default** target; mobile is opt-in via `--target emulator|simulator` and always drives an app you provide (`--app `), never a URL. **Local** mobile runs (`run`, `testmd run`, `testrun run`) need macOS Apple Silicon. **From any other machine** (Linux, Windows, Intel Mac, a Mac without Xcode/Android Studio), run saved mobile `_test.md` files on the cloud grid with `kane-cli testrun run --remote --device-name "" --os-version ` — the grid boots the emulator/simulator on a HyperExecute macOS host (the account needs a HyperExecute plan with macOS runners). Never tell a non-Mac user mobile is impossible: point them at `--remote`. +- **Mobile**: drive a native app on a virtual Android emulator or iOS simulator instead of the browser → Read `references/mobile.md` first. Desktop (the browser) stays the **default** target; mobile is opt-in via `--target emulator|simulator` and always drives an app you provide (`--app `), never a URL. **Local** mobile runs (`run`, `testmd run`, `testrun run`) need macOS Apple Silicon. **From any other machine** (Linux, Windows, Intel Mac, a Mac without Xcode/Android Studio), run saved mobile `_test.md` files on the cloud grid with `kane-cli testrun run --remote --device-name "" --os-version < /dev/null` — the grid boots the emulator/simulator on a HyperExecute macOS host (the account needs a HyperExecute plan with macOS runners). Never tell a non-Mac user mobile is impossible: point them at `--remote`. **Every run, always:** follow §1 above. @@ -187,7 +188,7 @@ When the user's request involves a browser — or writing test cases: kane-cli run "" --agent [options] ``` -> The `run` subcommand is **mandatory**. `kane-cli ""` (no `run`) does **not** work — unknown first tokens exit `2` with a "did you mean" suggestion. Same rule applies to `kane-cli testmd run …` and `kane-cli generate …`. +> The `run` subcommand is **mandatory**. `kane-cli ""` (no `run`) does **not** work — unknown first tokens exit `2` with a "did you mean" suggestion. Same rule applies to `kane-cli testmd run …`. `--agent` is mandatory — it switches stdout to NDJSON. Most-used flags: @@ -209,7 +210,14 @@ Other flags (`--global-context`, `--local-context`, `--cdp-endpoint`, `--allow-m **Exit codes:** `0` passed · `1` failed · `2` auth/infra error · `3` timeout/cancelled. -**Unresolved variables (0.8.12+):** every `{{name}}` in the objective must have a value before the run starts. If one does not, the run **refuses before anything launches** — exit `2`, and with `--agent` a single `{"type":"error","code":"unresolved_variables", ...}` event carrying `variables[]` (`name`, `reason`: `value_missing` = the key exists in `file` with no value · `not_declared` = the key is in no file, add it to `suggested_file`; `used_by[]`). **This is terminal — do not retry the same command.** Either ask the user for the values, or write them yourself (`--variables '{"name":{"value":"…"}}'`, or `{"name":{"value":""}}` stubs into `suggested_file` for the user to fill), then run again. There is no bypass flag. Never checked: `{{smart.*}}`/`{{environment.*}}`/`{{secrets.*}}`/`{{totp.*}}`, and names an earlier step stores (`store … as 'x'`). Numbers in a variable file count as values (loaded as strings); booleans do not. +**Unresolved variables (0.8.15+):** kane-cli checks every `{{name}}` in the objective before the run starts. A name with no value is a warning: the run goes ahead and types the name as written, unless a step sets it first. With `--agent` the warning is one `{"type":"warning","code":"unresolved_variables", ...}` event before the first progress frame. It carries `suggested_file` and `variables[]`: `name`; `reason` (`value_missing` = the key exists in `file` with no value · `not_declared` = the key is in no file, add it to `suggested_file`); `used_by[]`. Never checked: an explicit `{{global.*}}` (it resolves from Test Manager at run time), `{{smart.*}}`, `{{environment.*}}`, `{{secrets.*}}`, `{{totp.*}}`, and names an earlier step stores (`store … as 'x'`). Numbers in a variable file count as values (loaded as strings); booleans do not. + +**Fill the variables before any run.** Do this before `run`, `testmd run` and `testrun run`, and always before `--remote`, which books a grid job. A missing value is asked for before the run; the first-run rule that nothing is asked before the first result does not cover it. + +1. Collect the names with no value: the `warning`, a dry-run plan's `unresolved[]` rows, or a design run's `variables_declared` and `variables_summary` rows. Design rows list what design declared, not every missing value, and a dry run's `valid: true` says nothing about values: the dry run's `warning` is the check. +2. A value you already have, because the user said it or the requirement document states it, you write into that key in the file the event names (for `run`, `--variables '{"name":{"value":"…"}}'` also works), and you tell the user what you filled and where. +3. The rest you ask for once, in one message, using each variable's `description` when design gave one. Plain values (a URL, an email, a user name) the user gives you here or adds to the file, their choice. Secrets, which are any row with `secret: true`, any description that names a credential, and any name containing `password`, `secret`, `token` or `key`, the user fills in the file and tells you when done; you never ask for the value in chat, never echo it, and report names and file paths only. +4. Then a fresh `--dry-run` of the exact selection: a valid plan with no `warning` is the check. A member whose rows are all filled may run while the others wait. Never fill a placeholder just to silence the check; a throwaway value is right only when the description asks for one, such as a deliberately wrong password. A frontmatter declaration with an empty value also silences the check, so look at the values a step relies on, not only at the warning. A name still empty is typed into the page as written; if a run then failed at that step, say so. ### Examples @@ -298,7 +306,7 @@ Action → extraction → assertion in one objective: Stdout is NDJSON, one event per line. On kane-cli 0.8.17+ every line also carries `v` (contract version, `1`) and `ts` (when it was emitted), and the first line is `{"type":"stream_start","cli_version":…,"surface":"run"|"testmd"|"testrun"}`. Ignore fields and event types you do not know: new ones can appear in any release. There are two shapes: - **Progress events** (most events) have `step` (1-based), `status` (`running` at start, `done`/`failed` at completion), `remark` — and **no `type` field**. -- **Typed events** have a `type` field: `project_folder_auto_defaulted` (run-startup gate, fires before any progress when no project/folder is configured), `bifurcation`, `child_agent_start`, `child_agent_end`, `ask_user`, `error` (an `error` with `code: "unresolved_variables"` is a pre-run refusal and the **only** line — no `run_end` follows; handle per §3), and finally `run_end`. +- **Typed events** have a `type` field: `project_folder_auto_defaulted` (run-startup gate, fires before any progress when no project/folder is configured), `bifurcation`, `child_agent_start`, `child_agent_end`, `ask_user`, `error`, `warning` (pre-run `code: "unresolved_variables"` — the run continues; handle per §3), and finally `run_end`. Parsing strategy: @@ -313,51 +321,21 @@ For one-shot `run`, build post-run logic on `run_end` and process exit. Saved te For full event schemas (`bifurcation` flow fields, `child_agent_*`, `ask_user` semantics, `cancel`/`user_response` outbound events, complete `run_end` field list), Read `references/parsing.md`. -`kane-cli generate` (§6) emits a **different** stream — every line is typed `generate_*` (no untyped progress lines), terminated by `generate_done`. Its schema is in `references/generate-parsing.md`. - The assurance conversational commands (`context ingest`/`context extract`, `design tests`, `maintain reconcile`, `cover`) do NOT take `--agent` — they take **`--mode agent`** and speak their own typed stream ending in `done` (open vocabulary — tolerate unknown event types; on 0.7.2+ the stream is strict — every stdout line parses, stderr silent — while on 0.7.1 a merged `context ingest` prints a few prose receipt lines BEFORE the stream — skip to the first `{` line, harmless on 0.7.2+ — and a landing-phase ingest failure ends with prose + exit 1/2 and no stream at all: a refusal, not a crash; on 0.7.2+ those failures ride the stream as `error` + `done`); **for those commands only, exit `3` means paused-and-resumable, not timeout** — schema in `references/assurance-parsing.md`, behavior in `references/assurance.md`. The context sync verbs (`kane-cli context sync`, `kane-cli context push`, `kane-cli context pull`, `kane-cli context clone`; kane-cli 0.8.14+) take `--mode agent` the same way and speak a `sync_*` family ending in `done`; there too exit `3` means a decision or a pull is needed, not a failure — `references/context-sync.md`. `kane-cli context sync setup` refuses without a terminal (`TTY_REQUIRED`): agents use `kane-cli context sync add` and `kane-cli context clone`. `kane-cli testrun run` also emits its own typed stream (`testrun_plan` … terminal `testrun_done`) — schema in `references/testrun.md`. `kane-cli testmd run` may additionally emit `test_md_evidence_ingest` (replay evidence published) and `test_md_bundle_sync` (test bundle synced) — informational; describe in plain language, never surface raw names. The post-run evidence hint (`` evidence: view locally with `kane-cli evidence serve ` ``) is a **stderr** text line, not a stdout event — don't try to parse it from the NDJSON stream; see `references/evidence.md` for how to act on it. --- -## 6. Generate test cases (authoring — no browser) - -`kane-cli generate` authors **Test Scenarios → Test Cases** from a plain-language description. It does **not** drive a browser. **Use it whenever a task needs quick test cases or scenarios from a description — don't hand-author them in chat or a file.** (Requirement documents + coverage accounting → assurance instead: `references/assurance.md`.) Reach for it to: turn a feature / requirement description into a test suite; expand or refine coverage (more edge cases, negative paths, a narrower focus); or save the Functional cases as runnable `_test.md` and hand them to `kane-cli testmd run`. Full details + event schema: **Read `references/generate.md`**. - -Three explicit modes, each runs **one turn then exits**: - -| Mode | Command | -|---|---| -| **New** | `kane-cli generate "" --agent` | -| **Refine** | `kane-cli generate "" --refine --req --agent` | -| **Save** | `kane-cli generate --save --req --agent` → writes runnable `_test.md` | - -**Launch + present** — same as §1: use `Bash` (not Monitor), emit "Generating test cases…" before launch, then parse the output when it returns. Generate is a **quick single turn** — it exits on its own at `generate_done`. - -**After Bash returns**, parse the NDJSON and present only what matters: - -| Show | Event | How | -|------|-------|-----| -| **The deliverable** | `generate_snapshot` | Present scenarios + cases (see below) | -| **Clarifications** | `generate_clarification` | Surface the question — it needs an answer | -| **Save results** | `generate_save_result` | List files written | -| **Errors** | `error` | Surface the message | -| **Skip everything else** | `generate_thinking`, `generate_progress`, `generate_chat`, `generate_start` | Noise — don't narrate | - -At `generate_done`, **present the result adaptively**: -- **≤ ~30 cases** → a nested tree: each scenario, then its cases tagged Positive / Negative / Edge. -- **more than that** → a summary line + a bulleted scenario list (title + case count); expand a scenario's cases only when asked. - -Then offer the next commands from the terminal line's Refine / Save hints (they carry the request id) — don't hand-build them. - -**Clarification → refine (do not skip):** if the turn ends with a clarification, that's **exit 0 — not an error**. Act on it: answer it yourself, or ask your own user, then **re-invoke** `kane-cli generate "" --refine --req --agent`. Never drop a clarification. +## 6. Test cases from a description (no requirement document) -**Attach files:** `--files a,b,c` adds local files (docs / images / PDF / CSV — up to 10, ≤ 50 MB each) as generation context on a **new** or **`--refine`** turn (not `--save`); each emits a `generate_upload` line before `generate_start`. Details in `references/generate.md`. +When the user describes what to test in chat and has no document, use the description as the requirement. -**Save is Functional-only:** `--save` writes only **Functional** cases to `_test.md` (under `/.testmuai/tests` by default). Non-functional cases (Security, Performance, …) are generated and shown but not saved. Run saved files with **`kane-cli testmd run`** (`references/testmd.md`) — that's the generate → testmd pipeline. +1. Write the user's words to `requirements/.md` in the project, as given, one heading per feature, nothing invented. Tell the user the file exists and that the tests will cite it. +2. `kane-cli context ingest requirements/.md --mode agent`, review the extracted use-cases at the checkpoint, then `kane-cli design tests …`, exactly as `references/assurance.md` describes. +3. Present the designed tests, gaps, warnings and the variables that need values (§3 **Fill the variables before any run**). The designed tests get their own review checkpoint before any run (`references/assurance.md`). Then hand the set to `kane-cli testrun run … < /dev/null` (`references/assurance.md`, the authoring bridge). -Internal event/field names (`generate_snapshot`, `request_id`, …) are for parsing only — never show them to the user (§5 rule). Wire schema: `references/generate-parsing.md`. +A draft in chat cannot be run, and a hand-written `_test.md` carries no requirement link and no coverage accounting; the deliverable is the designed test. --- @@ -368,7 +346,6 @@ Internal event/field names (`generate_snapshot`, `request_id`, …) are for pars | User wants to save/persist/re-run a test | `references/testmd.md` | | Run a suite of saved `_test.md` tests as one batch | `references/testrun.md` | | Run a suite on the cloud grid (`--remote`), incl. mobile suites from any machine | `references/testrun.md` §Remote + `references/mobile.md` §Remote | -| You need quick test cases or scenarios from a description | `references/generate.md` | | User has requirement docs (PRD/spec) → designed suite, coverage, or suite upkeep | `references/assurance.md` | | Need the assurance NDJSON event schema (`--mode agent`) | `references/assurance-parsing.md` | | Share the context store with a team, join one, keep stores level, or resolve a sync conflict (kane-cli 0.8.14+) | `references/context-sync.md` | @@ -376,7 +353,6 @@ Internal event/field names (`generate_snapshot`, `request_id`, …) are for pars | View, share, validate, or merge evidence packs | `references/evidence.md` | | Multiple independent browser tasks | `references/parallel.md` | | Need full NDJSON event schema (`run`) | `references/parsing.md` | -| Need the `generate` NDJSON event schema | `references/generate-parsing.md` | | Browse / create projects or folders, or parse the auto-default event | `references/test-manager.md` | | Start of every session: preflight, the ready card, sign-in | `references/ready-check.md` | | A person's first session: run first, the tour, three choices | `references/first-run.md` | @@ -388,7 +364,7 @@ Internal event/field names (`generate_snapshot`, `request_id`, …) are for pars ## Command-specific completion -The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). `generate` emits `generate_done`. Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. +The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. Progress is for live display: count only `done`/`failed` completions, retaining child and execution context when step indices repeat. diff --git a/.claude/skills/kane-cli/references/assurance-parsing.md b/.claude/skills/kane-cli/references/assurance-parsing.md index a69d44e..9562a82 100644 --- a/.claude/skills/kane-cli/references/assurance-parsing.md +++ b/.claude/skills/kane-cli/references/assurance-parsing.md @@ -41,8 +41,8 @@ A stream that ends **without** `done` means the process crashed — outcome unkn | `held` / `update_held` *(0.7.1+)* | items held for the user's review instead of committed: `source_id` + `count` + `reason` / `count` + `targets[]` | surface the count and that review happens at resume/`context review` | | `commit` | what landed: counts + `minted[]` (`cid` + `logical_id`); extract adds `proposal_id` | translate ("5 use-cases extracted"); `logical_id` slugs are how you reference nodes later | | `receipt` | per-phase commit receipt (design; extract also emits one at its commits): `commit_n`, `phase`, `committed[]`, `reused`, `rejected[]`, `warnings[]`, `next`, and (design only) `parity` | surface non-empty `rejected[]` and `warnings[]` in plain language; meaningful reuse is worth one line | -| `variables_declared` *(0.8.12+)* | design: stubs written by a phase commit — `file` (pool file) + `variables[]` (`name`, `description`, `secret`); only names that existed in no variable file | surface as a to-do ("N variables need values — in ") | -| `variables_summary` *(0.8.12+)* | design, end of run: every declared name still needing a value — same shape | repeat the to-do in the closing summary; fills gate authoring | +| `variables_declared` *(0.8.12+)* | design: the stubs the tests phase wrote, sent just before that phase's `commit` event — `file` (pool file) + `variables[]` (`name`, `description`, `secret`); only names that existed in no variable file | surface as a to-do ("N variables need values — in "), then fill them before any run of these tests (SKILL.md §3) | +| `variables_summary` *(0.8.12+)* | design, after `session_complete` on a clean completion only (a paused or refused run never sends it): every name this run declared, same shape, no re-check of the pool | repeat the to-do in the closing summary; on a pause build it from `variables_declared` alone; unfilled names are typed as written when the test is authored | | `message_sent` | `--message` delivered: `sid`, `chars` | confirmation only | | `panel_resolved` *(0.7.1+)* | a `--answer` flag landed on a pending question: `id`, `by`, `via` | confirmation only | | `ask_deferred` *(0.7.1+)* | `--with-source` set the pending batch aside: `source_id`, `cid`, `questions` (count) | tell the user the questions were deferred while the agent reads the new source | @@ -99,4 +99,4 @@ One payload event carrying the full `--json` document — `coverage` for `cover` | `3` | **paused and resumable** — not a failure; run the pause loop. Includes crash-pauses (0.7.1+). On the sync verbs (0.8.14+) `done{status: "paused"}` means decisions are waiting: after a rebase walk the rebase is open (answer with `--answer`); after `kane-cli context sync doctor --abort` it is closed and the unanswered decisions stayed in the backup — the `sync_rebase_done` before it says which (`status` `paused` or `aborted`); `paused` with no `sync_rebase_done` means doctor set aside a rebase whose saved state it could not read — `sync_doctor.detail` says the next `kane-cli context sync` reapplies the saved backup. `done{status: "refused", exit_code: 3}` means a pull or a rebase is needed first — `references/context-sync.md` §4 | | `130` | force-interrupted — for extract and design, resumable only if a `session_paused` event arrived; on the sync verbs the rebase state is on disk: `kane-cli context sync doctor --mode agent` shows it, `kane-cli context sync --mode agent` resumes it | -Reminder: this exit-3 meaning is **specific to these assurance commands**. `run` / `testmd` / `testrun` / `generate` keep their own meanings (3 = timeout/cancelled). +Reminder: this exit-3 meaning is **specific to these assurance commands**. `run` / `testmd` / `testrun` keep their own meanings (3 = timeout/cancelled). diff --git a/.claude/skills/kane-cli/references/assurance.md b/.claude/skills/kane-cli/references/assurance.md index cab6a28..05b188a 100644 --- a/.claude/skills/kane-cli/references/assurance.md +++ b/.claude/skills/kane-cli/references/assurance.md @@ -2,9 +2,8 @@ # Assurance — Agent Surface -When the user has **requirements** — a PRD, a spec, acceptance notes — and wants tests designed from them, wants to know what's covered, or wants the suite kept current, use the **assurance commands** (`kane-cli context`, `design`, `cover`, `maintain`). Do not hand-write the tests, and do not reach for `generate`: +When the user has **requirements** — a PRD, a spec, acceptance notes — and wants tests designed from them, wants to know what's covered, or wants the suite kept current, use the **assurance commands** (`kane-cli context`, `design`, `cover`, `maintain`). Do not hand-write the tests: -- `kane-cli generate` = quick scenarios/cases from a one-line description. No requirement linkage. - **Assurance** = tests derived from the actual documents, every claim cited, every test permanently tagged with the acceptance criteria it verifies, coverage measured against requirements. Use it whenever the user cares about "what exactly is covered, and how do we know?" Everything here works over a local store (`.context/` in the project directory) that the commands create and manage themselves. @@ -19,7 +18,7 @@ kane-cli context review --verdicts --json # 2. CHECKPOINT: us kane-cli design tests --use-case --mode agent --max 8 # 3. design ACs, scenarios, tests kane-cli context review --verdicts --json # 4. CHECKPOINT: user approves the design kane-cli testmd run .testmuai/tests/_test.md --agent # 5. author each kept test once (real browser) -kane-cli testrun run --match 't-' # 6. batch replays from then on +kane-cli testrun run --match 't-' < /dev/null # 6. batch replays from then on kane-cli cover gaps # 7. designed % × proven % + per-use-case debt kane-cli maintain reconcile --from --source-id --mode agent # when a source changes (§11) ``` @@ -35,12 +34,12 @@ Extract, design, and reconcile call the KaneAI service and consume credits; ever ## 2. The pause loop — exit 3 is a pause, NOT a failure -**Scope: this rule applies to `context extract`/`context ingest`, `design tests`, and `maintain reconcile` (§11).** (For `run`/`testmd`/`testrun`/`generate`, exit 3 still means timeout/cancelled.) +**Scope: this rule applies to `context extract`/`context ingest`, `design tests`, and `maintain reconcile` (§11).** (For `run`/`testmd`/`testrun`, exit 3 still means timeout/cancelled.) These commands take **`--mode agent`** — not `--agent`; they reject that flag, and a bare non-TTY invocation exits `2` asking for an explicit mode. In `--mode agent`, **every question pauses the run** (0.8.8+ — earlier CLIs auto-answered low/medium-risk questions with their recommended defaults and paused only on high risk): - The run exits `3`, emits `session_paused` with the session id, the questions in full (text, options, the recommended one, risk, rationale), and the verbatim resume command. -- **Never drop a pause** (same rule as generate clarifications). Answer it: if your own context clearly resolves the question, answer it yourself; otherwise surface the question — with its options and recommendation — to your user and get their answer. +- **Never drop a pause.** Answer it: if your own context clearly resolves the question, answer it yourself; otherwise surface the question — with its options and recommendation — to your user and get their answer. - **0.7.1+ sessions are durable from the first turn**: a crash that left a checkpoint exits `3` with a `session_paused` carrying `crashed: true` (no `pending_questions`) and the resume command — exit 3 always means "resumable". A crash before anything durable was saved still exits `1`; check `context sessions --json` before retrying anything paid. Three ways to resume: @@ -119,13 +118,13 @@ kane-cli design tests --use-case --mode agent --max 8 - There is **no `--because` flag on `design tests`** — interactively the session collects the redesign reason itself; headless `--force` proceeds with an auto-stamped reason. - `--phase ` (0.7.1+) re-enters a design at a phase, re-seeded from the committed earlier phases; missing predecessors exit 2 with the commands to run first in `next`. - Output: acceptance criteria, scenarios, exactly one test per scenario — written as runnable files under `.testmuai/tests/*_test.md`, each assert step tagged with the criteria it verifies. Plus **gaps** (recorded, ranked missing pieces) and **warnings** (e.g. a test claiming more criteria than its check asserts). Citations are verified against the pinned source text before commit (0.7.1+) — a `CITE_UNVERIFIED` error means a citation could not be verified even after repair. -- **Variables (0.8.12+).** Design reads every `*.json` in the user's variable directories (names + has-value only, never values) and reuses a name that exists before inventing one. Each name it invents is declared: an empty stub lands in `.testmuai/variables/assurance.json`, never a value, never over an existing key. The stream tells you which: `variables_declared` at the commit that wrote them, `variables_summary` at the end (schema in `references/assurance-parsing.md`). **Surface these as a to-do** — "2 variables need values: login_email, login_password — in .testmuai/variables/assurance.json" — because a designed test refuses to author until they are filled (SKILL.md §3, Unresolved variables). Fill them yourself only if the user gave you the values. +- **Variables (0.8.12+).** Design reads every `*.json` in the user's variable directories (names + has-value only, never values) and reuses a name that exists before inventing one. Each name it invents is declared: an empty stub lands in `.testmuai/variables/assurance.json`, never a value, never over an existing key. The stream tells you which: `variables_declared` in the tests phase, just before its `commit` event, and `variables_summary` after `session_complete` on a clean completion (a paused run sends only the first; schema in `references/assurance-parsing.md`). **Surface these as a to-do** — "2 variables need values: login_email, login_password — in .testmuai/variables/assurance.json" — because a designed test authored before they are filled types the placeholders as written (SKILL.md §3, Unresolved variables). Before any run of these tests, fill them: SKILL.md §3 **Fill the variables before any run**, asking with each row's `description`. - **Present tests, gaps, warnings, AND variables needing values** — first-class output, not noise. Then go to the review checkpoint (§4) before any authoring. - `kane-cli design explain ` replays *why* a test exists (technique, boundary values, criteria) with zero AI cost — use it when the user asks "why this test?". ## 6. The authoring bridge — from designed files to batch runs -A freshly designed test has never been executed. On 0.8.4+ hand the set straight to `kane-cli testrun run`: unauthored members classify as `author`, the run authors them in a real browser, and afterwards the authored and replayed evidence consolidates into one published execution — best-effort: when consolidation cannot complete, evidence stays split rather than lost. `--from-context` (0.8.4+) selects members by assurance test ids and follows edit supersessions. Designed tests may carry `{{variables}}` for values the requirements never pinned (a store URL, a product name) — supply them per `references/testmd.md`. `kane-cli testmd run --agent` remains the single-test authoring path, and on pre-0.8.4 CLIs it is REQUIRED first — `testrun` there refuses never-authored members (`missing_meta`). Evidence packs seal per `references/evidence.md`. +A freshly designed test has never been executed. On 0.8.4+ hand the set straight to `kane-cli testrun run`: unauthored members classify as `author`, the run authors them in a real browser, and afterwards the authored and replayed evidence consolidates into one published execution — best-effort: when consolidation cannot complete, evidence stays split rather than lost. `--from-context` (0.8.4+) selects members by assurance test ids and follows edit supersessions. Designed tests may carry `{{variables}}` for values the requirements never pinned (a store URL, a product name) — fill them first (SKILL.md §3), or the run types the placeholder as written. `kane-cli testmd run --agent` remains the single-test authoring path, and on pre-0.8.4 CLIs it is REQUIRED first — `testrun` there refuses never-authored members (`missing_meta`). Evidence packs seal per `references/evidence.md`. ### 6.1 A designed test is the design — do not edit it by hand diff --git a/.claude/skills/kane-cli/references/cards.md b/.claude/skills/kane-cli/references/cards.md index 9904459..b035268 100644 --- a/.claude/skills/kane-cli/references/cards.md +++ b/.claude/skills/kane-cli/references/cards.md @@ -16,6 +16,7 @@ Every result is an emoji table. A one-line "Test passed" instead of the card is - **`🟡 Didn't start` is not `🔴 Failed`.** When nothing ran, say what to fix. - **Secret-looking values never go in chat.** For a missing value whose name contains `password`, `secret`, `token` or `key`, add an empty entry to the variables file for the person to fill. Ask in chat only for plain values (a URL, a user name). - If the run's output carried an update notice, add one quiet last line under the card: `kane-cli is available.` +- **Variables with no value go on the card.** When the run's output carried the `unresolved_variables` warning, add a `⚠️ **Variables**` row before ➡️ Next naming each one. ## 2. Run, passed @@ -72,7 +73,7 @@ Exit code `1`, or `status: "failed"`. Show the failing step's screenshot under t ## 4. Didn't start -Exit code `2`: nothing ran and no credits were used. Causes include missing variable values, no start URL, sign-in or setup errors, a test file that does not parse, an invalid suite plan, a cloud grid refusal. +Exit code `2`: nothing ran and no credits were used. Causes include no start URL, sign-in or setup errors, a test file that does not parse, an invalid suite plan, a cloud grid refusal. ```markdown | | | diff --git a/.claude/skills/kane-cli/references/debug.md b/.claude/skills/kane-cli/references/debug.md index 7d8faf0..372d188 100644 --- a/.claude/skills/kane-cli/references/debug.md +++ b/.claude/skills/kane-cli/references/debug.md @@ -66,7 +66,7 @@ For `run` objectives and plain `_test.md` files. For a designed test, read the s | 🎯 Agent clicks wrong element | Ambiguous UI, multiple similar elements | Be more specific: "click the **blue** 'Submit' button in the **checkout form**" | | 👁️ Agent says done but didn't finish | Objective too vague | Add explicit assertions: "assert the confirmation page shows order number" | | 💀 Exit code 2, no steps | Auth, TMS credential exchange, or Chrome failure | Check `kane-cli whoami`, verify Chrome is available | -| ❓ Exit code 2 with "did you mean …" | Bare-objective shortcut — agent ran `kane-cli ""` without the `run` subcommand | Re-invoke as `kane-cli run "" --agent` (same rule for `testmd run` / `generate`) | +| ❓ Exit code 2 with "did you mean …" | Bare-objective shortcut — agent ran `kane-cli ""` without the `run` subcommand | Re-invoke as `kane-cli run "" --agent` (same rule for `testmd run`) | | 📤 Upload silently fails after configuring a project/folder by hand | Saved ID is invalid (typo, deleted, no access) | No action needed — the next run detects the 4xx and auto-defaults a working project/folder. To rebind manually: `kane-cli config project` (TTY picker) or `kane-cli projects list` → `kane-cli config project ` (see `references/test-manager.md`) | | ⏱️ Exit code 3 | Timeout or cancelled | Increase `--timeout` or `--max-steps`, or split into smaller objectives | | 🚫 "CDP endpoint not reachable" | Chrome not running | Let kane-cli manage Chrome (remove `--cdp-endpoint`) | diff --git a/.claude/skills/kane-cli/references/generate-parsing.md b/.claude/skills/kane-cli/references/generate-parsing.md deleted file mode 100644 index d767ebe..0000000 --- a/.claude/skills/kane-cli/references/generate-parsing.md +++ /dev/null @@ -1,65 +0,0 @@ - - -# Reading `generate --agent` Output - -> **Internal reference only.** The event types and field names below are for you to parse programmatically. **Never expose them to the user** — present plain-language scenarios/cases per `references/generate.md`, never `generate_snapshot`, `request_id`, `generate_done`, or raw JSON. - -With `--agent`, `kane-cli generate` writes **one JSON object per line** to **stdout**. Unlike `run` (which has untyped `step`/`status` progress lines), **every generate line is typed** — it always has a `type` field. That makes parsing simpler: - -``` -for each line of NDJSON: - parse JSON, switch on obj.type - if obj.type === "generate_done" → terminal event, stop parsing - if obj.type === "generate_snapshot" → the deliverable (full scenarios + cases) - if obj.type === "generate_clarification" → turn ended awaiting an answer (still exit 0) - else → progress / informational -``` - -Build post-turn logic on **`generate_done`** (terminal, stable schema) and read the result from **`generate_snapshot`** (emitted once, at turn end). - -## Event types - -Generate-specific (typed `generate_*`): - -| `type` | Key fields | Meaning → what to do | -|---|---|---| -| `generate_upload` | `file`, `index`, `total`, `status` | Only when `--files` is used. Emitted once per attached file **before `generate_start`**, while the file uploads; `status` goes `"uploading"` → `"done"` (or `"failed"`). Progress only — narrate "attaching …" or ignore. A failed upload surfaces as an `error` + non-zero exit. | -| `generate_start` | `request_id`, `objective_chars`, `scenario_limit`, `per_scenario_limit`, `is_refine` | Turn began. **Capture `request_id`** — it's the handle for every later `--refine` / `--save`. `is_refine` distinguishes a new request from a continuation. | -| `generate_thinking` | `took_ms` | Liveness only. Narrate "thinking…" or ignore. | -| `generate_progress` | `pct` | Milestone (25 / 50 / 75 / 100). Optional progress display — not a completion signal. | -| `generate_snapshot` | `scenario_count`, `case_count`, `scenarios[]` | **The deliverable** — full scenarios, each with its cases. Each case carries `title`, `polarity` (`"p"`/`"n"`/`"e"` = Positive / Negative / Edge), `category` (`"Functional"`, `"Security"`, …), `priority`. Present it per `generate.md`. Emitted exactly once. | -| `generate_clarification` | `text` | The generator needs an answer; the turn ended (exit 0) awaiting it. **Not an error** — answer via `--refine --req` (see `generate.md`). | -| `generate_chat` | `text` | The model's prose reply (e.g. what a refine changed). Show as info. | -| `generate_save_result` | `suite_dir`, `saved`, `fell_back`, `warning?` | `--save` wrote files. `saved` = files written; `fell_back` = cases written as prose because they couldn't be expanded; `warning` e.g. `"no functional test cases"`. | -| `generate_done` | `request_id`, `status`, `scenario_count`, `case_count`, `refine_hint`, `save_hint`, `suite_dir?` | **Terminal line.** Branch on `status`; use `refine_hint` / `save_hint` **verbatim** as the next command. `suite_dir` present only after a `--save`. | - -Shared with `run` (NOT generate-specific — don't treat as part of the generate contract): - -| `type` | Fields | Meaning | -|---|---|---| -| `project_folder_auto_defaulted` | resolved project + folder (id, name) | Run-startup gate auto-resolved a project/folder when none was configured (or the cached one was stale/invalid). Fires before `generate_start`. Translate to plain language (see `references/test-manager.md`). | -| `error` | `message` | A failure occurred; pair with the exit code to decide retry vs abort. | -| `update_available` | `current`, `latest`, `severity` | A newer kane-cli exists. Informational; emitted before `generate_start`. | - -## Terminal `generate_done` - -Example after a **`--save`** run (note `suite_dir` — a new/refine turn omits it): - -```json -{ - "type": "generate_done", - "request_id": "23271", - "status": "completed", - "scenario_count": 3, - "case_count": 11, - "refine_hint": "kane-cli generate \"\" --refine --req 23271", - "save_hint": "kane-cli generate --save --req 23271", - "suite_dir": ".testmuai/tests/checkout-23271" -} -``` - -`status`: `completed` (exit 0 — including a turn that ended with a clarification) · `failed` or `ended` (exit 1) · `stopped` (exit 3). `suite_dir` is present only after a `--save`. The hints carry no `--agent` flag, but it is auto-on for non-TTY callers (agents/pipes), so re-invoking a hint verbatim still yields NDJSON. - -## Exit codes - -`0` completed · `1` failed · `2` error (auth / setup / transport, or an invalid flag combination) · `3` stopped / cancelled · `130` interrupted (Ctrl-C). Interactive `--refine`/`--save` continuity is by re-invocation with `--req` — there is no stdin relay. diff --git a/.claude/skills/kane-cli/references/generate.md b/.claude/skills/kane-cli/references/generate.md deleted file mode 100644 index d947c1e..0000000 --- a/.claude/skills/kane-cli/references/generate.md +++ /dev/null @@ -1,139 +0,0 @@ - - -# Generating Test Cases with `kane-cli generate` - -`kane-cli generate` turns a plain-language description of *what to test* into structured **Test Scenarios** (logical groupings) each containing **Test Cases** (typed Positive / Negative / Edge). It calls the AI Test Case Generator — **no browser is launched**. The result is a tree of scenarios + cases you present to the user, refine conversationally, and optionally save as runnable `_test.md` files. - -**Use this whenever a task needs test cases or scenarios written — don't hand-author them in chat or a scratch file.** Reach for it to: turn a requirement / feature description into a test suite; expand or refine coverage (more edge cases, negative paths, a narrower or broader focus); or save the Functional cases as runnable `_test.md`. It is **not** for driving a browser — that's `kane-cli run` (§3 of SKILL.md). - -> For the full web-product picture (dashboards, issue-link inputs), see the public docs: . The CLI takes a **text objective**, optionally with local files attached via **`--files`** (see "Attaching files" below). - -## The three modes — one turn per invocation, then exit - -There is **no interactive session**. Each invocation runs exactly one generation turn and exits. Continuity across turns is carried by a **request id** (`--req `) that the previous turn's terminal line hands back. - -| Mode | Command | Notes | -|---|---|---| -| **New** | `kane-cli generate "" --agent` | Starts a fresh request. Capture the request id from the terminal line. | -| **Refine** | `kane-cli generate "" --refine --req --agent` | Adjusts an existing request. `--refine` **and** `--req` required; needs a change description. | -| **Save** | `kane-cli generate --save --req [--out ] --agent` | Writes the request's Functional cases to `_test.md`. No new turn, takes no objective. | - -`--refine` and `--save` always run headless (even from a terminal). - -### Flags - -| Flag | Purpose | -|---|---| -| `--agent` | Typed NDJSON on stdout (auto-on when stdin is not a TTY) | -| `--req ` | The request id to `--refine` or `--save` | -| `--out ` | Save target — **only** with `--save`; default `/.testmuai/tests` | -| `--name ` | Names the run and the saved suite folder | -| `--scenario-limit ` / `--per-scenario-limit ` | Cap scenarios / cases-per-scenario | -| `--memory` | Use the memory layer — reuse relevant existing cases, reduce duplicates | -| `--files ` | Comma-separated local files to attach as context (new / refine only — see "Attaching files") | -| `--project ` / `--folder ` | Test Manager project / folder | -| `--username` / `--access-key` | Auth (same as `run`) | - -If neither `--project`/`--folder` nor a saved project/folder is set when generation starts, kane-cli auto-resolves one headlessly and emits a `project_folder_auto_defaulted` event before `generate_start`. Translate it to a one-line note for the user — full handling lives in `references/test-manager.md`. - -## Attaching files - -Pass local files as extra context with `--files ` on a **new** generation or a **`--refine`** (not `--save`). The generator reads them and reflects them in the scenarios + cases — attach a spec, a screenshot of the UI, a PDF / Word doc, or a CSV of inputs. - -```bash -kane-cli generate "test the login flow described in the attached spec" --files ./login-spec.pdf,./wireframe.png --agent -``` - -- **Supported types** — documents (`.txt .json .xml .csv .pdf .docx .xlsx`), images (`.jpg .jpeg .png .gif .bmp .webp`), audio (`.mp3 .wav .m4a`), video (`.mp4 .mov .webm .mpeg .mpga`). -- **Limits** — up to **10 files**, each **≤ 50 MB**. -- **Validated as a set, up front** — if any path is missing, an unsupported type, too large, or over the count, the whole command is rejected (exit `2`) **before anything is sent** and the offending paths are listed; fix and re-run. Files outside the current directory are allowed but flagged with a warning on stderr. -- **`new` / `--refine` only** — combining `--files` with `--save` exits `2`. - -Under `--agent`, each file emits a `generate_upload` line (`status` `uploading` → `done`) **before** `generate_start` — see `references/generate-parsing.md`. *(Interactively in the TUI, type `@` in the generate prompt to attach a file inline.)* - -## Presenting a result (adaptive) - -The terminal data carries the full scenarios + cases. **Present it based on size**: - -- **≤ ~30 cases → a nested tree** (scenario, then each case with its type tag): - ``` - ✓ Generated 3 scenarios · 11 cases (request 23271) - - ▸ Login - - Valid credentials [Positive] - - Wrong password [Negative] - - Empty fields [Edge] - ▸ Checkout - - Guest checkout [Positive] - - Expired card [Negative] - ... - ``` -- **more than ~30 cases → a summary + scenario list** (cases on request): - ``` - ✓ Generated 6 scenarios · 84 cases (request 23271) - - • Login (12 cases) - • Checkout (20 cases) - • Cart management (14 cases) - ... - ``` - -Always end with the **next-step commands the terminal line provides** (Refine / Save) — they already carry the request id, so don't hand-build them. - -## Clarifications — act on them, never drop them - -If a turn ends with a **clarification question**, that is **success (exit 0)**, not an error — the generator needs an answer before it can continue. You must act on it: - -1. Read the question. -2. **Decide** — answer it yourself from context, **or** surface it to your user and get an answer. -3. **Re-invoke** with the answer as a refine: - ```bash - kane-cli generate "" --refine --req --agent - ``` - -## The refine → save → run loop - -```bash -# 1. New request -kane-cli generate "checkout flow on a shopping site" --agent -# → terminal line carries request id 23271 + Refine/Save hints - -# 2. Refine (repeat as needed) -kane-cli generate "also cover an expired card and an out-of-stock item" --refine --req 23271 --agent - -# 3. Save the Functional cases as runnable _test.md -kane-cli generate --save --req 23271 --agent -# → /.testmuai/tests///_test.md - -# 4. Run / replay them -kane-cli testmd run .testmuai/tests///_test.md --agent -``` - -**Save is Functional-only.** `--save` writes only test cases whose category is **Functional** — those are the ones runnable as `_test.md`. Non-functional cases (Security, Performance, etc.) are generated and shown in the result but are **not** written; saving a request with no Functional cases writes nothing and says so. Saved files are ordinary `_test.md` tests — see `references/testmd.md` for running, editing, and replay. This is the **generate → testmd** pipeline: author cases here, run them there. - -## Exit codes - -| Code | Meaning | -|---|---| -| 0 | Turn completed (including a turn that ended with a clarification) | -| 1 | Generation failed (or ended) | -| 2 | Error — auth / setup / transport, or an invalid flag combination | -| 3 | Generation stopped / cancelled | -| 130 | Interrupted (Ctrl-C) | - -Invalid flag combinations exit `2` with a message on stderr. The full set: - -- `--refine` and `--save` together -- `--refine` without `--req` -- `--refine` without a change description -- `--refine` combined with `--out` (`--out` is save-only) -- `--save` without `--req` -- `--save` with a description (it takes none) -- `--out` without `--save` -- `--files` with `--save` (files attach to a new generation or a refine, not a save) -- `--req` without `--refine` or `--save` -- a new generation with no description - -## Reading the output - -`--agent` emits one typed JSON object per line. For the full event schema and the parse strategy, Read **`references/generate-parsing.md`**. As with all `--agent` output, the field names are for parsing only — **never show them to the user**; present plain-language scenarios and cases (per "Presenting a result" above). diff --git a/.claude/skills/kane-cli/references/mobile.md b/.claude/skills/kane-cli/references/mobile.md index 222a8c4..eca231e 100644 --- a/.claude/skills/kane-cli/references/mobile.md +++ b/.claude/skills/kane-cli/references/mobile.md @@ -9,10 +9,10 @@ Desktop (the browser) is the **default** target and the primary use of kane-cli. | Where the device runs | Host | Commands | |---|---|---| | **Local** (a simulator/emulator on this machine) | **macOS on Apple Silicon (arm64) only** — not Intel Macs, Linux, or Windows | `run --target …`, `testmd run`, `testrun run` | -| **Cloud grid** (a virtual device on a HyperExecute macOS host) | **Any machine** — no Xcode / Android Studio needed. The account needs a HyperExecute plan with macOS runners | `testrun run --remote …` — see §Remote below | +| **Cloud grid** (a virtual device on a HyperExecute macOS host) | **Any machine** — no Xcode / Android Studio needed. The account needs a HyperExecute plan with macOS runners | `testrun run --remote … < /dev/null` — see §Remote below | - **Desktop stays the default.** The `--target` axis is what selects mobile. Leave it off and you get the browser. -- If the user is not on mac-arm64, local mobile is not an option — **offer the grid**: save the objective as a `_test.md` (`target: emulator|simulator` + `app:`) and run it with `kane-cli testrun run --remote --device-name "" --os-version `. Do not tell them mobile is unavailable. +- If the user is not on mac-arm64, local mobile is not an option — **offer the grid**: save the objective as a `_test.md` (`target: emulator|simulator` + `app:`) and run it with `kane-cli testrun run --remote --device-name "" --os-version < /dev/null`. Do not tell them mobile is unavailable. ## The three targets @@ -111,7 +111,7 @@ Everything else about `_test.md` (step bodies, replay/cascade, commands) is unch `kane-cli testrun run` accepts mobile `_test.md` members (0.8.7+): -- **Locally** (mac-arm64 with the setup above): `kane-cli testrun run --device-name "" --os-version ` — the device as `kane-cli devices list --target ` prints it, or the members' own `device_name:`/`os_version:`. +- **Locally** (mac-arm64 with the setup above): `kane-cli testrun run --device-name "" --os-version < /dev/null` — the device as `kane-cli devices list --target ` prints it, or the members' own `device_name:`/`os_version:`. - **On the cloud grid, from any machine**: add `--remote` — next section. Full testrun flags, events, and rollup: `references/testrun.md`. ## Remote: mobile suites on the cloud grid (`testrun run --remote`) @@ -121,9 +121,9 @@ One command turns a mobile suite into a HyperExecute job on a macOS host that bo ```bash kane-cli plugin install remote-execution # once; then `kane-cli plugin doctor remote-execution` kane-cli devices list --target emulator --remote --agent # grid catalog: name + os_versions per row -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 --dry-run # validate + resolve device, no job -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 -kane-cli testrun run tests/ios/ --remote --device-name "iPhone 15" --os-version 17.5 +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 --dry-run < /dev/null # validate + resolve device, no job +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 < /dev/null +kane-cli testrun run --match '^tests/ios/' --remote --device-name "iPhone 15" --os-version 17.5 < /dev/null ``` **Always `--dry-run` first** — it runs the remote preflight and resolves the device against the catalog at no cost. Use `Bash` with a long timeout (up to 600000 ms) for the real run: device setup + app install + members take several minutes. diff --git a/.claude/skills/kane-cli/references/parallel.md b/.claude/skills/kane-cli/references/parallel.md index 5bc51bb..0866044 100644 --- a/.claude/skills/kane-cli/references/parallel.md +++ b/.claude/skills/kane-cli/references/parallel.md @@ -4,7 +4,7 @@ For multiple independent browser tasks, decompose and run in parallel using the Agent tool. -> **Saved tests? Use testrun instead.** If the tasks are committed `_test.md` files, do NOT hand-roll parallelism — `kane-cli testrun run --parallel N` gives you isolated Chromes, a pooled scheduler, one rollup, and one evidence pack. Read `references/testrun.md`. This reference is for **ad-hoc `run` objectives** only. +> **Saved tests? Use testrun instead.** If the tasks are committed `_test.md` files, do NOT hand-roll parallelism — `kane-cli testrun run --parallel N < /dev/null` gives you isolated Chromes, a pooled scheduler, one rollup, and one evidence pack. Read `references/testrun.md`. This reference is for **ad-hoc `run` objectives** only. ## When to Split diff --git a/.claude/skills/kane-cli/references/parsing.md b/.claude/skills/kane-cli/references/parsing.md index 2e6aac9..ec3952e 100644 --- a/.claude/skills/kane-cli/references/parsing.md +++ b/.claude/skills/kane-cli/references/parsing.md @@ -55,24 +55,25 @@ These are **untyped** — they have no `type` field. Do **not** key on `event.ty | Event (`type` field) | Key Fields | Purpose | |-------|-----------|---------| -| `project_folder_auto_defaulted` | resolved project + folder (id, name) | Run-startup gate auto-resolved a project/folder when none was configured (or the cached one was stale/invalid). Fires before any progress event on `run` / `testmd run` / `generate`. Translate to plain language (see `references/test-manager.md`). | +| `project_folder_auto_defaulted` | resolved project + folder (id, name) | Run-startup gate auto-resolved a project/folder when none was configured (or the cached one was stale/invalid). Fires before any progress event on `run` / `testmd run`. Translate to plain language (see `references/test-manager.md`). | | `bifurcation` | `flows[]`, `count` | Agent split objective into sub-flows | | `child_agent_start` | `child_id`, `objective`, `parent_step` | Child agent spawned | | `child_agent_end` | `child_id`, `success`, `steps_taken`, `summary` | Child agent finished | | `ask_user` | `question`, `step_index`, `options?` | Agent needs user input | -| `error` | `message`, `code?` | Error occurred. With `code: "unresolved_variables"` it is the pre-run refusal — the only event of the run, no `run_end` follows; schema below. | +| `error` | `message` | Error occurred | +| `warning` | `code`, `message`, … | *(0.8.15+)* Before any progress event: `code: "unresolved_variables"` — a `{{name}}` had no value; the run continues. Schema below. | | `test_md_evidence_ingest` | `status: "ok"\|"failed"`, `evidence_id`, `stage?` (failure only) | `testmd run` only: a replay's evidence pack published to the dashboard. Informational. | | `test_md_bundle_sync` | `status: "ok"\|"failed"`, `commit_id`, `bytes?` (success) / `stage?` (failure) | `testmd run` / `testmd sync`: test bundle pushed to the cloud after an authored commit. Informational. | | `testrun_*` family | see `references/testrun.md` | Emitted only by `kane-cli testrun run`; terminal event is `testrun_done`, not `run_end`. | **Note:** The `run` stream has no `run_start` event; startup metadata or errors can precede progress. -### `error` with `code: "unresolved_variables"` (0.8.12+) +### `warning` with `code: "unresolved_variables"` (0.8.15+) -Emitted by `run`, `testmd run` and (after `testrun_plan`) `testrun run` when an authored step references a `{{name}}` that has no value. Nothing was dispatched; exit code `2`; stderr is silent. +Emitted by `run`, `testmd run` and (after `testrun_plan`) `testrun run` when an authored step references a `{{name}}` that has no value. The run goes ahead: the name is typed as written unless a step sets it first. The warning does not change the exit code. A run given a Test Manager dataset row can emit a second `warning` with the same code for `${x}` names that are no column of that row, still before any progress. ```json -{"type":"error","code":"unresolved_variables","message":"2 variable(s) have no value — nothing was dispatched", +{"type":"warning","code":"unresolved_variables","message":"2 variable(s) have no value — typed as written unless a step sets them first", "suggested_file":".testmuai/variables/variables.json", "variables":[ {"name":"checkout_url","reason":"not_declared","used_by":[{"file":"objective","step":1}]}, @@ -82,14 +83,12 @@ Emitted by `run`, `testmd run` and (after `testrun_plan`) `testrun run` when an | Field | Meaning | |---|---| -| `variables[].reason` | `value_missing` — the key exists in `variables[].file` with an empty value · `not_declared` — the key is in no variable file | +| `variables[].reason` | `value_missing` — the key exists in `variables[].file` with an empty value · `not_declared` — the key is in no variable file · `not_a_dataset_column` — a `${x}` that is no column of the Test Manager dataset the run was given (`variables[].dataset` names it) | | `variables[].file` | the pool file that holds the empty key (`value_missing` only); `--variables` when it came inline | | `variables[].used_by[]` | `{file, step}` — `file` is `objective` for `kane-cli run`, else the test file (flattened step index) | | `suggested_file` | where to add a `not_declared` key: `.testmuai/variables/assurance.json` inside an assurance store, `variables.json` otherwise | -Terminal: do not re-run the same command. Supply values (`--variables`, or fill the file) and run again. - -**Note:** `ask_user` is auto-disabled when stdin is not a TTY. Since agents typically run kane-cli as a subprocess, ask_user events will not be emitted. Write objectives that don't require interactive input. +Report it with the run's outcome. If a later step failed on a literal placeholder, point at this. ## Parsing Strategy for one-shot `run` @@ -159,6 +158,6 @@ To cancel a run: ## Command-specific completion -The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). `generate` emits `generate_done`. Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. +The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. Progress is for live display: count only `done`/`failed` completions, retaining child and execution context when step indices repeat. diff --git a/.claude/skills/kane-cli/references/setup-and-config.md b/.claude/skills/kane-cli/references/setup-and-config.md index b876526..ace5e83 100644 --- a/.claude/skills/kane-cli/references/setup-and-config.md +++ b/.claude/skills/kane-cli/references/setup-and-config.md @@ -82,7 +82,7 @@ Variables parameterize objectives with reusable values and secrets. Use `{{key}} **Values (0.8.12+):** a string is a value when non-empty; a number is accepted and loaded as its string; a boolean, object or array is not a value. An empty `value` is a declared-but-unfilled key. -**Before a run (0.8.12+):** every `{{name}}` an authored step references must have a value — `run`, `testmd run` and `testrun run` refuse before launching anything otherwise (exit `2`; with `--agent`, one `error` event with `code: "unresolved_variables"` — SKILL.md §3). No bypass flag. +**Before a run (0.8.15+):** `run`, `testmd run` and `testrun run` check every `{{name}}` an authored step references before launching anything. A name with no value is a warning (with `--agent`, one `warning` event with `code: "unresolved_variables"`; SKILL.md §3): the run goes ahead and types the name as written unless a step sets it first. **`assurance.json`:** `kane-cli design tests` writes empty stubs (`{"name":{"value":"","secret":false,"description":"…"}}`) into `{cwd}/.testmuai/variables/assurance.json` for every variable it declares, never a value, and never over a key that already exists in any pool file. Fill those before authoring the designed tests. diff --git a/.claude/skills/kane-cli/references/test-manager.md b/.claude/skills/kane-cli/references/test-manager.md index 91653cf..65e9030 100644 --- a/.claude/skills/kane-cli/references/test-manager.md +++ b/.claude/skills/kane-cli/references/test-manager.md @@ -104,7 +104,7 @@ To use the result for subsequent runs, persist with `kane-cli config project ` | Filter candidates by project-relative path regex | — | +| `--match ` | Filter candidates by project-relative path regex | The path is as the OS writes it: `tests/app/` on macOS and Linux, `tests\app\` on Windows. Quote the regex with double quotes in cmd.exe; single quotes are literal there. | | `--tags ` | ANY-match on frontmatter `tags:` (repeatable or comma-separated, case-insensitive) | — | | `--parallel ` | Worker count; each desktop worker gets an isolated Chrome with a fresh temp profile | `1` | | `--on-failure ` | `continue` (run everything) \| `fail-fast` (stop dispatching new members after a failure) | `continue` | @@ -45,16 +45,16 @@ All members must share one org + project. *(0.8.4+)* Members need **not** be aut | `org_mismatch` | Different organisation than the other tests | "Check `kane-cli testmd status ` — it belongs to another org" | | `project_mismatch` | Different project than the other tests | "Run it separately or per-project" | -If **any** member fails preflight, the plan is invalid: nothing runs, exit `2`. Suggest `--dry-run` to preview the plan cheaply before a big run. *(0.8.12+)* Preflight also checks variables: a member whose authored steps reference a `{{name}}` with no value fails with `unresolved_variables` — `testrun run` has no `--variables` flag, so fill the pool file (`.testmuai/variables/*.json`) or the member's own `variables:` frontmatter. +If **any** member fails preflight, the plan is invalid: nothing runs, exit `2`. Suggest `--dry-run` to preview the plan cheaply before a big run. *(0.8.15+)* Preflight also checks variables. A `{{name}}` with no value is a warning: the member stays in the plan (`valid` does not look at variables), and the run types the name as written unless a step sets it first. The check covers every step of the member, replayed ones included. `testrun run` has no `--variables` flag, so fill the pool file (`.testmuai/variables/*.json`) or the member's own `variables:` frontmatter first. ## Remote: the suite as one HyperExecute job (`--remote`) -`kane-cli testrun run --remote` ships the cwd as the job payload, provisions a grid runtime on a **HyperExecute macOS runner** — Chrome for web members, a **virtual Android emulator or iOS simulator** for mobile members — runs every member there as its own headless `testmd run`, and brings the recordings (`output-/`) and the sealed evidence pack back into the project. It works **from any machine** with nothing local but Node and the plugin (no Chrome needed); the account needs a HyperExecute plan with macOS runners and the plugin (`kane-cli plugin install remote-execution`; check with `kane-cli plugin doctor remote-execution`). Auth is a LambdaTest username + access key — an OAuth profile is exchanged automatically. Not the same as `--ws-endpoint`, which attaches a remote browser to a run that still executes locally. +`kane-cli testrun run --remote < /dev/null` ships the cwd as the job payload, provisions a grid runtime on a **HyperExecute macOS runner** — Chrome for web members, a **virtual Android emulator or iOS simulator** for mobile members — runs every member there as its own headless `testmd run`, and brings the recordings (`output-/`) and the sealed evidence pack back into the project. It works **from any machine** with nothing local but Node and the plugin (no Chrome needed); the account needs a HyperExecute plan with macOS runners and the plugin (`kane-cli plugin install remote-execution`; check with `kane-cli plugin doctor remote-execution`). Auth is a LambdaTest username + access key — an OAuth profile is exchanged automatically. Not the same as `--ws-endpoint`, which attaches a remote browser to a run that still executes locally. ```bash -kane-cli testrun run --tags smoke --remote --dry-run # web suite: validate, dispatch nothing -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 # Android suite on the grid -kane-cli testrun run tests/ios/ --remote --device-name "iPhone 15" --os-version 17.5 # iOS suite on the grid +kane-cli testrun run --tags smoke --remote --dry-run < /dev/null # web suite: validate, dispatch nothing +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 < /dev/null # Android suite on the grid +kane-cli testrun run --match '^tests/ios/' --remote --device-name "iPhone 15" --os-version 17.5 < /dev/null # iOS suite on the grid ``` - **Always `--dry-run` first.** It runs the normal preflight plus the **remote preflight** and resolves the device against the grid catalog (`kane-cli devices list --target emulator|simulator --remote --agent`) without creating a job. @@ -87,7 +87,7 @@ All typed; stdout; one JSON object per line. **Local completion: `testrun_done`. | `type` | Payload | Notes | |---|---|---| -| `testrun_plan` | `members: [{path, test_id?, tags, failure?}]`, `valid`, `parallel`, `parallel_clamped?` | If `valid: false`, treat as immediate failure — report each member's `failure` reason and stop expecting more events. *(0.8.12+)* `failure: "unresolved_variables"` means a member references a `{{name}}` with no value; one `error` event with `code: "unresolved_variables"` follows the plan (schema in `references/parsing.md`) and lists every such name across members — surface it, do not retry. | +| `testrun_plan` | `members: [{path, test_id?, tags, failure?, unresolved?}]`, `valid`, `parallel`, `parallel_clamped?` | If `valid: false`, treat as immediate failure — report each member's `failure` reason; a `warning` may still follow before the process exits. *(0.8.15+)* `unresolved[]` lists the member's `{{name}}`s with no value; `valid` does not look at it. One `warning` event with `code: "unresolved_variables"` follows the plan (schema in `references/parsing.md`) and the run goes ahead. | | `testrun_start` | `execution_id`, `members` (paths), `parallel` | | | `testrun_member_start` | `path`, `test_id?`, *(0.8.17+)* `session_id`, `log_path` | A saved test started. `log_path` is the absolute path of that test's own event log (see **Each test's own log** below). | | `testrun_member_end` | `path`, `test_id?`, `status`, `duration_s`, *(0.8.17+)* `session_id`, `log_path`, `failure?: {message, step_index?}` | `status` ∈ `passed \| failed \| broken \| interrupted`. `failure` is present when the test did not pass: use it for the "where" and "why" of the failed-tests table. | @@ -129,6 +129,7 @@ for each line: if type === "testrun_done" → capture suite outcome; remote runs keep reading if type === "remote_done" → capture remote status, exit and sessions_path if type === "testrun_plan" && !valid → report offenders, expect exit 2 + if type === "warning" && code === "unresolved_variables" → name the variables with no value (references/parsing.md); the run continues; ignore other warning codes if type === "testrun_member_end" → note per-member outcome if type === "testrun_summary" → capture totals for the rollup else → informational; narrate sparingly @@ -164,7 +165,7 @@ Local suites containing any mobile member require `--parallel 1`; larger values Healing is enabled by default (three shrinking replay windows, then re-authoring of authorable steps). `--no-adaptive-heal` disables it. Retired `--retry`/`--retry-count` only print a notice and have no effect. Replay-only recorded steps retain their recordings even during healing. -NDJSON selection uses stdin, not stdout: run `kane-cli testrun run < /dev/null` for automation launched from a terminal. Dry-run validates a plan, not runtime authentication or browser/device readiness. Always observe process exit, including paths without a normal completion event. +NDJSON selection uses stdin, not stdout: every `kane-cli testrun run` line ends in `< /dev/null` (bash and zsh on macOS, Linux and Git Bash; `< NUL` in cmd.exe; from PowerShell run it through cmd: `cmd /c "kane-cli testrun run … < NUL"`). Check the first stdout line: on 0.8.17+ it is `{"type":"stream_start"…}`. A prose plan there instead means the CLI is in terminal mode: it prints the human view and, after a real run, opens an evidence table that waits until `q` or Esc, so the process never exits on its own. Stop it and rerun with the redirect. An empty stdout with a message on stderr is a usage error: read it. Dry-run validates a plan, not runtime authentication or browser/device readiness. Always observe process exit, including paths without a normal completion event. ### Remote behavior still requiring verification diff --git a/README.md b/README.md index cd5f6ad..225ddfd 100644 --- a/README.md +++ b/README.md @@ -91,7 +91,7 @@ kane-cli launches your locally installed Google Chrome (stable channel) via the kane-cli can also run tests against an iOS Simulator or Android Emulator. It is off by default, so your web runs are unaffected. - **Locally** — macOS Apple Silicon (arm64) only. Install the platform tooling you already use: **Xcode** (for iOS), or **Android Studio** with one `arm64-v8a` AVD (for Android); sign in and install kane-cli's managed test tooling: `kane-cli login && kane-cli doctor --target simulator --install`; then `kane-cli run "" --target simulator --app ./MyApp.zip`. -- **On the cloud grid** — from Linux, Windows, or any Mac, with no mobile tooling: `kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14` runs a saved mobile suite on a HyperExecute emulator or simulator (needs a LambdaTest plan with HyperExecute macOS runners and `kane-cli plugin install remote-execution`). +- **On the cloud grid** — from Linux, Windows, or any Mac, with no mobile tooling: `kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14` runs a saved mobile suite on a HyperExecute emulator or simulator (needs a LambdaTest plan with HyperExecute macOS runners and `kane-cli plugin install remote-execution`). > Full setup and prerequisites: [Mobile testing](docs/user-guide/mobile/overview.md) · [Remote runs on the cloud grid](docs/user-guide/remote-execution.md). @@ -201,7 +201,7 @@ kane-cli run " and …'>" --ag Three rules: -1. Use `--agent` for `run`, `testmd run`, and `generate`. `testrun run` has no `--agent`: it emits NDJSON when **stdin** is not a TTY (use `< /dev/null` for terminal automation). Assurance conversational commands use `--mode agent`. Parse command-specific completion events and process exit. +1. Use `--agent` for `run` and `testmd run`. `testrun run` has no `--agent`: it emits NDJSON only when **stdin** is not a TTY, so every `testrun run` line an agent writes ends in `< /dev/null` (`< NUL` in cmd.exe; from PowerShell, `cmd /c "… < NUL"`); the first stdout line is `stream_start` (0.8.17+); a prose plan there means terminal mode, which after a real run waits on an evidence table until `q` or Esc. Assurance conversational commands use `--mode agent`. Parse command-specific completion events and process exit. 2. **Use the "store as" pattern** for extraction. `"go to example.com, store the page title as 'page_title'"` — never `"read the page title"`. 3. **One objective = one task.** If a flow has more than ~15 steps, split it into multiple `kane-cli run` calls and run them in parallel. diff --git a/docs/user-guide/assurance/automation.md b/docs/user-guide/assurance/automation.md index 0f5f217..9c7bd48 100644 --- a/docs/user-guide/assurance/automation.md +++ b/docs/user-guide/assurance/automation.md @@ -62,8 +62,8 @@ With `--mode agent`, stdout speaks a versioned NDJSON vocabulary — envelope `{ | `held` / `update_held` *(0.7.1)* | items were **held** for your review instead of committed (`source_id`, `count`, `reason` / `count`, `targets[]`) — the `--trust hold` and degraded-detection paths | | `commit` | what landed: counts + `minted[]` (`cid` + `logical_id`); extract adds `proposal_id` | | `receipt` | per-phase commit receipt (design; extract emits one at its commits too) — `commit_n`, `phase`, `committed[]`, `warnings[]`, a human-readable `next` hint, and (design only) `parity` | -| `variables_declared` *(0.8.12)* | design: the stubs a phase commit wrote — `file` (the pool file) + `variables[]` (`name`, `description`, `secret`). Only names that existed in no variable file; fires after the `commit` that minted the tests | -| `variables_summary` *(0.8.12)* | design, at the end of the run: every name still needing a value — same shape as `variables_declared` | +| `variables_declared` *(0.8.12)* | design: the stubs a phase commit wrote — `file` (the pool file) + `variables[]` (`name`, `description`, `secret`). Only names that existed in no variable file; fires in the tests phase, just before its `commit` event | +| `variables_summary` *(0.8.12)* | design, after `session_complete` on a clean completion: every name this run declared, same shape as `variables_declared`. A paused run does not send it | | `message_sent` | your `--message` was delivered: `sid`, `chars` | | `panel_resolved` *(0.7.1)* | a pending question was answered by a `--answer` flag: `id`, `by`, `via` | | `ask_deferred` *(0.7.1)* | a pending question batch was set aside because `--with-source` landed a new source first: `source_id`, `cid`, `questions` (count) | diff --git a/docs/user-guide/assurance/design.md b/docs/user-guide/assurance/design.md index 09a873f..49184cf 100644 --- a/docs/user-guide/assurance/design.md +++ b/docs/user-guide/assurance/design.md @@ -67,7 +67,7 @@ Confirm count check: 1 (equals) — the stated promise: after adding one in-stoc Three things to notice: - **`@verifies` tags** bind each assert step to the acceptance criteria it proves. This is the link [`kane-cli cover`](./coverage.md) measures against — captured at authoring time, permanent, auditable. -- **`{{variables}}`** appear wherever the requirements didn't pin a value (the store URL, a known in-stock product). Design reads every `*.json` in your variable directories first and reuses a name you already have; each name it has to invent is **declared** — an empty stub lands in `.testmuai/variables/assurance.json`, never a value, and a key that already exists anywhere is never overwritten. You hear about it twice: a `⚒ declared 2 new variables — store_url · in_stock_product (values needed)` line as the phase commits, and a closing `2 variable(s) need values — see .testmuai/variables/assurance.json`. Fill the values before authoring — a run refuses to start on a name with no value. See [Variables & context](../variables-and-context.md#variables-declared-by-kane-cli-design-tests). +- **`{{variables}}`** appear wherever the requirements didn't pin a value (the store URL, a known in-stock product). Design reads every `*.json` in your variable directories first and reuses a name you already have; each name it has to invent is **declared** — an empty stub lands in `.testmuai/variables/assurance.json`, never a value, and a key that already exists anywhere is never overwritten. You hear about it twice: a `⚒ declared 2 new variables — store_url · in_stock_product (values needed)` line as the phase commits, and a closing `2 variable(s) need values — see .testmuai/variables/assurance.json`. Fill the values before authoring — a run warns on a name with no value and types it as written. An agent fills the values it already has from your documents or your messages and asks you for the rest; a secret value goes straight into the file, never through chat. See [Variables & context](../variables-and-context.md#variables-declared-by-kane-cli-design-tests). - The `assurance:` frontmatter links the file to its design entry in the graph, so coverage lookups are exact even after the file moves. Alongside the tests, the run records the ACs and scenarios themselves, the wiring between them, **gap nodes** with full context for everything it could not resolve, and a rationale sidecar per test (under `.context/design/rationale/`) that `design explain` replays. You'll also see **warnings** at commit time — for example when a test's `@verifies` list claims more criteria than its written check actually asserts. diff --git a/docs/user-guide/cicd.md b/docs/user-guide/cicd.md index f6d3196..3860a88 100644 --- a/docs/user-guide/cicd.md +++ b/docs/user-guide/cicd.md @@ -10,7 +10,7 @@ These patterns apply to every CI system; the recipes below differ only in how th - Always pass `--timeout `. A hung run cannot be allowed to block the pipeline. - Authenticate with `--username` and `--access-key` from CI secrets. Do not call `kane-cli login` in CI — that flow opens a browser for OAuth and will not work on a runner. - Provide a start URL. A CI run has no interactive prompt, so if neither the objective names a site, the `--url` flag is passed, nor a `default_url` is configured, the run fails fast instead of waiting for input. The simplest options are to start the objective with the site ("Go to https://… and …") or pass `--url `. To deliberately start from the browser's current page instead, add `--allow-missing-url`. See [Default start URL](./configuration.md#default-start-url). -- Load test data with `--variables-file `. Check the file into your repo (without secret values), or generate it before the step. A `{{name}}` with no value fails the job with exit `2` **before** any browser starts, with a receipt naming the variable — so fill values from CI secrets in that step, not later. +- Load test data with `--variables-file `. Check the file into your repo (without secret values), or generate it before the step. A `{{name}}` with no value does not change the exit code by itself: kane-cli warns before the browser starts and types the name as written, and a step that then fails on the literal text fails the job as usual. Fill values from CI secrets in that step. The warning is a `warning` line on stdout with `code: "unresolved_variables"`; fail the step on it if a missing value should block the pipeline. - **Project and folder are optional.** If you want uploads filed under a specific Test Manager project/folder, pre-configure with `kane-cli config project ` / `kane-cli config folder ` (use `kane-cli projects list` / `kane-cli folders list` to find the IDs). If you skip this, kane-cli auto-defaults a project/folder on first run and reports which one it picked — see [test-manager-integration.md](./test-manager-integration.md). - Check the exit code. The mapping is documented in [running tests](./running-tests.md#exit-codes); the short form is `0` passed, `1` failed, `2` error, `3` timeout or cancellation. - Run whole suites with one command. If your repo has committed `_test.md` tests, prefer one `testrun` invocation over a shell loop: @@ -27,12 +27,12 @@ These patterns apply to every CI system; the recipes below differ only in how th kane-cli plugin install remote-execution # a web suite on 4 grid runners - kane-cli testrun run tests/web/ --remote --parallel 4 \ + kane-cli testrun run --match '^tests/web/' --remote --parallel 4 \ --username "$LT_USERNAME" --access-key "$LT_ACCESS_KEY" \ --on-failure fail-fast # a mobile suite, from a Linux runner - kane-cli testrun run tests/app/ --remote \ + kane-cli testrun run --match '^tests/app/' --remote \ --device-name "Pixel 7" --os-version 14 \ --username "$LT_USERNAME" --access-key "$LT_ACCESS_KEY" \ --on-failure fail-fast diff --git a/docs/user-guide/mobile/overview.md b/docs/user-guide/mobile/overview.md index a75cea0..dc03700 100644 --- a/docs/user-guide/mobile/overview.md +++ b/docs/user-guide/mobile/overview.md @@ -59,7 +59,7 @@ kane-cli run "Add the first item to the cart" --app ./builds/app-debug.apk # a saved test, or a whole suite of them kane-cli testmd run tests/checkout_test.md -kane-cli testrun run tests/app/ --device-name "Pixel 7 API 35" --os-version 15 +kane-cli testrun run --match '^tests/app/' --device-name "Pixel 7 API 35" --os-version 15 ``` Pick a device with `--device-name` and `--os-version` as `kane-cli devices list --target emulator|simulator` prints them, or save defaults with `kane-cli config set-device-name` / `set-os-version`. In the interactive TUI, switch targets with `/mobile` and `/desktop`. For the full flag list and the app formats each target accepts, see [Running tests](../running-tests.md). @@ -71,8 +71,8 @@ Skip the local setup entirely: `kane-cli testrun run --remote` sends your mobile ```bash kane-cli plugin install remote-execution # once kane-cli devices list --target emulator --remote # what the grid can provision -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 --dry-run -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 --dry-run +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 ``` Three things differ from a local run: the device comes from the **grid catalog** (`devices list … --remote`), one job runs **one platform** (emulator members on one Android version; simulator members on one HyperExecute pool), and a **local build** is uploaded from your machine before dispatch and handed to the grid as an `APP…` id. The details — prerequisites, the app rules, and what one job can hold — are in [Remote runs on the cloud grid](../remote-execution.md). diff --git a/docs/user-guide/remote-execution.md b/docs/user-guide/remote-execution.md index f0c4a10..2c897eb 100644 --- a/docs/user-guide/remote-execution.md +++ b/docs/user-guide/remote-execution.md @@ -10,8 +10,8 @@ Remote runs cover both kinds of test: ```bash kane-cli plugin install remote-execution # once kane-cli testrun run --tags smoke --remote --parallel 4 # a web suite on 4 grid runners -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 # an Android suite -kane-cli testrun run tests/ios/ --remote --device-name "iPhone 15" --os-version 17.5 +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 # an Android suite +kane-cli testrun run --match '^tests/ios/' --remote --device-name "iPhone 15" --os-version 17.5 ``` > `--remote` is not the same as `--ws-endpoint` / `--cdp-endpoint`. Those attach a **remote browser** to a run that still executes on your machine (`kane-cli run`, `kane-cli testmd run`). `--remote` moves the **whole suite** to the grid: kane-cli itself runs there, and nothing but Node and the plugin is needed locally. @@ -67,7 +67,7 @@ A web suite needs nothing beyond the prerequisites: the grid runner has Chrome, ```bash kane-cli plugin install remote-execution -kane-cli testrun run tests/web/ --remote --parallel 4 +kane-cli testrun run --match '^tests/web/' --remote --parallel 4 ``` What you see back is a normal `testrun` summary; the only extra lines are the dispatch and the job link. A member that authors on the grid comes back with its `output-/` recordings, so the next run — local or remote — replays them. Commit those recordings as you would after a local run. @@ -154,11 +154,11 @@ npm install -g @testmuai/kane-cli kane-cli plugin install remote-execution # a web suite -kane-cli testrun run tests/web/ --remote --parallel 4 \ +kane-cli testrun run --match '^tests/web/' --remote --parallel 4 \ --username "$LT_USERNAME" --access-key "$LT_ACCESS_KEY" # a mobile suite -kane-cli testrun run tests/app/ --remote \ +kane-cli testrun run --match '^tests/app/' --remote \ --device-name "Pixel 7" --os-version 14 \ --username "$LT_USERNAME" --access-key "$LT_ACCESS_KEY" ``` diff --git a/docs/user-guide/running-tests.md b/docs/user-guide/running-tests.md index a777fdd..2ce6a5a 100644 --- a/docs/user-guide/running-tests.md +++ b/docs/user-guide/running-tests.md @@ -195,13 +195,13 @@ The customer-facing flags accepted by `kane-cli run`: ### Unresolved variables -Every `{{name}}` in the objective must have a value before the run starts. A name with no value stops the run there — no browser, no session — with exit code `2` and a receipt naming each variable and what it needs. There is no flag to bypass it: fill the value or remove the reference. With `--agent`, the refusal is a single typed event: +kane-cli checks every `{{name}}` in the objective before the run starts. A name with no value is a warning: the run goes ahead and types the name as written, unless a step sets it first. On a TTY the warning is the receipt shown in [Variables and context](./variables-and-context.md#before-a-run-unresolved-variables); with `--agent` it is a single typed event before the first progress frame: ```json -{"type":"error","code":"unresolved_variables","message":"2 variable(s) have no value — nothing was dispatched","suggested_file":".testmuai/variables/variables.json","variables":[{"name":"checkout_url","reason":"not_declared","used_by":[{"file":"objective","step":1}]},{"name":"login_password","reason":"value_missing","file":".testmuai/variables/variables.json","used_by":[{"file":"objective","step":1}]}]} +{"type":"warning","code":"unresolved_variables","message":"2 variable(s) have no value — typed as written unless a step sets them first","suggested_file":".testmuai/variables/variables.json","variables":[{"name":"checkout_url","reason":"not_declared","used_by":[{"file":"objective","step":1}]},{"name":"login_password","reason":"value_missing","file":".testmuai/variables/variables.json","used_by":[{"file":"objective","step":1}]}]} ``` -`reason` is `value_missing` (the key exists in `file`, with no value) or `not_declared` (the key is in no file; `suggested_file` is where to add it). Names an earlier step stores, and the `{{smart.*}}` / `{{environment.*}}` / `{{secrets.*}}` / `{{totp.*}}` namespaces, are never checked. +`reason` is `value_missing` (the key exists in `file`, with no value), `not_declared` (the key is in no file; `suggested_file` is where to add it), or `not_a_dataset_column` (a `${x}` that is no column of the dataset row the run was given). The warning does not change the exit code. An explicit `{{global.*}}` reference (resolved from Test Manager at run time), names an earlier step stores, and the `{{smart.*}}` / `{{environment.*}}` / `{{secrets.*}}` / `{{totp.*}}` namespaces are never checked. For variables and context file behavior, see [./variables-and-context.md](./variables-and-context.md). For code export and the run mode toggle, see [./configuration.md](./configuration.md). diff --git a/docs/user-guide/testmd/running.md b/docs/user-guide/testmd/running.md index c3e3e77..703f67a 100644 --- a/docs/user-guide/testmd/running.md +++ b/docs/user-guide/testmd/running.md @@ -69,7 +69,7 @@ Every flag accepted by `kane-cli testmd run`: Most flags have a frontmatter counterpart with the same name (with underscores). Where both are set, the CLI flag wins — except for `variables`, which the file owns; see [overview.md](./overview.md#variables). -Before any step runs, every `{{name}}` the authored steps reference must have a value — from the file's own `variables:` frontmatter, a variable file, `--variables-file` or `--variables`. A name with no value stops the run before the browser starts, with exit `2` and a receipt naming the variable, the file waiting for its value, and the step. Replayed steps are not checked. See [Before a run](../variables-and-context.md#before-a-run-unresolved-variables). +Before any step runs, every `{{name}}` the authored steps reference is checked against the file's own `variables:` frontmatter, the variable files, `--variables-file` and `--variables`. A name with no value is a warning: the receipt names the variable, the file waiting for its value, and the step, and the run goes ahead, typing the name as written unless a step sets it first. Replayed steps are not checked. See [Before a run](../variables-and-context.md#before-a-run-unresolved-variables). ## How a run works @@ -249,7 +249,7 @@ kane-cli testmd run ./tests/checkout_test.md \ In a non-interactive run (stdin is not a TTY), there is no one to answer an interactive `ask_user` prompt, so kane-cli disables it: a step that would otherwise wait for input fails cleanly instead of blocking forever. Write test steps that do not depend on mid-run prompts when running in CI. -If the runner cannot have Chrome at all, run the tests as a suite on the cloud grid instead: `kane-cli testrun run tests/ --remote` executes every `testmd run` on a HyperExecute macOS runner and returns the recordings and evidence pack. See [Remote runs on the cloud grid](../remote-execution.md). +If the runner cannot have Chrome at all, run the tests as a suite on the cloud grid instead: `kane-cli testrun run --match '^tests/' --remote` executes every `testmd run` on a HyperExecute macOS runner and returns the recordings and evidence pack. See [Remote runs on the cloud grid](../remote-execution.md). Capture exit code in a shell script: diff --git a/docs/user-guide/testrun.md b/docs/user-guide/testrun.md index 69c8d17..56f05ba 100644 --- a/docs/user-guide/testrun.md +++ b/docs/user-guide/testrun.md @@ -14,7 +14,7 @@ Use `testrun` when you have a suite of committed tests to run together — night Members come either from explicit paths (each must end in `_test.md`) or, when no paths are given, from a recursive walk of the current directory. Two filters then apply, in order: -- **`--match `** — keep tests whose project-relative path matches the regex. +- **`--match `** — keep tests whose project-relative path matches the regex. `--match` sees the path as the OS writes it: `tests/app/` on macOS and Linux, `tests\app\` on Windows. Quote the regex with double quotes in cmd.exe; single quotes are literal there. - **`--tags `** — keep tests whose [`tags:` frontmatter](./testmd/overview.md#frontmatter) matches **any** of the given tags (case-insensitive). Repeat the flag or pass a comma-separated list; `--tags smoke,checkout` and `--tags smoke --tags checkout` are equivalent. Duplicates are removed and the final list runs in a stable order. @@ -36,7 +36,6 @@ A member can fail preflight for these reasons: |---|---|---| | `org_mismatch` | Belongs to a different organisation than the rest | Check with `kane-cli testmd status ` | | `project_mismatch` | Belongs to a different project than the rest | Check with `kane-cli testmd status `; run project-by-project | -| `unresolved_variables` *(0.8.12)* | An authored step references a `{{name}}` that has no value in any variable file or in the member's own `variables:` frontmatter | Fill the value in `.testmuai/variables/*.json` (the receipt names the file) or remove the reference. `testrun run` has no `--variables` flag | If any member fails preflight, the plan is invalid and **nothing runs** (exit `2`). The offenders print to stderr: @@ -46,10 +45,10 @@ error: plan invalid — 2 offending test(s): tests/other_project_test.md: project_mismatch ``` -Variable offenders get the full receipt instead of a one-line code — every unresolved name across the members, each with the test files and steps that use it: +A member whose authored steps reference a `{{name}}` with no value passes preflight. kane-cli prints a warning receipt — every such name across the members, each with the test files and steps that use it — and the run types the name as written unless a step sets it first: ``` -✗ 2 variables have no value — nothing was dispatched +warning: 2 variables have no value Not in any variables file other_key b_test.md step 1 @@ -58,10 +57,10 @@ Variable offenders get the full receipt instead of a one-line code — every unr Add them to .testmuai/variables/variables.json If {{name}} is literal page text, write \{{name}} to keep it as-is. - Fill the values and run again. + A step that sets a name first binds it; anything else is typed as written. Fill the values to bind them. ``` -In agent mode (stdin not a TTY) the same information arrives as one `error` event with `code: "unresolved_variables"` right after `testrun_plan` — see [Running tests](./running-tests.md#unresolved-variables) for the shape. +`testrun run` has no `--variables` flag: fill the pool file or the member's own `variables:` frontmatter. In agent mode (stdin not a TTY) the same content arrives as one `warning` event with `code: "unresolved_variables"` right after `testrun_plan`, whose members carry their rows — see [Running tests](./running-tests.md#unresolved-variables) for the shape. ## Mobile members @@ -71,8 +70,8 @@ A `_test.md` with a mobile [`target:`](./testmd/overview.md#mobile-target) (`emu - **On the cloud grid** (`--remote`), the suite runs on a virtual device on a HyperExecute macOS host, so it works **from any machine** — Linux, Windows, or a Mac with no Xcode or Android Studio. Pick the device from `kane-cli devices list --target emulator|simulator --remote`. One grid job runs one platform (emulator members on one Android version, simulator members on one HyperExecute pool), and a member's local build is uploaded from your machine before dispatch and handed to the grid as an `APP…` id. Everything else is in [Remote runs](./remote-execution.md). ```bash -kane-cli testrun run tests/app/ --device-name "Pixel 7 API 35" --os-version 15 # local emulators -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 # the grid +kane-cli testrun run --match '^tests/app/' --device-name "Pixel 7 API 35" --os-version 15 # local emulators +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 # the grid ``` ## Running @@ -110,7 +109,7 @@ Each desktop worker gets its **own isolated Chrome** with a fresh temporary prof kane-cli testrun run --tags smoke --parallel 4 --dry-run ``` -Exit `0` means the plan is valid; exit `2` reports planning refusals. This does not prove runtime readiness: local dry-run returns before runtime authentication and browser/device startup. Remote dry-run also validates the grid catalog but does not upload apps or dispatch a job. A real run can still fail after a valid plan. +Exit `0` means the plan is valid; exit `2` reports planning refusals. This does not prove runtime readiness: local dry-run returns before runtime authentication and browser/device startup. Remote dry-run also validates the grid catalog but does not upload apps or dispatch a job. A real run can still fail after a valid plan. A dry run prints the variable warning too, so you can fill values before the real run. ## Reading results @@ -137,7 +136,7 @@ Add `--remote` and the same selection runs as one **HyperExecute** job: Chrome o kane-cli plugin install remote-execution kane-cli testrun run --tags smoke --remote --dry-run # resolve everything, dispatch nothing kane-cli testrun run --tags smoke --remote --parallel 4 # a web suite on 4 grid runners -kane-cli testrun run tests/ios/ --remote --device-name "iPhone 15" --os-version 17.5 +kane-cli testrun run --match '^tests/ios/' --remote --device-name "iPhone 15" --os-version 17.5 ``` `--dry-run` runs the remote preflight too, so the selection is validated against the grid before any job exists. A web selection and a device selection cannot share one job — run them as two suites. This is different from `--ws-endpoint`, which attaches a remote browser to a run that still executes on your machine. The full guide — prerequisites, how a job runs, web suites, device catalog, app rules, what one job can hold, and the events — is [Remote runs on the cloud grid](./remote-execution.md). diff --git a/docs/user-guide/troubleshooting.md b/docs/user-guide/troubleshooting.md index 71abf0b..f698769 100644 --- a/docs/user-guide/troubleshooting.md +++ b/docs/user-guide/troubleshooting.md @@ -149,11 +149,11 @@ You have three options: ## "Variables not resolving" -Since 0.8.12 a run **refuses to start** when an authored step references a `{{name}}` that has no value — you get a receipt naming each variable, the file that is waiting for its value (or `Not in any variables file`), and the step that uses it; exit code `2`, nothing dispatched. Read the receipt first: it tells you whether to fill an existing key or add a new one, and which file. See [Before a run](./variables-and-context.md#before-a-run-unresolved-variables). +A run **warns before it starts** when an authored step references a `{{name}}` that has no value: the receipt names each variable, the file that is waiting for its value (or `Not in any variables file`), and the step that uses it. The run then goes ahead and types the name as written. Read the warning first: it tells you whether to fill an existing key or add a new one, and which file. See [Before a run](./variables-and-context.md#before-a-run-unresolved-variables). -If a `{{my_var}}` placeholder is nevertheless appearing **literally** in a browser action, one of three things is true: +If a `{{my_var}}` placeholder appears **literally** in a browser action, one of three things is true: -- The step is a **replay** — replayed steps resolve from their tape and from TMS, and a missing value there produces a warning line rather than a refusal. Fill the value and run again. +- The warning above named it and nothing set it — fill the value and run again (a **replay** step resolves from its tape and from TMS instead). - The reference is **escaped** — `\{{my_var}}` is typed as-is on purpose, for pages where the braces are real text. - The variable file is not being loaded at all. Check, in order: diff --git a/docs/user-guide/variables-and-context.md b/docs/user-guide/variables-and-context.md index ffd8797..986792c 100644 --- a/docs/user-guide/variables-and-context.md +++ b/docs/user-guide/variables-and-context.md @@ -133,10 +133,10 @@ Fill `value`. `kane-cli` never overwrites a key that already exists in any of yo ### Before a run: unresolved variables -`kane-cli run`, `kane-cli testmd run` and `kane-cli testrun run` check every `{{name}}` an authored step references **before anything starts** — no browser, no session. A name with no value stops the run there, with exit code `2`, and the receipt says what each one needs: +`kane-cli run`, `kane-cli testmd run` and `kane-cli testrun run` check every `{{name}}` an authored step references **before anything starts**. A name with no value is a warning: the run goes ahead and types the name as written, unless a step sets it first: ``` -✗ 3 variables have no value — nothing was dispatched +warning: 3 variables have no value Waiting for a value in .testmuai/variables/assurance.json storefront_sign_in_url step 1 @@ -148,12 +148,12 @@ Fill `value`. `kane-cli` never overwrites a key that already exists in any of yo Add it to .testmuai/variables/assurance.json If {{name}} is literal page text, write \{{name}} to keep it as-is. - Fill the values and run again. + A step that sets a name first binds it; anything else is typed as written. Fill the values to bind them. ``` -The pool file is named once per group. A test filename appears only when it is not the file you named — an `@import`ed unit and a testrun member both keep theirs, and a `kane-cli run` objective shows no location at all. A name that is in no file gets `Add it to ` when a pool file exists, or the JSON to create when there is none yet. +The pool file is named once per group. A test filename appears only when it is not the file you named — an `@import`ed unit and a testrun member both keep theirs, and a `kane-cli run` objective shows no location at all. A name that is in no file gets `Add it to ` when the pool file exists, or the JSON to create when there is none yet. -There is no flag to bypass the check: fill the value or remove the reference. Never checked: `{{smart.*}}`, `{{environment.*}}`, `{{secrets.*}}` and `{{totp.*}}` (resolved at run time); a name an earlier step stored (`store the price as 'price'`); a test's own frontmatter `variables:`; and replayed steps, which resolve from their tape and from TMS — a replay with a missing value gets a warning line, never a refusal. `${x}` is not a variable reference on this path. In agent mode the refusal is one typed `error` event with `code: "unresolved_variables"` — see [Running tests](./running-tests.md#unresolved-variables). +Never checked: an explicit `{{global.*}}` (it resolves from Test Manager at run time — only a bare `{{name}}` is a pool question), `{{smart.*}}`, `{{environment.*}}`, `{{secrets.*}}` and `{{totp.*}}` (resolved at run time); a name an earlier step stores (`store the price as 'price'`); a name the test's own frontmatter `variables:` declares, even with an empty value; and replayed steps, which resolve from their tape and from TMS (a `testrun run` preflight still lists a replayed step's names, because it checks the file before it knows which steps replay). `${x}` is a dataset column reference, not a pool reference; a `${x}` that is no column of the dataset row the run was given gets its own warning. In agent mode the same content is one `warning` event with `code: "unresolved_variables"` — see [Running tests](./running-tests.md#unresolved-variables). ## Context files diff --git a/skill-installer/skills/SKILL.md b/skill-installer/skills/SKILL.md index 3663572..81c6e03 100644 --- a/skill-installer/skills/SKILL.md +++ b/skill-installer/skills/SKILL.md @@ -1,18 +1,18 @@ --- name: kane-cli -description: Browser automation + AI test authoring via kane-cli - run browser objectives, generate & refine test scenarios/cases from a description, design requirement-linked test suites from a PRD/spec (assurance), parse NDJSON output, inspect logs, save runnable _test.md. Use for any task requiring a real browser (navigate, click, fill forms, test web UI, take screenshots), or to author test cases, quick cases from a description via kane-cli generate; a designed, coverage-accounted suite from requirement documents via the assurance commands. Never write test cases by hand. Also runs mobile app tests - native Android app on a virtual emulator or iOS app on a simulator via --target emulator|simulator (desktop browser stays the default target): locally on macOS Apple Silicon, or on the LambdaTest cloud grid from any machine via kane-cli testrun run --remote. Also shares the assurance store (.context/) with a team through a location (a GitHub repository, an S3-compatible bucket, or a folder) via kane-cli context sync, and resolves the decisions a sync rebase asks. +description: Browser automation + AI test authoring via kane-cli - run browser objectives, design requirement-linked test suites from a PRD/spec or from a description (assurance), parse NDJSON output, inspect logs, save runnable _test.md. Use for any task requiring a real browser (navigate, click, fill forms, test web UI, take screenshots), or to author test cases: a designed, coverage-accounted suite from requirement documents, or from a description the user gives you, via the assurance commands. Never write test cases by hand. Also runs mobile app tests - native Android app on a virtual emulator or iOS app on a simulator via --target emulator|simulator (desktop browser stays the default target): locally on macOS Apple Silicon, or on the LambdaTest cloud grid from any machine via kane-cli testrun run --remote. Also shares the assurance store (.context/) with a team through a location (a GitHub repository, an S3-compatible bucket, or a folder) via kane-cli context sync, and resolves the decisions a sync rebase asks. --- # Kane CLI — Browser Automation Skill -Use `kane-cli` for **any task that requires a real browser**: navigating websites, clicking elements, filling forms, searching, testing web UI, taking screenshots, or verifying deployments. Do NOT use Playwright, Puppeteer, or Selenium directly. Use `--agent` for `run`, `testmd run`, and `generate`. `testrun run` has no `--agent`: it emits NDJSON when **stdin** is not a TTY (use `< /dev/null` for terminal automation). Assurance conversational commands use `--mode agent`. +Use `kane-cli` for **any task that requires a real browser**: navigating websites, clicking elements, filling forms, searching, testing web UI, taking screenshots, or verifying deployments. Do NOT use Playwright, Puppeteer, or Selenium directly. Use `--agent` for `run` and `testmd run`. `testrun run` has no `--agent`: it emits NDJSON only when **stdin** is not a TTY, so every `testrun run` line you write ends in `< /dev/null` (bash and zsh: macOS, Linux, Git Bash; in cmd.exe write `< NUL`; from PowerShell run it through cmd: `cmd /c "kane-cli testrun run … < NUL"`). Check the first stdout line: on 0.8.17+ it is `{"type":"stream_start"…}`. A prose plan there instead means the CLI is in terminal mode: it shows the human view and, after a real run, opens an evidence table that waits until `q` or Esc, so the process never exits on its own; stop it and rerun with the redirect. An empty stdout with a message on stderr is a usage error: read it. Assurance conversational commands use `--mode agent`. -**Authoring test cases or scenarios?** Never write them by hand — kane-cli has two authoring pipelines, and the routing matters: +**Authoring test cases or scenarios?** Never write them by hand: every test case comes from the **assurance** commands — Read `references/assurance.md` first. -- The user describes what to test in a sentence or two, or wants quick scenario/case ideas → `kane-cli generate` (§6). -- The user has **requirement documents** (a PRD, a spec, acceptance notes) and wants a designed suite, requirement-linked coverage, or "what exactly is covered?" answers → the **assurance** commands — Read `references/assurance.md` first. +- The user has **requirement documents** (a PRD, a spec, acceptance notes) → ingest them, then design. +- The user only **describes** what to test, in chat → write their description, in their words, to a requirements file, ingest that file, then design (§6). The description is the requirement. -Don't draft test cases in chat or scratch files: both pipelines produce structured, refinable, runnable `_test.md` output. +Don't draft test cases in chat or scratch files: design produces structured, refinable, runnable `_test.md` output. --- @@ -50,7 +50,7 @@ On Windows PowerShell: `$env:KANE_CLI_USER_AGENT=''; kane-cli run **Keeping runs.** When the person's saved purpose is `suite` or `ask`, add `--name ` to every one-off `run`. A named run is recorded as a `_test.md` while it runs, so keeping it afterwards costs nothing, and a run launched without a name cannot be kept without running again. With `one-off`, leave the flag out. Details: `references/first-run.md` §4. -Bash blocks until kane-cli exits, then hands you the complete stdout. Parse it, summarize what happened, and present the result card. Wait for process completion on `testmd run` and `generate` too, but parse their own completion events: `test_md_done` and `generate_done`, respectively. An intermediate `run_end` does not finish a saved test. +Bash blocks until kane-cli exits, then hands you the complete stdout. Parse it, summarize what happened, and present the result card. Wait for process completion on `testmd run` too, but parse its own completion event, `test_md_done`. An intermediate `run_end` does not finish a saved test. Set a generous timeout (up to 600000ms) since browser runs can take a while. @@ -76,7 +76,8 @@ Progress events have `step`/`status`/`remark` fields and **no `type` field**. |------|-------------|-----| | **Failures** | Any step with `status: "failed"` | `Step failed: ` | | **Flow changes** | `bifurcation`, `child_agent_start`, `child_agent_end` | Plain-language one-liner (e.g. "The agent split the objective into 2 sub-tasks") | -| **Errors** | `error` typed events | `Error: `. The exception is `code: "unresolved_variables"`, which is a pre-run refusal, not a failure: see §3 **Unresolved variables** | +| **Errors** | `error` typed events | `Error: ` | +| **Unresolved variables** | `warning` with `code: "unresolved_variables"` | Before any progress: name the variables with no value and where a value goes (§3 **Unresolved variables**). The run continued | | **Overall progress** | All passing steps | One summary line: ` steps completed: <2–4 key actions from remarks>` | #### What to skip @@ -159,9 +160,9 @@ When the user's request involves a browser — or writing test cases: **What does the user want?** - A single one-shot browser task → build a `kane-cli run --agent` command (§3 + §4) - A test they want to save / re-run / commit → Read `references/testmd.md` first, then use `kane-cli testmd` -- Run a suite of saved tests (several `_test.md` at once) → Read `references/testrun.md` first, then use `kane-cli testrun run` -- Need test cases or scenarios from a short description — because the user asked, or because the task needs them (no browser) → **don't hand-write them**; Read `references/generate.md` first, then use `kane-cli generate` (§6) -- Has requirement documents (PRD/spec) and wants a designed suite, coverage accounting, or suite upkeep → Read `references/assurance.md` first — the assurance commands (`context`/`design`/`cover`/`maintain reconcile`, kane-cli 0.6.1+; several features need newer releases, up to 0.8.14+ — the reference marks each), NOT `generate` +- Run a suite of saved tests (several `_test.md` at once) → Read `references/testrun.md` first, then use `kane-cli testrun run … < /dev/null` +- Need test cases from a description the user gives in chat, with no document → **don't hand-write them**; write the description to a requirements file and design from it (§6) +- Has requirement documents (PRD/spec) and wants a designed suite, coverage accounting, or suite upkeep → Read `references/assurance.md` first — the assurance commands (`context`/`design`/`cover`/`maintain reconcile`, kane-cli 0.6.1+; several features need newer releases, up to 0.8.14+ — the reference marks each) - A **designed** test (its `_test.md` carries an `assurance:` block) failed a run → Read `references/assurance.md` §6.1 **before touching the file** — an edit is adopted as the next design version on the next run; classify the failure first (app bug → report it; requirement changed → reconcile; wording → redesign through the CLI; capability missing → stop), and edit a step of an existing designed test by hand only after reading `references/objectives-cookbook.md` - Share the context store with a team, join a teammate's, keep two stores level, or resolve a sync conflict → Read `references/context-sync.md` first — `kane-cli context sync`, `kane-cli context push`, `kane-cli context pull` and `kane-cli context clone` (kane-cli 0.8.14+); the store is shared through a location, never by copying or git-merging `.context/` - Multiple independent browser tasks → Read `references/parallel.md` first @@ -175,7 +176,7 @@ When the user's request involves a browser — or writing test cases: - The person wants to watch runs live, or asks about the status line → Read `references/live-strip.md` (Claude Code only) - You need the full NDJSON event schema (rare — §5's summary covers 90% of cases) → Read `references/parsing.md` - Compare / evaluate / justify kane-cli against another tool or approach (cost, tokens, effort, ROI) → Read `references/fair-evaluation.md` first — comparisons are only honest like-for-like across the test lifecycle -- **Mobile**: drive a native app on a virtual Android emulator or iOS simulator instead of the browser → Read `references/mobile.md` first. Desktop (the browser) stays the **default** target; mobile is opt-in via `--target emulator|simulator` and always drives an app you provide (`--app `), never a URL. **Local** mobile runs (`run`, `testmd run`, `testrun run`) need macOS Apple Silicon. **From any other machine** (Linux, Windows, Intel Mac, a Mac without Xcode/Android Studio), run saved mobile `_test.md` files on the cloud grid with `kane-cli testrun run --remote --device-name "" --os-version ` — the grid boots the emulator/simulator on a HyperExecute macOS host (the account needs a HyperExecute plan with macOS runners). Never tell a non-Mac user mobile is impossible: point them at `--remote`. +- **Mobile**: drive a native app on a virtual Android emulator or iOS simulator instead of the browser → Read `references/mobile.md` first. Desktop (the browser) stays the **default** target; mobile is opt-in via `--target emulator|simulator` and always drives an app you provide (`--app `), never a URL. **Local** mobile runs (`run`, `testmd run`, `testrun run`) need macOS Apple Silicon. **From any other machine** (Linux, Windows, Intel Mac, a Mac without Xcode/Android Studio), run saved mobile `_test.md` files on the cloud grid with `kane-cli testrun run --remote --device-name "" --os-version < /dev/null` — the grid boots the emulator/simulator on a HyperExecute macOS host (the account needs a HyperExecute plan with macOS runners). Never tell a non-Mac user mobile is impossible: point them at `--remote`. **Every run, always:** follow §1 above. @@ -187,7 +188,7 @@ When the user's request involves a browser — or writing test cases: kane-cli run "" --agent [options] ``` -> The `run` subcommand is **mandatory**. `kane-cli ""` (no `run`) does **not** work — unknown first tokens exit `2` with a "did you mean" suggestion. Same rule applies to `kane-cli testmd run …` and `kane-cli generate …`. +> The `run` subcommand is **mandatory**. `kane-cli ""` (no `run`) does **not** work — unknown first tokens exit `2` with a "did you mean" suggestion. Same rule applies to `kane-cli testmd run …`. `--agent` is mandatory — it switches stdout to NDJSON. Most-used flags: @@ -209,7 +210,14 @@ Other flags (`--global-context`, `--local-context`, `--cdp-endpoint`, `--allow-m **Exit codes:** `0` passed · `1` failed · `2` auth/infra error · `3` timeout/cancelled. -**Unresolved variables (0.8.12+):** every `{{name}}` in the objective must have a value before the run starts. If one does not, the run **refuses before anything launches** — exit `2`, and with `--agent` a single `{"type":"error","code":"unresolved_variables", ...}` event carrying `variables[]` (`name`, `reason`: `value_missing` = the key exists in `file` with no value · `not_declared` = the key is in no file, add it to `suggested_file`; `used_by[]`). **This is terminal — do not retry the same command.** Either ask the user for the values, or write them yourself (`--variables '{"name":{"value":"…"}}'`, or `{"name":{"value":""}}` stubs into `suggested_file` for the user to fill), then run again. There is no bypass flag. Never checked: `{{smart.*}}`/`{{environment.*}}`/`{{secrets.*}}`/`{{totp.*}}`, and names an earlier step stores (`store … as 'x'`). Numbers in a variable file count as values (loaded as strings); booleans do not. +**Unresolved variables (0.8.15+):** kane-cli checks every `{{name}}` in the objective before the run starts. A name with no value is a warning: the run goes ahead and types the name as written, unless a step sets it first. With `--agent` the warning is one `{"type":"warning","code":"unresolved_variables", ...}` event before the first progress frame. It carries `suggested_file` and `variables[]`: `name`; `reason` (`value_missing` = the key exists in `file` with no value · `not_declared` = the key is in no file, add it to `suggested_file`); `used_by[]`. Never checked: an explicit `{{global.*}}` (it resolves from Test Manager at run time), `{{smart.*}}`, `{{environment.*}}`, `{{secrets.*}}`, `{{totp.*}}`, and names an earlier step stores (`store … as 'x'`). Numbers in a variable file count as values (loaded as strings); booleans do not. + +**Fill the variables before any run.** Do this before `run`, `testmd run` and `testrun run`, and always before `--remote`, which books a grid job. A missing value is asked for before the run; the first-run rule that nothing is asked before the first result does not cover it. + +1. Collect the names with no value: the `warning`, a dry-run plan's `unresolved[]` rows, or a design run's `variables_declared` and `variables_summary` rows. Design rows list what design declared, not every missing value, and a dry run's `valid: true` says nothing about values: the dry run's `warning` is the check. +2. A value you already have, because the user said it or the requirement document states it, you write into that key in the file the event names (for `run`, `--variables '{"name":{"value":"…"}}'` also works), and you tell the user what you filled and where. +3. The rest you ask for once, in one message, using each variable's `description` when design gave one. Plain values (a URL, an email, a user name) the user gives you here or adds to the file, their choice. Secrets, which are any row with `secret: true`, any description that names a credential, and any name containing `password`, `secret`, `token` or `key`, the user fills in the file and tells you when done; you never ask for the value in chat, never echo it, and report names and file paths only. +4. Then a fresh `--dry-run` of the exact selection: a valid plan with no `warning` is the check. A member whose rows are all filled may run while the others wait. Never fill a placeholder just to silence the check; a throwaway value is right only when the description asks for one, such as a deliberately wrong password. A frontmatter declaration with an empty value also silences the check, so look at the values a step relies on, not only at the warning. A name still empty is typed into the page as written; if a run then failed at that step, say so. ### Examples @@ -298,7 +306,7 @@ Action → extraction → assertion in one objective: Stdout is NDJSON, one event per line. On kane-cli 0.8.17+ every line also carries `v` (contract version, `1`) and `ts` (when it was emitted), and the first line is `{"type":"stream_start","cli_version":…,"surface":"run"|"testmd"|"testrun"}`. Ignore fields and event types you do not know: new ones can appear in any release. There are two shapes: - **Progress events** (most events) have `step` (1-based), `status` (`running` at start, `done`/`failed` at completion), `remark` — and **no `type` field**. -- **Typed events** have a `type` field: `project_folder_auto_defaulted` (run-startup gate, fires before any progress when no project/folder is configured), `bifurcation`, `child_agent_start`, `child_agent_end`, `ask_user`, `error` (an `error` with `code: "unresolved_variables"` is a pre-run refusal and the **only** line — no `run_end` follows; handle per §3), and finally `run_end`. +- **Typed events** have a `type` field: `project_folder_auto_defaulted` (run-startup gate, fires before any progress when no project/folder is configured), `bifurcation`, `child_agent_start`, `child_agent_end`, `ask_user`, `error`, `warning` (pre-run `code: "unresolved_variables"` — the run continues; handle per §3), and finally `run_end`. Parsing strategy: @@ -313,51 +321,21 @@ For one-shot `run`, build post-run logic on `run_end` and process exit. Saved te For full event schemas (`bifurcation` flow fields, `child_agent_*`, `ask_user` semantics, `cancel`/`user_response` outbound events, complete `run_end` field list), Read `references/parsing.md`. -`kane-cli generate` (§6) emits a **different** stream — every line is typed `generate_*` (no untyped progress lines), terminated by `generate_done`. Its schema is in `references/generate-parsing.md`. - The assurance conversational commands (`context ingest`/`context extract`, `design tests`, `maintain reconcile`, `cover`) do NOT take `--agent` — they take **`--mode agent`** and speak their own typed stream ending in `done` (open vocabulary — tolerate unknown event types; on 0.7.2+ the stream is strict — every stdout line parses, stderr silent — while on 0.7.1 a merged `context ingest` prints a few prose receipt lines BEFORE the stream — skip to the first `{` line, harmless on 0.7.2+ — and a landing-phase ingest failure ends with prose + exit 1/2 and no stream at all: a refusal, not a crash; on 0.7.2+ those failures ride the stream as `error` + `done`); **for those commands only, exit `3` means paused-and-resumable, not timeout** — schema in `references/assurance-parsing.md`, behavior in `references/assurance.md`. The context sync verbs (`kane-cli context sync`, `kane-cli context push`, `kane-cli context pull`, `kane-cli context clone`; kane-cli 0.8.14+) take `--mode agent` the same way and speak a `sync_*` family ending in `done`; there too exit `3` means a decision or a pull is needed, not a failure — `references/context-sync.md`. `kane-cli context sync setup` refuses without a terminal (`TTY_REQUIRED`): agents use `kane-cli context sync add` and `kane-cli context clone`. `kane-cli testrun run` also emits its own typed stream (`testrun_plan` … terminal `testrun_done`) — schema in `references/testrun.md`. `kane-cli testmd run` may additionally emit `test_md_evidence_ingest` (replay evidence published) and `test_md_bundle_sync` (test bundle synced) — informational; describe in plain language, never surface raw names. The post-run evidence hint (`` evidence: view locally with `kane-cli evidence serve ` ``) is a **stderr** text line, not a stdout event — don't try to parse it from the NDJSON stream; see `references/evidence.md` for how to act on it. --- -## 6. Generate test cases (authoring — no browser) - -`kane-cli generate` authors **Test Scenarios → Test Cases** from a plain-language description. It does **not** drive a browser. **Use it whenever a task needs quick test cases or scenarios from a description — don't hand-author them in chat or a file.** (Requirement documents + coverage accounting → assurance instead: `references/assurance.md`.) Reach for it to: turn a feature / requirement description into a test suite; expand or refine coverage (more edge cases, negative paths, a narrower focus); or save the Functional cases as runnable `_test.md` and hand them to `kane-cli testmd run`. Full details + event schema: **Read `references/generate.md`**. - -Three explicit modes, each runs **one turn then exits**: - -| Mode | Command | -|---|---| -| **New** | `kane-cli generate "" --agent` | -| **Refine** | `kane-cli generate "" --refine --req --agent` | -| **Save** | `kane-cli generate --save --req --agent` → writes runnable `_test.md` | - -**Launch + present** — same as §1: use `Bash` (not Monitor), emit "Generating test cases…" before launch, then parse the output when it returns. Generate is a **quick single turn** — it exits on its own at `generate_done`. - -**After Bash returns**, parse the NDJSON and present only what matters: - -| Show | Event | How | -|------|-------|-----| -| **The deliverable** | `generate_snapshot` | Present scenarios + cases (see below) | -| **Clarifications** | `generate_clarification` | Surface the question — it needs an answer | -| **Save results** | `generate_save_result` | List files written | -| **Errors** | `error` | Surface the message | -| **Skip everything else** | `generate_thinking`, `generate_progress`, `generate_chat`, `generate_start` | Noise — don't narrate | - -At `generate_done`, **present the result adaptively**: -- **≤ ~30 cases** → a nested tree: each scenario, then its cases tagged Positive / Negative / Edge. -- **more than that** → a summary line + a bulleted scenario list (title + case count); expand a scenario's cases only when asked. - -Then offer the next commands from the terminal line's Refine / Save hints (they carry the request id) — don't hand-build them. - -**Clarification → refine (do not skip):** if the turn ends with a clarification, that's **exit 0 — not an error**. Act on it: answer it yourself, or ask your own user, then **re-invoke** `kane-cli generate "" --refine --req --agent`. Never drop a clarification. +## 6. Test cases from a description (no requirement document) -**Attach files:** `--files a,b,c` adds local files (docs / images / PDF / CSV — up to 10, ≤ 50 MB each) as generation context on a **new** or **`--refine`** turn (not `--save`); each emits a `generate_upload` line before `generate_start`. Details in `references/generate.md`. +When the user describes what to test in chat and has no document, use the description as the requirement. -**Save is Functional-only:** `--save` writes only **Functional** cases to `_test.md` (under `/.testmuai/tests` by default). Non-functional cases (Security, Performance, …) are generated and shown but not saved. Run saved files with **`kane-cli testmd run`** (`references/testmd.md`) — that's the generate → testmd pipeline. +1. Write the user's words to `requirements/.md` in the project, as given, one heading per feature, nothing invented. Tell the user the file exists and that the tests will cite it. +2. `kane-cli context ingest requirements/.md --mode agent`, review the extracted use-cases at the checkpoint, then `kane-cli design tests …`, exactly as `references/assurance.md` describes. +3. Present the designed tests, gaps, warnings and the variables that need values (§3 **Fill the variables before any run**). The designed tests get their own review checkpoint before any run (`references/assurance.md`). Then hand the set to `kane-cli testrun run … < /dev/null` (`references/assurance.md`, the authoring bridge). -Internal event/field names (`generate_snapshot`, `request_id`, …) are for parsing only — never show them to the user (§5 rule). Wire schema: `references/generate-parsing.md`. +A draft in chat cannot be run, and a hand-written `_test.md` carries no requirement link and no coverage accounting; the deliverable is the designed test. --- @@ -368,7 +346,6 @@ Internal event/field names (`generate_snapshot`, `request_id`, …) are for pars | User wants to save/persist/re-run a test | `references/testmd.md` | | Run a suite of saved `_test.md` tests as one batch | `references/testrun.md` | | Run a suite on the cloud grid (`--remote`), incl. mobile suites from any machine | `references/testrun.md` §Remote + `references/mobile.md` §Remote | -| You need quick test cases or scenarios from a description | `references/generate.md` | | User has requirement docs (PRD/spec) → designed suite, coverage, or suite upkeep | `references/assurance.md` | | Need the assurance NDJSON event schema (`--mode agent`) | `references/assurance-parsing.md` | | Share the context store with a team, join one, keep stores level, or resolve a sync conflict (kane-cli 0.8.14+) | `references/context-sync.md` | @@ -376,7 +353,6 @@ Internal event/field names (`generate_snapshot`, `request_id`, …) are for pars | View, share, validate, or merge evidence packs | `references/evidence.md` | | Multiple independent browser tasks | `references/parallel.md` | | Need full NDJSON event schema (`run`) | `references/parsing.md` | -| Need the `generate` NDJSON event schema | `references/generate-parsing.md` | | Browse / create projects or folders, or parse the auto-default event | `references/test-manager.md` | | Start of every session: preflight, the ready card, sign-in | `references/ready-check.md` | | A person's first session: run first, the tour, three choices | `references/first-run.md` | @@ -388,7 +364,7 @@ Internal event/field names (`generate_snapshot`, `request_id`, …) are for pars ## Command-specific completion -The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). `generate` emits `generate_done`. Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. +The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. Progress is for live display: count only `done`/`failed` completions, retaining child and execution context when step indices repeat. diff --git a/skill-installer/skills/references/assurance-parsing.md b/skill-installer/skills/references/assurance-parsing.md index a69d44e..9562a82 100644 --- a/skill-installer/skills/references/assurance-parsing.md +++ b/skill-installer/skills/references/assurance-parsing.md @@ -41,8 +41,8 @@ A stream that ends **without** `done` means the process crashed — outcome unkn | `held` / `update_held` *(0.7.1+)* | items held for the user's review instead of committed: `source_id` + `count` + `reason` / `count` + `targets[]` | surface the count and that review happens at resume/`context review` | | `commit` | what landed: counts + `minted[]` (`cid` + `logical_id`); extract adds `proposal_id` | translate ("5 use-cases extracted"); `logical_id` slugs are how you reference nodes later | | `receipt` | per-phase commit receipt (design; extract also emits one at its commits): `commit_n`, `phase`, `committed[]`, `reused`, `rejected[]`, `warnings[]`, `next`, and (design only) `parity` | surface non-empty `rejected[]` and `warnings[]` in plain language; meaningful reuse is worth one line | -| `variables_declared` *(0.8.12+)* | design: stubs written by a phase commit — `file` (pool file) + `variables[]` (`name`, `description`, `secret`); only names that existed in no variable file | surface as a to-do ("N variables need values — in ") | -| `variables_summary` *(0.8.12+)* | design, end of run: every declared name still needing a value — same shape | repeat the to-do in the closing summary; fills gate authoring | +| `variables_declared` *(0.8.12+)* | design: the stubs the tests phase wrote, sent just before that phase's `commit` event — `file` (pool file) + `variables[]` (`name`, `description`, `secret`); only names that existed in no variable file | surface as a to-do ("N variables need values — in "), then fill them before any run of these tests (SKILL.md §3) | +| `variables_summary` *(0.8.12+)* | design, after `session_complete` on a clean completion only (a paused or refused run never sends it): every name this run declared, same shape, no re-check of the pool | repeat the to-do in the closing summary; on a pause build it from `variables_declared` alone; unfilled names are typed as written when the test is authored | | `message_sent` | `--message` delivered: `sid`, `chars` | confirmation only | | `panel_resolved` *(0.7.1+)* | a `--answer` flag landed on a pending question: `id`, `by`, `via` | confirmation only | | `ask_deferred` *(0.7.1+)* | `--with-source` set the pending batch aside: `source_id`, `cid`, `questions` (count) | tell the user the questions were deferred while the agent reads the new source | @@ -99,4 +99,4 @@ One payload event carrying the full `--json` document — `coverage` for `cover` | `3` | **paused and resumable** — not a failure; run the pause loop. Includes crash-pauses (0.7.1+). On the sync verbs (0.8.14+) `done{status: "paused"}` means decisions are waiting: after a rebase walk the rebase is open (answer with `--answer`); after `kane-cli context sync doctor --abort` it is closed and the unanswered decisions stayed in the backup — the `sync_rebase_done` before it says which (`status` `paused` or `aborted`); `paused` with no `sync_rebase_done` means doctor set aside a rebase whose saved state it could not read — `sync_doctor.detail` says the next `kane-cli context sync` reapplies the saved backup. `done{status: "refused", exit_code: 3}` means a pull or a rebase is needed first — `references/context-sync.md` §4 | | `130` | force-interrupted — for extract and design, resumable only if a `session_paused` event arrived; on the sync verbs the rebase state is on disk: `kane-cli context sync doctor --mode agent` shows it, `kane-cli context sync --mode agent` resumes it | -Reminder: this exit-3 meaning is **specific to these assurance commands**. `run` / `testmd` / `testrun` / `generate` keep their own meanings (3 = timeout/cancelled). +Reminder: this exit-3 meaning is **specific to these assurance commands**. `run` / `testmd` / `testrun` keep their own meanings (3 = timeout/cancelled). diff --git a/skill-installer/skills/references/assurance.md b/skill-installer/skills/references/assurance.md index cab6a28..05b188a 100644 --- a/skill-installer/skills/references/assurance.md +++ b/skill-installer/skills/references/assurance.md @@ -2,9 +2,8 @@ # Assurance — Agent Surface -When the user has **requirements** — a PRD, a spec, acceptance notes — and wants tests designed from them, wants to know what's covered, or wants the suite kept current, use the **assurance commands** (`kane-cli context`, `design`, `cover`, `maintain`). Do not hand-write the tests, and do not reach for `generate`: +When the user has **requirements** — a PRD, a spec, acceptance notes — and wants tests designed from them, wants to know what's covered, or wants the suite kept current, use the **assurance commands** (`kane-cli context`, `design`, `cover`, `maintain`). Do not hand-write the tests: -- `kane-cli generate` = quick scenarios/cases from a one-line description. No requirement linkage. - **Assurance** = tests derived from the actual documents, every claim cited, every test permanently tagged with the acceptance criteria it verifies, coverage measured against requirements. Use it whenever the user cares about "what exactly is covered, and how do we know?" Everything here works over a local store (`.context/` in the project directory) that the commands create and manage themselves. @@ -19,7 +18,7 @@ kane-cli context review --verdicts --json # 2. CHECKPOINT: us kane-cli design tests --use-case --mode agent --max 8 # 3. design ACs, scenarios, tests kane-cli context review --verdicts --json # 4. CHECKPOINT: user approves the design kane-cli testmd run .testmuai/tests/_test.md --agent # 5. author each kept test once (real browser) -kane-cli testrun run --match 't-' # 6. batch replays from then on +kane-cli testrun run --match 't-' < /dev/null # 6. batch replays from then on kane-cli cover gaps # 7. designed % × proven % + per-use-case debt kane-cli maintain reconcile --from --source-id --mode agent # when a source changes (§11) ``` @@ -35,12 +34,12 @@ Extract, design, and reconcile call the KaneAI service and consume credits; ever ## 2. The pause loop — exit 3 is a pause, NOT a failure -**Scope: this rule applies to `context extract`/`context ingest`, `design tests`, and `maintain reconcile` (§11).** (For `run`/`testmd`/`testrun`/`generate`, exit 3 still means timeout/cancelled.) +**Scope: this rule applies to `context extract`/`context ingest`, `design tests`, and `maintain reconcile` (§11).** (For `run`/`testmd`/`testrun`, exit 3 still means timeout/cancelled.) These commands take **`--mode agent`** — not `--agent`; they reject that flag, and a bare non-TTY invocation exits `2` asking for an explicit mode. In `--mode agent`, **every question pauses the run** (0.8.8+ — earlier CLIs auto-answered low/medium-risk questions with their recommended defaults and paused only on high risk): - The run exits `3`, emits `session_paused` with the session id, the questions in full (text, options, the recommended one, risk, rationale), and the verbatim resume command. -- **Never drop a pause** (same rule as generate clarifications). Answer it: if your own context clearly resolves the question, answer it yourself; otherwise surface the question — with its options and recommendation — to your user and get their answer. +- **Never drop a pause.** Answer it: if your own context clearly resolves the question, answer it yourself; otherwise surface the question — with its options and recommendation — to your user and get their answer. - **0.7.1+ sessions are durable from the first turn**: a crash that left a checkpoint exits `3` with a `session_paused` carrying `crashed: true` (no `pending_questions`) and the resume command — exit 3 always means "resumable". A crash before anything durable was saved still exits `1`; check `context sessions --json` before retrying anything paid. Three ways to resume: @@ -119,13 +118,13 @@ kane-cli design tests --use-case --mode agent --max 8 - There is **no `--because` flag on `design tests`** — interactively the session collects the redesign reason itself; headless `--force` proceeds with an auto-stamped reason. - `--phase ` (0.7.1+) re-enters a design at a phase, re-seeded from the committed earlier phases; missing predecessors exit 2 with the commands to run first in `next`. - Output: acceptance criteria, scenarios, exactly one test per scenario — written as runnable files under `.testmuai/tests/*_test.md`, each assert step tagged with the criteria it verifies. Plus **gaps** (recorded, ranked missing pieces) and **warnings** (e.g. a test claiming more criteria than its check asserts). Citations are verified against the pinned source text before commit (0.7.1+) — a `CITE_UNVERIFIED` error means a citation could not be verified even after repair. -- **Variables (0.8.12+).** Design reads every `*.json` in the user's variable directories (names + has-value only, never values) and reuses a name that exists before inventing one. Each name it invents is declared: an empty stub lands in `.testmuai/variables/assurance.json`, never a value, never over an existing key. The stream tells you which: `variables_declared` at the commit that wrote them, `variables_summary` at the end (schema in `references/assurance-parsing.md`). **Surface these as a to-do** — "2 variables need values: login_email, login_password — in .testmuai/variables/assurance.json" — because a designed test refuses to author until they are filled (SKILL.md §3, Unresolved variables). Fill them yourself only if the user gave you the values. +- **Variables (0.8.12+).** Design reads every `*.json` in the user's variable directories (names + has-value only, never values) and reuses a name that exists before inventing one. Each name it invents is declared: an empty stub lands in `.testmuai/variables/assurance.json`, never a value, never over an existing key. The stream tells you which: `variables_declared` in the tests phase, just before its `commit` event, and `variables_summary` after `session_complete` on a clean completion (a paused run sends only the first; schema in `references/assurance-parsing.md`). **Surface these as a to-do** — "2 variables need values: login_email, login_password — in .testmuai/variables/assurance.json" — because a designed test authored before they are filled types the placeholders as written (SKILL.md §3, Unresolved variables). Before any run of these tests, fill them: SKILL.md §3 **Fill the variables before any run**, asking with each row's `description`. - **Present tests, gaps, warnings, AND variables needing values** — first-class output, not noise. Then go to the review checkpoint (§4) before any authoring. - `kane-cli design explain ` replays *why* a test exists (technique, boundary values, criteria) with zero AI cost — use it when the user asks "why this test?". ## 6. The authoring bridge — from designed files to batch runs -A freshly designed test has never been executed. On 0.8.4+ hand the set straight to `kane-cli testrun run`: unauthored members classify as `author`, the run authors them in a real browser, and afterwards the authored and replayed evidence consolidates into one published execution — best-effort: when consolidation cannot complete, evidence stays split rather than lost. `--from-context` (0.8.4+) selects members by assurance test ids and follows edit supersessions. Designed tests may carry `{{variables}}` for values the requirements never pinned (a store URL, a product name) — supply them per `references/testmd.md`. `kane-cli testmd run --agent` remains the single-test authoring path, and on pre-0.8.4 CLIs it is REQUIRED first — `testrun` there refuses never-authored members (`missing_meta`). Evidence packs seal per `references/evidence.md`. +A freshly designed test has never been executed. On 0.8.4+ hand the set straight to `kane-cli testrun run`: unauthored members classify as `author`, the run authors them in a real browser, and afterwards the authored and replayed evidence consolidates into one published execution — best-effort: when consolidation cannot complete, evidence stays split rather than lost. `--from-context` (0.8.4+) selects members by assurance test ids and follows edit supersessions. Designed tests may carry `{{variables}}` for values the requirements never pinned (a store URL, a product name) — fill them first (SKILL.md §3), or the run types the placeholder as written. `kane-cli testmd run --agent` remains the single-test authoring path, and on pre-0.8.4 CLIs it is REQUIRED first — `testrun` there refuses never-authored members (`missing_meta`). Evidence packs seal per `references/evidence.md`. ### 6.1 A designed test is the design — do not edit it by hand diff --git a/skill-installer/skills/references/cards.md b/skill-installer/skills/references/cards.md index 9904459..b035268 100644 --- a/skill-installer/skills/references/cards.md +++ b/skill-installer/skills/references/cards.md @@ -16,6 +16,7 @@ Every result is an emoji table. A one-line "Test passed" instead of the card is - **`🟡 Didn't start` is not `🔴 Failed`.** When nothing ran, say what to fix. - **Secret-looking values never go in chat.** For a missing value whose name contains `password`, `secret`, `token` or `key`, add an empty entry to the variables file for the person to fill. Ask in chat only for plain values (a URL, a user name). - If the run's output carried an update notice, add one quiet last line under the card: `kane-cli is available.` +- **Variables with no value go on the card.** When the run's output carried the `unresolved_variables` warning, add a `⚠️ **Variables**` row before ➡️ Next naming each one. ## 2. Run, passed @@ -72,7 +73,7 @@ Exit code `1`, or `status: "failed"`. Show the failing step's screenshot under t ## 4. Didn't start -Exit code `2`: nothing ran and no credits were used. Causes include missing variable values, no start URL, sign-in or setup errors, a test file that does not parse, an invalid suite plan, a cloud grid refusal. +Exit code `2`: nothing ran and no credits were used. Causes include no start URL, sign-in or setup errors, a test file that does not parse, an invalid suite plan, a cloud grid refusal. ```markdown | | | diff --git a/skill-installer/skills/references/debug.md b/skill-installer/skills/references/debug.md index 7d8faf0..372d188 100644 --- a/skill-installer/skills/references/debug.md +++ b/skill-installer/skills/references/debug.md @@ -66,7 +66,7 @@ For `run` objectives and plain `_test.md` files. For a designed test, read the s | 🎯 Agent clicks wrong element | Ambiguous UI, multiple similar elements | Be more specific: "click the **blue** 'Submit' button in the **checkout form**" | | 👁️ Agent says done but didn't finish | Objective too vague | Add explicit assertions: "assert the confirmation page shows order number" | | 💀 Exit code 2, no steps | Auth, TMS credential exchange, or Chrome failure | Check `kane-cli whoami`, verify Chrome is available | -| ❓ Exit code 2 with "did you mean …" | Bare-objective shortcut — agent ran `kane-cli ""` without the `run` subcommand | Re-invoke as `kane-cli run "" --agent` (same rule for `testmd run` / `generate`) | +| ❓ Exit code 2 with "did you mean …" | Bare-objective shortcut — agent ran `kane-cli ""` without the `run` subcommand | Re-invoke as `kane-cli run "" --agent` (same rule for `testmd run`) | | 📤 Upload silently fails after configuring a project/folder by hand | Saved ID is invalid (typo, deleted, no access) | No action needed — the next run detects the 4xx and auto-defaults a working project/folder. To rebind manually: `kane-cli config project` (TTY picker) or `kane-cli projects list` → `kane-cli config project ` (see `references/test-manager.md`) | | ⏱️ Exit code 3 | Timeout or cancelled | Increase `--timeout` or `--max-steps`, or split into smaller objectives | | 🚫 "CDP endpoint not reachable" | Chrome not running | Let kane-cli manage Chrome (remove `--cdp-endpoint`) | diff --git a/skill-installer/skills/references/generate-parsing.md b/skill-installer/skills/references/generate-parsing.md deleted file mode 100644 index d767ebe..0000000 --- a/skill-installer/skills/references/generate-parsing.md +++ /dev/null @@ -1,65 +0,0 @@ - - -# Reading `generate --agent` Output - -> **Internal reference only.** The event types and field names below are for you to parse programmatically. **Never expose them to the user** — present plain-language scenarios/cases per `references/generate.md`, never `generate_snapshot`, `request_id`, `generate_done`, or raw JSON. - -With `--agent`, `kane-cli generate` writes **one JSON object per line** to **stdout**. Unlike `run` (which has untyped `step`/`status` progress lines), **every generate line is typed** — it always has a `type` field. That makes parsing simpler: - -``` -for each line of NDJSON: - parse JSON, switch on obj.type - if obj.type === "generate_done" → terminal event, stop parsing - if obj.type === "generate_snapshot" → the deliverable (full scenarios + cases) - if obj.type === "generate_clarification" → turn ended awaiting an answer (still exit 0) - else → progress / informational -``` - -Build post-turn logic on **`generate_done`** (terminal, stable schema) and read the result from **`generate_snapshot`** (emitted once, at turn end). - -## Event types - -Generate-specific (typed `generate_*`): - -| `type` | Key fields | Meaning → what to do | -|---|---|---| -| `generate_upload` | `file`, `index`, `total`, `status` | Only when `--files` is used. Emitted once per attached file **before `generate_start`**, while the file uploads; `status` goes `"uploading"` → `"done"` (or `"failed"`). Progress only — narrate "attaching …" or ignore. A failed upload surfaces as an `error` + non-zero exit. | -| `generate_start` | `request_id`, `objective_chars`, `scenario_limit`, `per_scenario_limit`, `is_refine` | Turn began. **Capture `request_id`** — it's the handle for every later `--refine` / `--save`. `is_refine` distinguishes a new request from a continuation. | -| `generate_thinking` | `took_ms` | Liveness only. Narrate "thinking…" or ignore. | -| `generate_progress` | `pct` | Milestone (25 / 50 / 75 / 100). Optional progress display — not a completion signal. | -| `generate_snapshot` | `scenario_count`, `case_count`, `scenarios[]` | **The deliverable** — full scenarios, each with its cases. Each case carries `title`, `polarity` (`"p"`/`"n"`/`"e"` = Positive / Negative / Edge), `category` (`"Functional"`, `"Security"`, …), `priority`. Present it per `generate.md`. Emitted exactly once. | -| `generate_clarification` | `text` | The generator needs an answer; the turn ended (exit 0) awaiting it. **Not an error** — answer via `--refine --req` (see `generate.md`). | -| `generate_chat` | `text` | The model's prose reply (e.g. what a refine changed). Show as info. | -| `generate_save_result` | `suite_dir`, `saved`, `fell_back`, `warning?` | `--save` wrote files. `saved` = files written; `fell_back` = cases written as prose because they couldn't be expanded; `warning` e.g. `"no functional test cases"`. | -| `generate_done` | `request_id`, `status`, `scenario_count`, `case_count`, `refine_hint`, `save_hint`, `suite_dir?` | **Terminal line.** Branch on `status`; use `refine_hint` / `save_hint` **verbatim** as the next command. `suite_dir` present only after a `--save`. | - -Shared with `run` (NOT generate-specific — don't treat as part of the generate contract): - -| `type` | Fields | Meaning | -|---|---|---| -| `project_folder_auto_defaulted` | resolved project + folder (id, name) | Run-startup gate auto-resolved a project/folder when none was configured (or the cached one was stale/invalid). Fires before `generate_start`. Translate to plain language (see `references/test-manager.md`). | -| `error` | `message` | A failure occurred; pair with the exit code to decide retry vs abort. | -| `update_available` | `current`, `latest`, `severity` | A newer kane-cli exists. Informational; emitted before `generate_start`. | - -## Terminal `generate_done` - -Example after a **`--save`** run (note `suite_dir` — a new/refine turn omits it): - -```json -{ - "type": "generate_done", - "request_id": "23271", - "status": "completed", - "scenario_count": 3, - "case_count": 11, - "refine_hint": "kane-cli generate \"\" --refine --req 23271", - "save_hint": "kane-cli generate --save --req 23271", - "suite_dir": ".testmuai/tests/checkout-23271" -} -``` - -`status`: `completed` (exit 0 — including a turn that ended with a clarification) · `failed` or `ended` (exit 1) · `stopped` (exit 3). `suite_dir` is present only after a `--save`. The hints carry no `--agent` flag, but it is auto-on for non-TTY callers (agents/pipes), so re-invoking a hint verbatim still yields NDJSON. - -## Exit codes - -`0` completed · `1` failed · `2` error (auth / setup / transport, or an invalid flag combination) · `3` stopped / cancelled · `130` interrupted (Ctrl-C). Interactive `--refine`/`--save` continuity is by re-invocation with `--req` — there is no stdin relay. diff --git a/skill-installer/skills/references/generate.md b/skill-installer/skills/references/generate.md deleted file mode 100644 index d947c1e..0000000 --- a/skill-installer/skills/references/generate.md +++ /dev/null @@ -1,139 +0,0 @@ - - -# Generating Test Cases with `kane-cli generate` - -`kane-cli generate` turns a plain-language description of *what to test* into structured **Test Scenarios** (logical groupings) each containing **Test Cases** (typed Positive / Negative / Edge). It calls the AI Test Case Generator — **no browser is launched**. The result is a tree of scenarios + cases you present to the user, refine conversationally, and optionally save as runnable `_test.md` files. - -**Use this whenever a task needs test cases or scenarios written — don't hand-author them in chat or a scratch file.** Reach for it to: turn a requirement / feature description into a test suite; expand or refine coverage (more edge cases, negative paths, a narrower or broader focus); or save the Functional cases as runnable `_test.md`. It is **not** for driving a browser — that's `kane-cli run` (§3 of SKILL.md). - -> For the full web-product picture (dashboards, issue-link inputs), see the public docs: . The CLI takes a **text objective**, optionally with local files attached via **`--files`** (see "Attaching files" below). - -## The three modes — one turn per invocation, then exit - -There is **no interactive session**. Each invocation runs exactly one generation turn and exits. Continuity across turns is carried by a **request id** (`--req `) that the previous turn's terminal line hands back. - -| Mode | Command | Notes | -|---|---|---| -| **New** | `kane-cli generate "" --agent` | Starts a fresh request. Capture the request id from the terminal line. | -| **Refine** | `kane-cli generate "" --refine --req --agent` | Adjusts an existing request. `--refine` **and** `--req` required; needs a change description. | -| **Save** | `kane-cli generate --save --req [--out ] --agent` | Writes the request's Functional cases to `_test.md`. No new turn, takes no objective. | - -`--refine` and `--save` always run headless (even from a terminal). - -### Flags - -| Flag | Purpose | -|---|---| -| `--agent` | Typed NDJSON on stdout (auto-on when stdin is not a TTY) | -| `--req ` | The request id to `--refine` or `--save` | -| `--out ` | Save target — **only** with `--save`; default `/.testmuai/tests` | -| `--name ` | Names the run and the saved suite folder | -| `--scenario-limit ` / `--per-scenario-limit ` | Cap scenarios / cases-per-scenario | -| `--memory` | Use the memory layer — reuse relevant existing cases, reduce duplicates | -| `--files ` | Comma-separated local files to attach as context (new / refine only — see "Attaching files") | -| `--project ` / `--folder ` | Test Manager project / folder | -| `--username` / `--access-key` | Auth (same as `run`) | - -If neither `--project`/`--folder` nor a saved project/folder is set when generation starts, kane-cli auto-resolves one headlessly and emits a `project_folder_auto_defaulted` event before `generate_start`. Translate it to a one-line note for the user — full handling lives in `references/test-manager.md`. - -## Attaching files - -Pass local files as extra context with `--files ` on a **new** generation or a **`--refine`** (not `--save`). The generator reads them and reflects them in the scenarios + cases — attach a spec, a screenshot of the UI, a PDF / Word doc, or a CSV of inputs. - -```bash -kane-cli generate "test the login flow described in the attached spec" --files ./login-spec.pdf,./wireframe.png --agent -``` - -- **Supported types** — documents (`.txt .json .xml .csv .pdf .docx .xlsx`), images (`.jpg .jpeg .png .gif .bmp .webp`), audio (`.mp3 .wav .m4a`), video (`.mp4 .mov .webm .mpeg .mpga`). -- **Limits** — up to **10 files**, each **≤ 50 MB**. -- **Validated as a set, up front** — if any path is missing, an unsupported type, too large, or over the count, the whole command is rejected (exit `2`) **before anything is sent** and the offending paths are listed; fix and re-run. Files outside the current directory are allowed but flagged with a warning on stderr. -- **`new` / `--refine` only** — combining `--files` with `--save` exits `2`. - -Under `--agent`, each file emits a `generate_upload` line (`status` `uploading` → `done`) **before** `generate_start` — see `references/generate-parsing.md`. *(Interactively in the TUI, type `@` in the generate prompt to attach a file inline.)* - -## Presenting a result (adaptive) - -The terminal data carries the full scenarios + cases. **Present it based on size**: - -- **≤ ~30 cases → a nested tree** (scenario, then each case with its type tag): - ``` - ✓ Generated 3 scenarios · 11 cases (request 23271) - - ▸ Login - - Valid credentials [Positive] - - Wrong password [Negative] - - Empty fields [Edge] - ▸ Checkout - - Guest checkout [Positive] - - Expired card [Negative] - ... - ``` -- **more than ~30 cases → a summary + scenario list** (cases on request): - ``` - ✓ Generated 6 scenarios · 84 cases (request 23271) - - • Login (12 cases) - • Checkout (20 cases) - • Cart management (14 cases) - ... - ``` - -Always end with the **next-step commands the terminal line provides** (Refine / Save) — they already carry the request id, so don't hand-build them. - -## Clarifications — act on them, never drop them - -If a turn ends with a **clarification question**, that is **success (exit 0)**, not an error — the generator needs an answer before it can continue. You must act on it: - -1. Read the question. -2. **Decide** — answer it yourself from context, **or** surface it to your user and get an answer. -3. **Re-invoke** with the answer as a refine: - ```bash - kane-cli generate "" --refine --req --agent - ``` - -## The refine → save → run loop - -```bash -# 1. New request -kane-cli generate "checkout flow on a shopping site" --agent -# → terminal line carries request id 23271 + Refine/Save hints - -# 2. Refine (repeat as needed) -kane-cli generate "also cover an expired card and an out-of-stock item" --refine --req 23271 --agent - -# 3. Save the Functional cases as runnable _test.md -kane-cli generate --save --req 23271 --agent -# → /.testmuai/tests///_test.md - -# 4. Run / replay them -kane-cli testmd run .testmuai/tests///_test.md --agent -``` - -**Save is Functional-only.** `--save` writes only test cases whose category is **Functional** — those are the ones runnable as `_test.md`. Non-functional cases (Security, Performance, etc.) are generated and shown in the result but are **not** written; saving a request with no Functional cases writes nothing and says so. Saved files are ordinary `_test.md` tests — see `references/testmd.md` for running, editing, and replay. This is the **generate → testmd** pipeline: author cases here, run them there. - -## Exit codes - -| Code | Meaning | -|---|---| -| 0 | Turn completed (including a turn that ended with a clarification) | -| 1 | Generation failed (or ended) | -| 2 | Error — auth / setup / transport, or an invalid flag combination | -| 3 | Generation stopped / cancelled | -| 130 | Interrupted (Ctrl-C) | - -Invalid flag combinations exit `2` with a message on stderr. The full set: - -- `--refine` and `--save` together -- `--refine` without `--req` -- `--refine` without a change description -- `--refine` combined with `--out` (`--out` is save-only) -- `--save` without `--req` -- `--save` with a description (it takes none) -- `--out` without `--save` -- `--files` with `--save` (files attach to a new generation or a refine, not a save) -- `--req` without `--refine` or `--save` -- a new generation with no description - -## Reading the output - -`--agent` emits one typed JSON object per line. For the full event schema and the parse strategy, Read **`references/generate-parsing.md`**. As with all `--agent` output, the field names are for parsing only — **never show them to the user**; present plain-language scenarios and cases (per "Presenting a result" above). diff --git a/skill-installer/skills/references/mobile.md b/skill-installer/skills/references/mobile.md index 222a8c4..eca231e 100644 --- a/skill-installer/skills/references/mobile.md +++ b/skill-installer/skills/references/mobile.md @@ -9,10 +9,10 @@ Desktop (the browser) is the **default** target and the primary use of kane-cli. | Where the device runs | Host | Commands | |---|---|---| | **Local** (a simulator/emulator on this machine) | **macOS on Apple Silicon (arm64) only** — not Intel Macs, Linux, or Windows | `run --target …`, `testmd run`, `testrun run` | -| **Cloud grid** (a virtual device on a HyperExecute macOS host) | **Any machine** — no Xcode / Android Studio needed. The account needs a HyperExecute plan with macOS runners | `testrun run --remote …` — see §Remote below | +| **Cloud grid** (a virtual device on a HyperExecute macOS host) | **Any machine** — no Xcode / Android Studio needed. The account needs a HyperExecute plan with macOS runners | `testrun run --remote … < /dev/null` — see §Remote below | - **Desktop stays the default.** The `--target` axis is what selects mobile. Leave it off and you get the browser. -- If the user is not on mac-arm64, local mobile is not an option — **offer the grid**: save the objective as a `_test.md` (`target: emulator|simulator` + `app:`) and run it with `kane-cli testrun run --remote --device-name "" --os-version `. Do not tell them mobile is unavailable. +- If the user is not on mac-arm64, local mobile is not an option — **offer the grid**: save the objective as a `_test.md` (`target: emulator|simulator` + `app:`) and run it with `kane-cli testrun run --remote --device-name "" --os-version < /dev/null`. Do not tell them mobile is unavailable. ## The three targets @@ -111,7 +111,7 @@ Everything else about `_test.md` (step bodies, replay/cascade, commands) is unch `kane-cli testrun run` accepts mobile `_test.md` members (0.8.7+): -- **Locally** (mac-arm64 with the setup above): `kane-cli testrun run --device-name "" --os-version ` — the device as `kane-cli devices list --target ` prints it, or the members' own `device_name:`/`os_version:`. +- **Locally** (mac-arm64 with the setup above): `kane-cli testrun run --device-name "" --os-version < /dev/null` — the device as `kane-cli devices list --target ` prints it, or the members' own `device_name:`/`os_version:`. - **On the cloud grid, from any machine**: add `--remote` — next section. Full testrun flags, events, and rollup: `references/testrun.md`. ## Remote: mobile suites on the cloud grid (`testrun run --remote`) @@ -121,9 +121,9 @@ One command turns a mobile suite into a HyperExecute job on a macOS host that bo ```bash kane-cli plugin install remote-execution # once; then `kane-cli plugin doctor remote-execution` kane-cli devices list --target emulator --remote --agent # grid catalog: name + os_versions per row -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 --dry-run # validate + resolve device, no job -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 -kane-cli testrun run tests/ios/ --remote --device-name "iPhone 15" --os-version 17.5 +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 --dry-run < /dev/null # validate + resolve device, no job +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 < /dev/null +kane-cli testrun run --match '^tests/ios/' --remote --device-name "iPhone 15" --os-version 17.5 < /dev/null ``` **Always `--dry-run` first** — it runs the remote preflight and resolves the device against the catalog at no cost. Use `Bash` with a long timeout (up to 600000 ms) for the real run: device setup + app install + members take several minutes. diff --git a/skill-installer/skills/references/parallel.md b/skill-installer/skills/references/parallel.md index 5bc51bb..0866044 100644 --- a/skill-installer/skills/references/parallel.md +++ b/skill-installer/skills/references/parallel.md @@ -4,7 +4,7 @@ For multiple independent browser tasks, decompose and run in parallel using the Agent tool. -> **Saved tests? Use testrun instead.** If the tasks are committed `_test.md` files, do NOT hand-roll parallelism — `kane-cli testrun run --parallel N` gives you isolated Chromes, a pooled scheduler, one rollup, and one evidence pack. Read `references/testrun.md`. This reference is for **ad-hoc `run` objectives** only. +> **Saved tests? Use testrun instead.** If the tasks are committed `_test.md` files, do NOT hand-roll parallelism — `kane-cli testrun run --parallel N < /dev/null` gives you isolated Chromes, a pooled scheduler, one rollup, and one evidence pack. Read `references/testrun.md`. This reference is for **ad-hoc `run` objectives** only. ## When to Split diff --git a/skill-installer/skills/references/parsing.md b/skill-installer/skills/references/parsing.md index 2e6aac9..ec3952e 100644 --- a/skill-installer/skills/references/parsing.md +++ b/skill-installer/skills/references/parsing.md @@ -55,24 +55,25 @@ These are **untyped** — they have no `type` field. Do **not** key on `event.ty | Event (`type` field) | Key Fields | Purpose | |-------|-----------|---------| -| `project_folder_auto_defaulted` | resolved project + folder (id, name) | Run-startup gate auto-resolved a project/folder when none was configured (or the cached one was stale/invalid). Fires before any progress event on `run` / `testmd run` / `generate`. Translate to plain language (see `references/test-manager.md`). | +| `project_folder_auto_defaulted` | resolved project + folder (id, name) | Run-startup gate auto-resolved a project/folder when none was configured (or the cached one was stale/invalid). Fires before any progress event on `run` / `testmd run`. Translate to plain language (see `references/test-manager.md`). | | `bifurcation` | `flows[]`, `count` | Agent split objective into sub-flows | | `child_agent_start` | `child_id`, `objective`, `parent_step` | Child agent spawned | | `child_agent_end` | `child_id`, `success`, `steps_taken`, `summary` | Child agent finished | | `ask_user` | `question`, `step_index`, `options?` | Agent needs user input | -| `error` | `message`, `code?` | Error occurred. With `code: "unresolved_variables"` it is the pre-run refusal — the only event of the run, no `run_end` follows; schema below. | +| `error` | `message` | Error occurred | +| `warning` | `code`, `message`, … | *(0.8.15+)* Before any progress event: `code: "unresolved_variables"` — a `{{name}}` had no value; the run continues. Schema below. | | `test_md_evidence_ingest` | `status: "ok"\|"failed"`, `evidence_id`, `stage?` (failure only) | `testmd run` only: a replay's evidence pack published to the dashboard. Informational. | | `test_md_bundle_sync` | `status: "ok"\|"failed"`, `commit_id`, `bytes?` (success) / `stage?` (failure) | `testmd run` / `testmd sync`: test bundle pushed to the cloud after an authored commit. Informational. | | `testrun_*` family | see `references/testrun.md` | Emitted only by `kane-cli testrun run`; terminal event is `testrun_done`, not `run_end`. | **Note:** The `run` stream has no `run_start` event; startup metadata or errors can precede progress. -### `error` with `code: "unresolved_variables"` (0.8.12+) +### `warning` with `code: "unresolved_variables"` (0.8.15+) -Emitted by `run`, `testmd run` and (after `testrun_plan`) `testrun run` when an authored step references a `{{name}}` that has no value. Nothing was dispatched; exit code `2`; stderr is silent. +Emitted by `run`, `testmd run` and (after `testrun_plan`) `testrun run` when an authored step references a `{{name}}` that has no value. The run goes ahead: the name is typed as written unless a step sets it first. The warning does not change the exit code. A run given a Test Manager dataset row can emit a second `warning` with the same code for `${x}` names that are no column of that row, still before any progress. ```json -{"type":"error","code":"unresolved_variables","message":"2 variable(s) have no value — nothing was dispatched", +{"type":"warning","code":"unresolved_variables","message":"2 variable(s) have no value — typed as written unless a step sets them first", "suggested_file":".testmuai/variables/variables.json", "variables":[ {"name":"checkout_url","reason":"not_declared","used_by":[{"file":"objective","step":1}]}, @@ -82,14 +83,12 @@ Emitted by `run`, `testmd run` and (after `testrun_plan`) `testrun run` when an | Field | Meaning | |---|---| -| `variables[].reason` | `value_missing` — the key exists in `variables[].file` with an empty value · `not_declared` — the key is in no variable file | +| `variables[].reason` | `value_missing` — the key exists in `variables[].file` with an empty value · `not_declared` — the key is in no variable file · `not_a_dataset_column` — a `${x}` that is no column of the Test Manager dataset the run was given (`variables[].dataset` names it) | | `variables[].file` | the pool file that holds the empty key (`value_missing` only); `--variables` when it came inline | | `variables[].used_by[]` | `{file, step}` — `file` is `objective` for `kane-cli run`, else the test file (flattened step index) | | `suggested_file` | where to add a `not_declared` key: `.testmuai/variables/assurance.json` inside an assurance store, `variables.json` otherwise | -Terminal: do not re-run the same command. Supply values (`--variables`, or fill the file) and run again. - -**Note:** `ask_user` is auto-disabled when stdin is not a TTY. Since agents typically run kane-cli as a subprocess, ask_user events will not be emitted. Write objectives that don't require interactive input. +Report it with the run's outcome. If a later step failed on a literal placeholder, point at this. ## Parsing Strategy for one-shot `run` @@ -159,6 +158,6 @@ To cancel a run: ## Command-specific completion -The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). `generate` emits `generate_done`. Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. +The `run_end` parsing strategy applies to one-shot `run` only. For `testmd run`, collect `test_md_done.overall_status`, `duration_s`, `session_id`, and optional `share_url`; embedded `run_end` events can finish individual steps. Local suites emit `testrun_done`; dispatched remote suites then emit `remote_done` (retain `status`, `exit`, `sessions_path`). Assurance conversational agent streams end in `done`; review/read verbs have their own contracts. Always check process exit too: early refusal, invalid plan or dry-run can exit without the normal completion event. Progress is for live display: count only `done`/`failed` completions, retaining child and execution context when step indices repeat. diff --git a/skill-installer/skills/references/setup-and-config.md b/skill-installer/skills/references/setup-and-config.md index b876526..ace5e83 100644 --- a/skill-installer/skills/references/setup-and-config.md +++ b/skill-installer/skills/references/setup-and-config.md @@ -82,7 +82,7 @@ Variables parameterize objectives with reusable values and secrets. Use `{{key}} **Values (0.8.12+):** a string is a value when non-empty; a number is accepted and loaded as its string; a boolean, object or array is not a value. An empty `value` is a declared-but-unfilled key. -**Before a run (0.8.12+):** every `{{name}}` an authored step references must have a value — `run`, `testmd run` and `testrun run` refuse before launching anything otherwise (exit `2`; with `--agent`, one `error` event with `code: "unresolved_variables"` — SKILL.md §3). No bypass flag. +**Before a run (0.8.15+):** `run`, `testmd run` and `testrun run` check every `{{name}}` an authored step references before launching anything. A name with no value is a warning (with `--agent`, one `warning` event with `code: "unresolved_variables"`; SKILL.md §3): the run goes ahead and types the name as written unless a step sets it first. **`assurance.json`:** `kane-cli design tests` writes empty stubs (`{"name":{"value":"","secret":false,"description":"…"}}`) into `{cwd}/.testmuai/variables/assurance.json` for every variable it declares, never a value, and never over a key that already exists in any pool file. Fill those before authoring the designed tests. diff --git a/skill-installer/skills/references/test-manager.md b/skill-installer/skills/references/test-manager.md index 91653cf..65e9030 100644 --- a/skill-installer/skills/references/test-manager.md +++ b/skill-installer/skills/references/test-manager.md @@ -104,7 +104,7 @@ To use the result for subsequent runs, persist with `kane-cli config project ` | Filter candidates by project-relative path regex | — | +| `--match ` | Filter candidates by project-relative path regex | The path is as the OS writes it: `tests/app/` on macOS and Linux, `tests\app\` on Windows. Quote the regex with double quotes in cmd.exe; single quotes are literal there. | | `--tags ` | ANY-match on frontmatter `tags:` (repeatable or comma-separated, case-insensitive) | — | | `--parallel ` | Worker count; each desktop worker gets an isolated Chrome with a fresh temp profile | `1` | | `--on-failure ` | `continue` (run everything) \| `fail-fast` (stop dispatching new members after a failure) | `continue` | @@ -45,16 +45,16 @@ All members must share one org + project. *(0.8.4+)* Members need **not** be aut | `org_mismatch` | Different organisation than the other tests | "Check `kane-cli testmd status ` — it belongs to another org" | | `project_mismatch` | Different project than the other tests | "Run it separately or per-project" | -If **any** member fails preflight, the plan is invalid: nothing runs, exit `2`. Suggest `--dry-run` to preview the plan cheaply before a big run. *(0.8.12+)* Preflight also checks variables: a member whose authored steps reference a `{{name}}` with no value fails with `unresolved_variables` — `testrun run` has no `--variables` flag, so fill the pool file (`.testmuai/variables/*.json`) or the member's own `variables:` frontmatter. +If **any** member fails preflight, the plan is invalid: nothing runs, exit `2`. Suggest `--dry-run` to preview the plan cheaply before a big run. *(0.8.15+)* Preflight also checks variables. A `{{name}}` with no value is a warning: the member stays in the plan (`valid` does not look at variables), and the run types the name as written unless a step sets it first. The check covers every step of the member, replayed ones included. `testrun run` has no `--variables` flag, so fill the pool file (`.testmuai/variables/*.json`) or the member's own `variables:` frontmatter first. ## Remote: the suite as one HyperExecute job (`--remote`) -`kane-cli testrun run --remote` ships the cwd as the job payload, provisions a grid runtime on a **HyperExecute macOS runner** — Chrome for web members, a **virtual Android emulator or iOS simulator** for mobile members — runs every member there as its own headless `testmd run`, and brings the recordings (`output-/`) and the sealed evidence pack back into the project. It works **from any machine** with nothing local but Node and the plugin (no Chrome needed); the account needs a HyperExecute plan with macOS runners and the plugin (`kane-cli plugin install remote-execution`; check with `kane-cli plugin doctor remote-execution`). Auth is a LambdaTest username + access key — an OAuth profile is exchanged automatically. Not the same as `--ws-endpoint`, which attaches a remote browser to a run that still executes locally. +`kane-cli testrun run --remote < /dev/null` ships the cwd as the job payload, provisions a grid runtime on a **HyperExecute macOS runner** — Chrome for web members, a **virtual Android emulator or iOS simulator** for mobile members — runs every member there as its own headless `testmd run`, and brings the recordings (`output-/`) and the sealed evidence pack back into the project. It works **from any machine** with nothing local but Node and the plugin (no Chrome needed); the account needs a HyperExecute plan with macOS runners and the plugin (`kane-cli plugin install remote-execution`; check with `kane-cli plugin doctor remote-execution`). Auth is a LambdaTest username + access key — an OAuth profile is exchanged automatically. Not the same as `--ws-endpoint`, which attaches a remote browser to a run that still executes locally. ```bash -kane-cli testrun run --tags smoke --remote --dry-run # web suite: validate, dispatch nothing -kane-cli testrun run tests/app/ --remote --device-name "Pixel 7" --os-version 14 # Android suite on the grid -kane-cli testrun run tests/ios/ --remote --device-name "iPhone 15" --os-version 17.5 # iOS suite on the grid +kane-cli testrun run --tags smoke --remote --dry-run < /dev/null # web suite: validate, dispatch nothing +kane-cli testrun run --match '^tests/app/' --remote --device-name "Pixel 7" --os-version 14 < /dev/null # Android suite on the grid +kane-cli testrun run --match '^tests/ios/' --remote --device-name "iPhone 15" --os-version 17.5 < /dev/null # iOS suite on the grid ``` - **Always `--dry-run` first.** It runs the normal preflight plus the **remote preflight** and resolves the device against the grid catalog (`kane-cli devices list --target emulator|simulator --remote --agent`) without creating a job. @@ -87,7 +87,7 @@ All typed; stdout; one JSON object per line. **Local completion: `testrun_done`. | `type` | Payload | Notes | |---|---|---| -| `testrun_plan` | `members: [{path, test_id?, tags, failure?}]`, `valid`, `parallel`, `parallel_clamped?` | If `valid: false`, treat as immediate failure — report each member's `failure` reason and stop expecting more events. *(0.8.12+)* `failure: "unresolved_variables"` means a member references a `{{name}}` with no value; one `error` event with `code: "unresolved_variables"` follows the plan (schema in `references/parsing.md`) and lists every such name across members — surface it, do not retry. | +| `testrun_plan` | `members: [{path, test_id?, tags, failure?, unresolved?}]`, `valid`, `parallel`, `parallel_clamped?` | If `valid: false`, treat as immediate failure — report each member's `failure` reason; a `warning` may still follow before the process exits. *(0.8.15+)* `unresolved[]` lists the member's `{{name}}`s with no value; `valid` does not look at it. One `warning` event with `code: "unresolved_variables"` follows the plan (schema in `references/parsing.md`) and the run goes ahead. | | `testrun_start` | `execution_id`, `members` (paths), `parallel` | | | `testrun_member_start` | `path`, `test_id?`, *(0.8.17+)* `session_id`, `log_path` | A saved test started. `log_path` is the absolute path of that test's own event log (see **Each test's own log** below). | | `testrun_member_end` | `path`, `test_id?`, `status`, `duration_s`, *(0.8.17+)* `session_id`, `log_path`, `failure?: {message, step_index?}` | `status` ∈ `passed \| failed \| broken \| interrupted`. `failure` is present when the test did not pass: use it for the "where" and "why" of the failed-tests table. | @@ -129,6 +129,7 @@ for each line: if type === "testrun_done" → capture suite outcome; remote runs keep reading if type === "remote_done" → capture remote status, exit and sessions_path if type === "testrun_plan" && !valid → report offenders, expect exit 2 + if type === "warning" && code === "unresolved_variables" → name the variables with no value (references/parsing.md); the run continues; ignore other warning codes if type === "testrun_member_end" → note per-member outcome if type === "testrun_summary" → capture totals for the rollup else → informational; narrate sparingly @@ -164,7 +165,7 @@ Local suites containing any mobile member require `--parallel 1`; larger values Healing is enabled by default (three shrinking replay windows, then re-authoring of authorable steps). `--no-adaptive-heal` disables it. Retired `--retry`/`--retry-count` only print a notice and have no effect. Replay-only recorded steps retain their recordings even during healing. -NDJSON selection uses stdin, not stdout: run `kane-cli testrun run < /dev/null` for automation launched from a terminal. Dry-run validates a plan, not runtime authentication or browser/device readiness. Always observe process exit, including paths without a normal completion event. +NDJSON selection uses stdin, not stdout: every `kane-cli testrun run` line ends in `< /dev/null` (bash and zsh on macOS, Linux and Git Bash; `< NUL` in cmd.exe; from PowerShell run it through cmd: `cmd /c "kane-cli testrun run … < NUL"`). Check the first stdout line: on 0.8.17+ it is `{"type":"stream_start"…}`. A prose plan there instead means the CLI is in terminal mode: it prints the human view and, after a real run, opens an evidence table that waits until `q` or Esc, so the process never exits on its own. Stop it and rerun with the redirect. An empty stdout with a message on stderr is a usage error: read it. Dry-run validates a plan, not runtime authentication or browser/device readiness. Always observe process exit, including paths without a normal completion event. ### Remote behavior still requiring verification