Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion pstack/.cursor-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "pstack",
"displayName": "pstack",
"version": "0.11.3",
"version": "0.11.4",
"description": "if you want to go fast, go deep first. pstack helps you write less, but higher quality code. rigorous agent workflows you can parallelize with confidence.",
"author": {
"name": "Lauren Tan"
Expand Down
27 changes: 16 additions & 11 deletions pstack/skills/poteto-mode/playbooks/eval.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,26 +2,31 @@

**You own the experiment design. Plan, blind, run, synthesize.**

Evals test how a change affects agent behavior before promoting it: a new skill variant, a structural change, a prompt tweak. The failure mode is the observer effect. An agent that knows it's being evaluated behaves differently, so candidates must run blind.
Evals are blinded, one-shot bakeoffs for deciding whether to promote or reject a change: a new skill variant, a structural change, a prompt tweak. Each trial gets one clean attempt with no feedback or repair. This is not a standing regression suite or a CI merge gate.

**Non-negotiables for blinding:**
**Non-negotiables for blinding and isolation:**

- No `eval`, `test`, `judge`, `experiment`, `rubric`, `score`, `compare`, `benchmark`, `candidate`, or `arena` in any directory, file, or prompt the candidate sees.
- The candidate prompt looks like an organic user request. State the goal, not the meta. "build me a small todo cli" not "show me how you follow the principles chain".
- Candidate prompts must not name a skill, give a skill path, or say to apply/use/follow a skill. Plant skills in the workspace; the candidate discovers them.
- No chain-eliciting cues. Don't ask the candidate to list which skills, principles, or files they applied; that meta-prompt inflates citation behavior. Ask for design notes generally and grade chain-following from code shape, not self-report.
- Sanitize directory and slug names. Use project-shaped names a user might pick, not labels like `candidate-1` or `agent-a`.
- Sanitize directory, slug, and arm names. Use project-shaped names a user might pick, not labels like `candidate-1`, `agent-a`, `control`, or `skill-off`.
- Don't tell the candidate other candidates exist.
- The judge can know it's judging but sees outputs by sanitized label only, never by model name.
- Comparing two variants: one judge scores both sets in a single pass on one scale, blind to which set each came from. Two judge runs with different prompts don't compare, the calibration drifts.
- Start each trial in a fresh workspace and preferably a new session. Clear prior chat and give it no sibling memory. Never plant prior transcripts, judge notes, or sibling outputs.

**Steps:**

1. **Frame.** State what variant is under test and what behavior counts as success. Write the rubric (3-6 concrete criteria) for the judge only. Hold it back from candidates.
2. **Set up sanitized environments.** Per-candidate working dir with the variant in place. Plant any context an organic task would have: a project skeleton, the skills the candidate would naturally read.
3. **Author one organic prompt.** What a user would type. No leakage of what's being measured.
4. **Spawn N parallel candidates** on different models per the **arena** skill's Phase B. Each works in its own sanitized dir; same prompt to each.
5. **Spawn one blinded judge** on a different model family per the **arena** skill's Phase C. Judge sees outputs by sanitized label and the rubric, never a model name.
6. **Verify the chain from transcripts, not self-report.** Read each candidate's local transcript under the active workspace's `agent-transcripts/` directory (the system prompt names this path). Do not glob across `~/.cursor/projects/*/`; that crosses workspace boundaries and reads private chats from unrelated projects. Look at which files each candidate actually opened. Citing a principle is not reading its leaf skill, and reading it is not applying it. Grade chain-following from the files it really read plus the shape of the code, never from the candidate's own claims.
7. **Read every candidate output yourself** end to end. Compare to the judge's verdict. Disagreement means a model is biased or the rubric is ambiguous. Synthesize.
1. **Frame.** State the variant and the promote-or-reject claim. Write a judge-only rubric with 3-6 concrete criteria. Grade task success and the intended behavioral shape. Never make a turn-1 skill load, a particular file read, a citation, or "did the skill trigger?" a pass condition.
2. **Author an organic prompt set.** Include at least one task where the behavior should apply. If the variant changes a description, routing, sticky behavior, or when-to-apply rule, include at least one task where it should not engage and add false-positive cost to the rubric. Write what a user would type. Never name the behavioral tell the rubric grades, and never name or path a skill (if you measure dated headings, do not say "dated note" in the prompt). No other leakage of what is measured. If the task prompt itself is the target, write matched current and proposed versions here; otherwise every arm gets the same prompt.
3. **Build comparison arms.** Before editing, snapshot any prior skill contents the control will need. Variant gets the proposed skill, structure, or prompt. Control gets the current version. For skill presence or content changes, run both a prior-version control and a skill-absent arm unless absence is impossible. Never plant the ablation-target skill into a skill-absent control. Hold the project skeleton, model mix, and every non-target input constant. Controls are other sanitized labels. Promote only when the variant beats the prior control on the rubric without looking worse than absent on false positives.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Final promote bar contradicts ablation

High Severity · Logic Bug

Step 3 now requires evaluating against both a prior-version control and a skill-absent arm, with promotion gated on beating the prior control without looking worse than absent on false positives. However, Step 8 and the 'Reply' section still refer to a singular 'control,' which could lead to the full evaluation criteria from Step 3 being overlooked in the final synthesis.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit e330ed5. Configure here.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FP gate lacks negative prompts

Medium Severity · Logic Bug

Step 3’s promote rule requires the variant not look worse than the skill-absent arm on false positives for skill presence or content changes, but step 2 only authors negative organic prompts and false-positive rubric cost when the variant changes description, routing, sticky, or when-to-apply behavior. Pure content edits therefore hit an FP-vs-absent gate with no prompt set or rubric criteria that measure it.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit e330ed5. Configure here.

4. **Set up isolated trials.** Fresh per-trial workspace with only that arm's variant and organic-task context. Identical project skeleton across arms. A fresh workspace does not clear skills from workspace `.cursor/skills/`, user `~/.cursor/skills/`, or plugin installs: for a skill-absent arm, use a workspace-local isolation (or disable that is restored before any other arm runs). Never apply a shared user/plugin disable that also strips the skill from the variant arm. Preflight resolved sources and fail setup if the skill remains visible on a skill-absent arm or missing on a variant arm. For sticky, mode, description, or other always-on triggers, preflight that the variant reaches the candidate the way production does (reminder in context, description always loaded, and so on). If the harness cannot inject it that way, stop: the bakeoff is invalid for that variant class. Record each trial's workspace path and transcript ID as orchestrator-only metadata. Cheap deterministic preflights aid synthesis only; they never replace the blinded rubric.
Comment thread
cursor[bot] marked this conversation as resolved.
5. **Run 2-3 one-shot trials per prompt and arm.** Launch each runner directly in its recorded workspace with that arm's isolated context. Fan out in parallel with no shared grounding and no candidate-visible files across workspaces. Match model and trial pairings across arms. If the skill ships across models, use at least two model families; matched pairings on one family are not enough. Ask only for the organic task output, not a graft rationale. Missing output fails the trial. No retries, coaching, or repair. When budget binds, prefer 2 trials on fewer models over 1 on many.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Budget clashes with model rule

Medium Severity · Logic Bug

Step 5 newly requires at least two model families when a skill ships across models, while the same step still says that under budget pressure agents should prefer two trials on fewer models. There is no tie-break, so a verbatim reader can collapse to one family with two matched trials, then promote on evidence that does not generalize.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 98957c3. Configure here.

6. **Spawn one blinded judge** on a different model family after every trial finishes. In one pass, score every output by randomized sanitized label against the same rubric. Mark each criterion and output pass or fail. Programmatic checks may filter obvious fails before the judge; they do not decide promote or reject. Do not run the arena pick/graft workflow. This bakeoff ends at arm-level scoring.
7. **Inspect transcripts after scoring to explain how, not to decide pass or fail.** Read only the recorded transcript for each trial from that workspace's transcript directory (normally `~/.cursor/projects/<trial-workspace-slug>/agent-transcripts/`), using the session or transcript ID from setup. Derive the slug from the recorded workspace path. Do not glob across `~/.cursor/projects/*/` or open unregistered workspaces. Transcripts verify isolation and explain the output. They are not a pass gate.
8. **Read the outputs yourself.** At small N, read every output end to end. At large N, read every fail plus a stated random sample of passes; silent skim is not enough. Report pass rates by arm and prompt, then compare with the judge. Promote only when the variant beats the control overall without adding false positives. Otherwise reject. Explain disagreements as judge bias, contamination, or rubric ambiguity.
Comment thread
cursor[bot] marked this conversation as resolved.

**Reply:** variant under test, rubric, per-candidate notes, judge's verdict, your synthesis, and a recommendation for whether to promote the variant.
**Related:** Shipped skills may keep a separate standing regression pack of 5-20 cases. It is distinct from this bakeoff.

**Reply:** variant and control, prompt set, rubric, trial pass rates, per-candidate notes, judge's verdict, your synthesis, and the promote-or-reject decision.
Loading