Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
98 changes: 73 additions & 25 deletions plugins/thread-briefs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,35 +6,75 @@ every thread a short, durable **brief**, generated outside the working chat:

- **goal** — what the thread is actually trying to achieve
- **currentState** — what exists now, including half-done work
- **nextStep** — the single most concrete next action, or empty when nobody owes
the thread one. An open PR, a patch carried on a fork, or a workaround still in
place is owed; open-ended watching is not
- **nextStepActor** — who has to take it: `me`, `agent`, or `other`. The one
judgement the status needs that the prose cannot supply, since "test it and
tell me" and "keep going" read alike
- **blockedOn** — the party or artifact it is waiting on, when someone could go
chase it
- **nextStep** — the most useful next action, if there is one. Descriptive only:
a finished thread may still carry a suggestion here
- **nextStepActor** — who would take it: `me`, `agent`, or `other`
- **blockedOn** — the party or artifact outside the thread it is waiting on, when
someone could go chase it
- **constraints** — facts learned in the thread that would break a naive re-plan
- **title** — a 4–6 word name for the work, which can optionally replace bb's
own thread title

Plus a derived **stage** (discovery / planning / implementation / review) and
**status** (working / waiting-on-me / waiting-on-other / done) — where `working`
comes from bb's live thread state and the other three from the brief.

The two are orthogonal — stage says how far the work has got, status says who
owes the next move — but they are answered by one model call over one transcript,
and nothing in that call holds them to agreeing. So the parser reconciles them:
an empty `nextStep` means nobody owes the thread an action, which is only true
once the work is made, so a `stage` of `implementation` beside one is read as
`review`. Without it an agent's closing summary of what it built lands as
"Implementation — Done". A pinned stage is exempt; a pin is returned as given.
Plus a **stage** (discovery / planning / implementation / review) and a
**status** (working / waiting-on-me / waiting-on-other / done). `working` comes
from bb's live thread state; the other three statuses and the stage are asked of
the model directly.

The status question is one sentence: *assume you do whatever the thread asks of
you — is the task then finished, or does the thread have more to do?* Yes is
`done`, even when a step is left that only you can take ("approve PR #12", "run
the rollout check"), because nothing more will happen in the thread either way.
No, because your answer or go-ahead starts more work here, is `waiting-on-me`.
No, because something outside the thread has to act first, is
`waiting-on-other`, and names that thing in `blockedOn`.

### Why status is asked for, not derived

Status used to be derived: an empty `nextStep` and an empty `blockedOn` meant
done. That made Done a side effect of the model leaving two strings blank, while
the same prompt asked for "the single most concrete next action" — and a model
asked for one finds one. Four rules accumulated in the prompt pulling the
boundary one way and the other, and on the bb-dylan server 11 of 21 Done threads
were Done only because someone had pinned them by hand. Reading each of those 11
against its transcript gave three causes:

- an offer or optional check recorded as the step ("want me to file an issue?",
"if you want the hash confirmed, run…");
- a `nextStep` or `blockedOn` carried over from the previous brief, which was fed
back as the starting point and outlived the turns that retired it;
- a step that was the user's, done outside the thread.

So the model now answers the status itself, with one definition and one example
per value; `nextStep` no longer decides anything; and the previous brief fed back
into the next summary carries only the fields that should hold still — title,
goal, currentState, constraints. `nextStep` and `blockedOn` are re-read from the
transcript every time.

The parser applies one rule on top, which holds whatever the model thinks: a
brief that names a `blockedOn` is never `done`. An unreadable status falls back
to `waiting-on-me`, the reading whose mistake is cheap. There are no other
field-against-field corrections — stage and status are stored as answered.

A brief written before this change has no stored status and keeps the old
derivation until its thread is next summarized. They are deliberately not
re-summarized in bulk: every old idle thread that now read done would be past
the archive threshold already, and the next sweep would take them all at once.

### Measuring a prompt change

`eval/` holds the harness that produced the numbers behind this. `export.ts`
freezes threads from a running server into fixtures (outline, last message, the
stored brief); `run.ts` runs a checkout's prompt over them against the live model
and reports agreement with hand-given labels, false Dones and missed Dones, with
and without the previous brief fed back. Point `--src` at a `git worktree` of
`main` to score the old prompt against the same set. Fixtures are real
transcripts and bb-plugins is public, so they live outside the repository.

Either can be pinned by hand in the Brief panel, anchored to the thread's
activity cursor so the pin retires on the next real turn. The status pin is what
closes a thread whose next step was carried out somewhere the transcript cannot
see — "reload a client and confirm the panel opens" leaves nothing for a summary
to read, so the derivation would say `waiting-on-me` forever.
see — a go-ahead you gave in another thread, a PR you merged on github.com —
leaves nothing for a summary to read.

## Install

Expand Down Expand Up @@ -159,7 +199,7 @@ takes it over; turning grouping off deletes the three sections and restores the
sidebar preferences it changed, but cannot put a hand-made placement back.

**The side panel** — a **Brief** tab holding the full five fields (empty ones
are skipped), the derived status, when it was last summarized, status and stage
are skipped), the status, when it was last summarized, status and stage
controls for the manual overrides, and Re-summarize. The **Brief** button in the thread
header opens it; so does the panel's own new-tab launcher, under Actions. It
works the same on mobile and desktop — on a compact viewport the host reveals
Expand Down Expand Up @@ -325,8 +365,8 @@ obvious:
a statement that the model was right, and pinning it there would leave a pin
that does nothing until it silently expires.
- Dragging **out of** Done pins `waiting-on-me` rather than clearing the status
pin, because a done reading can come from the derivation as well as from a pin —
and clearing in that case would hand the card straight back to a derivation
pin, because a done reading can come from the model as well as from a pin —
and clearing in that case would hand the card straight back to a model reading
that still says done, snapping it into the column you just dragged it out of.
- **No stage** is not a drop target in either direction.

Expand Down Expand Up @@ -385,6 +425,14 @@ The [grey ring](#where-briefs-show-up) is the warning: both go through the same
sweep will take once the second threshold passes. A day of grey is the notice
period.

Done is not "nothing left anywhere": a done thread may still carry a step that is
yours alone, such as approving a PR, and it is archived on the same clock. That
is deliberate. The step stays on the card in the Done section for two days and
the hover label says "archiving soon" for the second; a thread archived anyway is
one click from back, and un-archiving it is final (below). Waiting on you would
have kept it in the sidebar indefinitely, which is the failure this status was
rewritten to fix.

Every other rule is a reason *not* to archive, which is the right default for a
sweep that runs unattended — a thread wrongly left in the sidebar costs a glance,
a thread wrongly archived costs a search for something you believe you left on
Expand Down Expand Up @@ -413,7 +461,7 @@ so a sweep running up to an hour late is invisible.
| Concern | Mechanism |
| --- | --- |
| Trigger | `bb.events.on("thread.idle")` + a per-thread quiet-period debounce, with a `*/10 * * * *` sweep as the backstop. A thread with no brief yet skips the quiet period, and is summarized from `thread.active` as well — mid-turn, so a long first turn is not spent briefless |
| Summarizer input | `threads.conversationOutline()` head + tail with the middle elided, `threads.output()` for the last message in full, and the previous brief |
| Summarizer input | `threads.conversationOutline()` head + tail with the middle elided, `threads.output()` for the last message in full, and the previous brief's title, goal, currentState and constraints |
| Storage | `bb.storage.kv`, one row per thread at `brief:<threadId>` |
| Sidebar glyph | a content script's `experimental_setThreadRowStatus`, fed by an `experimental_appOverlay` that owns the rpc + realtime subscription |
| Ring artwork | `app.experimental_icons.register`, one inline SVG per stage plus the done ring, in every palette colour, plus one grey done ring — since a row status takes an icon *name* and not a component, every combination has to be registered at init, before any project is known |
Expand Down
4 changes: 2 additions & 2 deletions plugins/thread-briefs/board.ts
Original file line number Diff line number Diff line change
Expand Up @@ -646,8 +646,8 @@ export function countByStatus(
/**
* The one thing the status badge cannot say.
*
* `deriveStatus` collapses a next step the *agent* could take by itself into
* `waiting-on-me`, because the nudge is ours to give — so two cards reading
* `waiting-on-me` covers a thread the *agent* could carry on by itself as well
* as one that needs your answer, because the nudge is ours to give — so two cards reading
* "Waiting on you" can want completely different amounts of work from you. On a
* board, where the whole task is choosing between them, that difference is worth
* a word.
Expand Down
63 changes: 46 additions & 17 deletions plugins/thread-briefs/brief.test.ts
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
import { describe, expect, it } from "vitest";
import {
deriveStatus,
legacyStatus,
effectiveStage,
effectiveStatus,
idleFor,
Expand Down Expand Up @@ -40,47 +40,76 @@ const stored = (overrides: Partial<StoredBrief> = {}): StoredBrief => ({
...overrides,
});

describe("deriveStatus", () => {
describe("effectiveStatus", () => {
it("reports the status the summarizer judged", () => {
// Done with a suggestion still standing: nextStep no longer decides.
expect(effectiveStatus(stored({ modelStatus: "done" }))).toBe("done");
expect(
effectiveStatus(
stored({ modelStatus: "waiting-on-me", fields: { ...stored().fields, nextStep: "" } }),
),
).toBe("waiting-on-me");
});

it("reads a brief written before status was asked for the old way", () => {
// Not backfilled: such a brief keeps its reading until the thread is
// next summarized.
expect(effectiveStatus(stored())).toBe("waiting-on-me");
expect(
effectiveStatus(stored({ fields: { ...stored().fields, nextStep: "" } })),
).toBe("done");
});

it("lets a pin in force win over the model", () => {
expect(
effectiveStatus(
stored({ modelStatus: "waiting-on-me", statusOverride: "done", statusOverrideSeq: 50 }),
),
).toBe("done");
});
});

describe("legacyStatus", () => {
it("reports done only when nothing is outstanding and nothing blocking", () => {
expect(deriveStatus({ nextStep: " ", blockedOn: "" })).toBe("done");
expect(legacyStatus({ nextStep: " ", blockedOn: "" })).toBe("done");
});

it("does not call a blocked thread done, whatever the next step says", () => {
// The prompt promises a non-empty nextStep whenever anything is
// outstanding. This is the guard for when it does not deliver one.
expect(deriveStatus({ nextStep: "", blockedOn: "Review from Dylan" })).toBe(
expect(legacyStatus({ nextStep: "", blockedOn: "Review from Dylan" })).toBe(
"waiting-on-other",
);
expect(deriveStatus({ nextStep: " ", blockedOn: " Upstream fix " })).toBe(
expect(legacyStatus({ nextStep: " ", blockedOn: " Upstream fix " })).toBe(
"waiting-on-other",
);
});

it("reports waiting-on-other when something is blocking", () => {
expect(
deriveStatus({ nextStep: "Merge it", blockedOn: "Review from Dylan" }),
legacyStatus({ nextStep: "Merge it", blockedOn: "Review from Dylan" }),
).toBe("waiting-on-other");
});

it("falls back to waiting-on-me for unfinished, unblocked work", () => {
expect(deriveStatus({ nextStep: "Pick an approach", blockedOn: "" })).toBe(
expect(legacyStatus({ nextStep: "Pick an approach", blockedOn: "" })).toBe(
"waiting-on-me",
);
});

it("does not need a trailing question to report waiting-on-me", () => {
// An idle thread with work left needs a human look either way, so the
// question signal no longer changes the outcome.
expect(deriveStatus({ nextStep: "Keep going", blockedOn: "" })).toBe(
expect(legacyStatus({ nextStep: "Keep going", blockedOn: "" })).toBe(
"waiting-on-me",
);
});
});

describe("deriveStatus with an actor", () => {
describe("legacyStatus with an actor", () => {
it("treats an external actor as blocked even with no blockedOn text", () => {
expect(
deriveStatus({
legacyStatus({
nextStep: "Land the upstream PR",
blockedOn: "",
nextStepActor: "other",
Expand All @@ -90,7 +119,7 @@ describe("deriveStatus with an actor", () => {

it("reports waiting-on-me for a step only the user can take", () => {
expect(
deriveStatus({
legacyStatus({
nextStep: "Try it and say whether the glyph looks right",
blockedOn: "",
nextStepActor: "me",
Expand All @@ -100,7 +129,7 @@ describe("deriveStatus with an actor", () => {

it("reports waiting-on-me for a step the agent could take, since the nudge is ours", () => {
expect(
deriveStatus({
legacyStatus({
nextStep: "Keep porting the remaining call sites",
blockedOn: "",
nextStepActor: "agent",
Expand All @@ -110,21 +139,21 @@ describe("deriveStatus with an actor", () => {

it("lets done win over any actor, so a finished thread is never a prompt", () => {
expect(
deriveStatus({ nextStep: "", blockedOn: "", nextStepActor: "other" }),
legacyStatus({ nextStep: "", blockedOn: "", nextStepActor: "other" }),
).toBe("done");
});

it("preserves the actor-free behaviour when the actor is absent", () => {
// Every brief written before this field existed lands here.
expect(deriveStatus({ nextStep: "Keep going", blockedOn: "" })).toBe(
deriveStatus({
expect(legacyStatus({ nextStep: "Keep going", blockedOn: "" })).toBe(
legacyStatus({
nextStep: "Keep going",
blockedOn: "",
nextStepActor: undefined,
}),
);
expect(
deriveStatus({
legacyStatus({
nextStep: "Keep going",
blockedOn: "",
nextStepActor: undefined,
Expand Down Expand Up @@ -213,7 +242,7 @@ describe("status overrides", () => {
statusOverrideSeq: 50,
lastActivitySeen: 50,
});
expect(deriveStatus(brief.fields)).toBe("waiting-on-me");
expect(legacyStatus(brief.fields)).toBe("waiting-on-me");
expect(effectiveStatus(brief)).toBe("done");
expect(isStatusOverrideStale(brief)).toBe(false);
});
Expand Down
68 changes: 18 additions & 50 deletions plugins/thread-briefs/brief.ts
Original file line number Diff line number Diff line change
Expand Up @@ -57,73 +57,41 @@ export function isStatusOverrideStale(stored: StoredBrief): boolean {

/**
* The status this brief reports: the manual one while it holds, otherwise the
* derivation over the brief's own fields.
* one the summarizer judged.
*
* The override sits *in front of* {@link deriveStatus} rather than editing the
* fields it reads, because `renderTranscript` feeds the previous brief into the
* next summary as a starting point: a `nextStep` blanked in storage would
* simply be written back, where a pin is a separate fact the summarizer never
* sees and cannot undo.
*
* It exists for the one thing the derivation cannot see. A `nextStep` addressed
* to you and carried out *outside the thread* — reload a client, check a
* rollout, confirm a glyph — leaves no trace in the transcript, so no summary
* can retire it and re-summarizing reads the same unresolved instruction back.
* That thread is `waiting-on-me` forever unless you can say otherwise.
* The manual pin exists for the one thing no summary can see: a step carried
* out somewhere the transcript does not reach. The summarizer is asked whether
* the task is finished *assuming* you do what the thread asks of you, so a step
* that is only yours already reads done; the pin is for the rest — a thread
* waiting on your go-ahead that you settled elsewhere, or a reading you simply
* disagree with.
*/
export function effectiveStatus(stored: StoredBrief): StoredBriefStatus {
const override = stored.statusOverride ?? null;
if (override !== null && !isStatusOverrideStale(stored)) return override;
return deriveStatus({
nextStep: stored.fields.nextStep,
blockedOn: stored.fields.blockedOn,
nextStepActor: stored.fields.nextStepActor,
});
return stored.modelStatus ?? legacyStatus(stored.fields);
}

/**
* Status as far as a *stored* brief can tell. Mechanical rather than a model
* judgement, so it stays right between summaries:
*
* - nothing to do and nothing blocking → the work is done
* - blocked on something → waiting on someone else
* - otherwise → waiting on me
*
* `done` requires *both* fields empty. The summarizer's prompt already promises
* a non-empty `nextStep` whenever anything is outstanding, but testing
* `blockedOn` here makes "a blocked thread is not done" a guarantee of this
* code rather than of the prompt, so one wayward summary cannot put a green
* tick on a thread that is waiting for a review.
*
* `nextStepActor` is the one input the model has to judge: a next step only we
* can take ("test it", "decide X", "reply to Y") reads the same in prose as one
* the agent could take unprompted. An actor of `other` therefore means waiting
* on someone else even when the summarizer named no `blockedOn`.
*
* `waiting-on-me` is the fallback because an idle thread with unfinished work
* needs a human look by default — whether or not the agent's last turn happened
* to end in a question, and whether or not we know the actor.
* The status of a brief written before the summarizer was asked for one, read
* from its fields the way it used to be: nothing to do and nothing blocking is
* done, a blocker or an outside actor is waiting on someone else, and anything
* else is waiting on you.
*
* `working` is deliberately absent: it is live thread state, not a property of
* a brief, so it is applied per row by {@link rowDecoration}. The return type
* says so — narrowing to the stored statuses is what lets a caller that needs
* one, like the refresher's `writtenForStatus`, take this value without a cast.
* Only for old rows. Such a brief stays on this reading until its thread is
* next summarized, which writes a `modelStatus` that replaces it.
*/
export function deriveStatus(args: {
export function legacyStatus(fields: {
nextStep: string;
blockedOn: string;
nextStepActor?: NextStepActor | undefined;
}): StoredBriefStatus {
const nextStep = args.nextStep.trim();
const blockedOn = args.blockedOn.trim();
const nextStep = fields.nextStep.trim();
const blockedOn = fields.blockedOn.trim();
if (nextStep === "" && blockedOn === "") return "done";
if (blockedOn !== "" || args.nextStepActor === "other") {
if (blockedOn !== "" || fields.nextStepActor === "other") {
return "waiting-on-other";
}
// `agent` — an idle thread the agent could carry on by itself — has no status
// of its own yet, and collapses into waiting-on-me because the nudge is ours
// to give. If that turns out to be a common bucket in practice it earns its
// own status then, rather than being guessed at now.
return "waiting-on-me";
}

Expand Down
Loading
Loading