Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 9 additions & 1 deletion .cargo/mutants.toml
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# Functions the mutation gate cannot judge, because `cargo test` cannot reach
# them. Matched against the mutant names that `cargo mutants --list` prints.
#
# EXCLUSIONS: 60
# EXCLUSIONS: 61
#
# That number is checked by `scripts/test.sh`, so adding an entry means editing
# this line too. The point is not the count, it is that the list only ever grows
Expand Down Expand Up @@ -163,6 +163,13 @@
# `is_interview_participant`, `candidate_video_frames_go_to_gemini`,
# `should_send_video_frame`, `frame_to_rgba` and `encode_rgba_jpeg`.
#
# `handle_board_event` is the whiteboard's counterpart to `handle_media_event`:
# it takes a `ByteStreamOpened` room event, and the reader that event carries
# sits in a `TakeCell` whose constructor the SDK keeps private, so no test can
# build the one event it acts on. What it decides is tested where it is
# decided: `board_stream_refusal`, `strokes_from_attributes`, and `drain`,
# which reads any stream of chunks rather than the SDK's reader alone.
#
# `create_interview`'s two comparisons on the insert rowcount are the one pair
# here that is reachable, tested, and still unkillable, so they are named
# individually rather than by function: every other mutant in it is caught.
Expand Down Expand Up @@ -322,6 +329,7 @@ exclude_re = [
"record_live_usage",
"drain_live_usage",
"handle_media_event",
"handle_board_event",
"attach_audio",
"next_audio_frame",
"next_video_frame",
Expand Down
13 changes: 12 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,8 @@ LiveKit tokens, and runs the interviewer agent.
│ · editor + syntax colors ├─────────────────────▶│ (SFU) │
│ · problem panel, timer │ data channel └─────┬───────────────┘
│ · test runners │ code_update, control, │
│ · report + history │ test_results, report │
│ · whiteboard │ test_results, report, │
│ · report + history │ board_image │
└────────────┬───────────────┘ ▼
│ ┌──────────────────────────────────┐
│ /api/* │ Rust agent (LiveKit runner) │
Expand All @@ -33,6 +34,16 @@ the agent receives structured code rather than editor screenshots. Python and
JavaScript run locally; C, C++, and Java run through Compiler Explorer, so
source code leaves the browser for those three.

The lobby also offers a whiteboard interview, which takes the same problem bank
and the same six steps and swaps the editor and the test runner for a board.
Nothing runs: the candidate draws their examples and traces one by hand. The
board is exported as an image a moment after each stroke settles and reaches
the interviewer over its own byte stream on the same data channel, and
`read_board` puts the latest one back in front of it on request. The final
board is attached to the report request, so the reviewer grades the drawing
rather than an empty editor, and the recording keeps the drawing as the
strokes that made it, which is what lets the replay redraw any moment of it.

Audio and code snapshots stay in memory unless [recording](#recording) is
enabled, which is off by default. Candidate video reaches Gemini only with
`CODETRIAL_GEMINI_CANDIDATE_VIDEO_ENABLED=true`. Face-presence analysis runs in
Expand Down
3 changes: 2 additions & 1 deletion docs/interview-contract-versions.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,10 +8,11 @@ can select it.

## The active bundle

Bundle 25: live prompt 17, report prompt 15, rubric 1, report schema 2.
Bundle 26: live prompt 18, report prompt 16, rubric 1, report schema 2.

| Bundle | Introduced |
|---|---|
| 26 | Whiteboard interviews: the live prompt is written for the surface the candidate works on, so a whiteboard session is told it has no editor and no test runner, is given the six steps as drawn work ending in the complexity of the approach on the board, is offered `read_board` in place of `read_editor`, and asks a candidate whose speech stays unclear to write it on the board rather than as a code comment; `board_snapshot` joins the evidence sources and is the only one besides candidate speech a whiteboard session may record, while an editor session may not record it at all; the phases about written work are gated on strokes on the board rather than on characters in the editor. The report prompt follows the same surface: a whiteboard review is sent labeled images of every completed REACTO phase plus a changed final board, so clearing the live surface does not erase earlier evidence; it is told that nothing ran and that the Coding, Test and Optimizations phases were a hand trace, the cases named against the drawing, and the complexity they confirmed, and cites the board where the other cites the code and the test account. The rubric and the report schema are unchanged, so a report from either surface is scored the same way and against the same ten phases. |
| 25 | Candidates can keep the floor while thinking, reclaim it during a reply, and yield it early. Explicit spoken requests for thinking time in English, including one that follows an answer in the same sentence or is asked as a question, suppress generated replies and automatic nudges until the candidate speaks again or chooses to continue. A hold ends on its own at the five-minute warning, at the round transition, and after two silent minutes with one brief check-in; the interviewer is told that anything it said during the hold was not heard. A Continue within ten seconds of the last one releases the hold without a reply of its own. Thinking keeps editor, microphone and test evidence live, gives the interviewer test runs and edits as context it does not answer, and never extends the deadline. The default endpointing window is three seconds, and the page shows it filling while the candidate is silent; yielding ends the audio stream so the interviewer replies without waiting it out. |
| 24 | The Live main instructions drop repeated explanations and illustrative examples and keep every timer, round, evidence-source and hint restriction. The greeting answers only the platform's startup request, and missing history, a compression or a tool result is not a new interview. `end_interview` is called silently, before any acknowledgment or goodbye, and the platform supplies the closing. A cut `read_editor` page or a checkpoint excerpt does not show the whole buffer, so an implementation or technique is not called absent before the named lines are read. The `read_editor` description asks for only the code the current question needs that nothing has shown, from a known relevant line rather than a refill of the whole editor. The greeting no longer repeats the exercise's title and brief, which THE EXERCISE already carries and the greeting now points at; the framework headers drop a scoring premise the disclosure rule already covers; test-run reactions and the earlier-steps reminder state their rule once, more briefly; and the `end_interview` description no longer restates the instruction it sits beside. With a configured compression window, a silent checkpoint rebuilt from local state follows a detected cut: the chosen language, the current round, the evidence, a bounded transcript that keeps a long behavioral round's opening, a bounded test report and, in the coding round, the editor's opening and ending. Its next step applies to the next candidate input, not to the checkpoint itself. Omission alone does not close a behavioral round, repeat its question or establish that its follow-up is unused, and a refusal or request to finish supplies no STAR evidence. Under the same window, editor, hint and evidence tool answers carry the latest unanswered candidate utterance as quoted historical data, never as a new turn. |
| 23 | A candidate who hides the worked examples in the preflight sends `hideExamples` with the token request, and the live prompt then says no examples are on their screen: the interviewer never points them at one, says a clarification or hint clue that mentions an example with a case they proposed or one of its own, and in the Example step asks for their ordinary and boundary cases before offering a small example once they have tried or are stuck. A session that does not hide them gets the live prompt unchanged. |
Expand Down
25 changes: 23 additions & 2 deletions docs/recording-contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -649,6 +649,7 @@ server, is the ordering.
|---|---|---|
| `transcript` | what was said | no |
| `editor` | the code and its language | yes |
| `board` | what was drawn since the last one | no |
| `tests` | a run's results | no |
| `stage` | the clock and the interview phase | yes |
| `avatar` | what Jim is doing | yes |
Expand Down Expand Up @@ -839,6 +840,7 @@ so a producer from a later deploy does not stop a recording.
|---|---|---|
| `stage` | `{title, meta, remainingSeconds}` | the problem heading and the clock |
| `editor` | `{code, language}` | the code panel, as text |
| `board` | `{ops, checkpoint?}`, each op `{op: "stroke", color, width, points}` or `{op: "undo" \| "redo" \| "clear"}`; `checkpoint` is a completed REACTO phase id | the whiteboard, redrawn from every op so far, with completed phases named in the replay |
| `tests` | `{passed, failed, total}` | one line, red if anything failed |
| `avatar` | `{state}`, one of `speaking`, `thinking`, `listening` | Jim's expression and label |
| `transcript` | `{speaker, text}` | nothing here; the replay page renders it |
Expand All @@ -854,9 +856,28 @@ already arrived. A `413` or a `404` stops the producers for the rest of the
interview: over quota and withdrawn consent both mean everything after this is
refused.

The board is the one kind that does not restate itself, and that is what makes
a whiteboard interview replayable at all. One board exported as an image is
over a hundred kilobytes, which is past the per-payload ceiling on its own and
would spend the whole per-interview budget on a handful of frames; the same
board as the strokes that drew it is a few kilobytes and arrives as operations,
so any moment of the interview can be redrawn rather than the few that could be
photographed. `web/whiteboard.js` is the one model: the candidate draws on it,
the replay page and this template rebuild from it, and a stroke it refuses
while drawing is a stroke it refuses coming back off the wire.

When the interviewer banks a REACTO phase, the browser adds a board event even
if no stroke changed. That event names the phase and therefore freezes the
current point in the operation journal for review. The same moment is exported
as a JPEG with the phase id in its byte-stream header. The agent retains one
image per phase for the report model, so clearing the live board cannot erase
the Example or Approach evidence that came before it. Replay stores no JPEG:
it rebuilds each checkpoint from the operations it already has.

Cadence is where the per-interview budget goes. The editor rides the debounce
the agent's `code_update` already uses; the transcript is one event per spoken
turn rather than per chunk; the clock is restated every fifteen seconds, because
the agent's `code_update` already uses; the board rides the same settle that
sends the interviewer their image, batched so no event outgrows the per-payload
ceiling; the transcript is one event per spoken turn rather than per chunk; the clock is restated every fifteen seconds, because
every second would be twenty-seven hundred events for a number the viewer can
read off the video; and the interviewer's state is sent on the change rather
than on the participant event that happened to carry it.
Expand Down
76 changes: 74 additions & 2 deletions scripts/browser-check.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -36,8 +36,9 @@ function scenarioTitle(problemId) {
return scenario(problemId).title;
}

function interviewUrl(problemId) {
return `${process.env.BASE_URL}/interview?problem=${scenario(problemId).page}&duration=20`;
function interviewUrl(problemId, mode) {
const surface = mode ? `&mode=${mode}` : "";
return `${process.env.BASE_URL}/interview?problem=${scenario(problemId).page}&duration=20${surface}`;
}

async function checkEditorNewlines(page) {
Expand Down Expand Up @@ -396,6 +397,70 @@ function stopProcessGroup(child) {
}
}

/// The whiteboard interview, as far as a run with no credentials can take it.
///
/// The media gate is the assertion, and it is a stronger one than it looks.
/// `init()` builds the board and only then starts the preflight, so anything
/// that throws on the way up stops the page where it stood: the gate sits at
/// "Starting camera and microphone..." for ever and the browser never asks for
/// either device. That is what a `const` still inside its temporal dead zone
/// did here, and every check that reads source text stayed green through it,
/// because the source was right and the order it ran in was not.
async function checkWhiteboardInterview(page, pageErrors) {
const before = pageErrors.length;
await page.goto(interviewUrl("two-sum", "whiteboard"), {
waitUntil: "domcontentloaded",
});

// The board first, and the order is the point. `init()` builds it and then
// starts the preflight, so a page that threw on the way up leaves the gate
// disabled for ever and Playwright reports that as a two-minute click
// timeout on a button nobody can place. The pens are built in the same
// function, one line before the preflight, so asking for them first turns
// that into a sentence naming what stopped.
try {
await page.locator("#board-pens button").nth(3).waitFor({ timeout: 15000 });
} catch {
const status = await page.locator("#audio-check-status").textContent();
throw new Error(
`the board never finished building, so init() stopped before the media preflight it runs next: the gate says ${JSON.stringify(status)}`,
);
}
await page.locator("#board").waitFor({ state: "visible" });
const pens = await page.locator("#board-pens button").count();
if (pens !== 4)
throw new Error(`the board offered ${pens} pens rather than 4`);
await clearMediaGate(page);
const clear = page.getByRole("button", { name: "Clear board" });
if (!(await clear.isDisabled()))
throw new Error("an empty board offered to clear itself");
const bounds = await page.locator("#board").boundingBox();
if (!bounds) throw new Error("the visible board had no drawing bounds");
await page.mouse.move(bounds.x + 40, bounds.y + 40);
await page.mouse.down();
await page.mouse.move(bounds.x + 90, bounds.y + 90);
await page.mouse.up();
if (await clear.isDisabled())
throw new Error("the board could not be cleared after drawing");
await clear.click();
if (!(await clear.isDisabled()))
throw new Error("clearing the board left it non-empty");
if (await page.getByRole("button", { name: "Undo" }).isDisabled())
throw new Error("a cleared board could not be restored with Undo");
// The editor is removed rather than hidden, so its absence is what says the
// page understood which interview it is holding.
if (await page.locator(".editor-panel").count()) {
throw new Error("the editor panel survived into a whiteboard interview");
}
const raised = pageErrors.slice(before);
if (raised.length) {
throw new Error(`the whiteboard interview raised:\n${raised.join("\n")}`);
}
console.log(
"whiteboard: the board drew, cleared, and opened the media gate behind it",
);
}

/// The interview page gates the room join on local media, with no bypass, so
/// every run clears it the way a candidate would. Fake browser devices drive
/// the level meter.
Expand Down Expand Up @@ -690,6 +755,12 @@ async function isolateRustAgent(
}
});
page.on("pageerror", (error) => consoleErrors.push(String(error)));
// The same errors again, on a list of their own. They belong in
// `consoleErrors` for the diagnostics that quote it, and a flow that wants
// to assert nothing threw cannot ask that list: offline practice warns
// there legitimately, and so does the avatar when its model is absent.
const pageErrors = [];
page.on("pageerror", (error) => pageErrors.push(String(error)));
// "status of 500" with no URL is not a diagnosis, so pair every failing
// response with the thing that was being fetched.
page.on("response", (response) => {
Expand Down Expand Up @@ -1013,6 +1084,7 @@ async function isolateRustAgent(
}

if (mode === "offline") {
await checkWhiteboardInterview(page, pageErrors);
await page.goto(interviewUrl("two-sum"), {
waitUntil: "domcontentloaded",
});
Expand Down
18 changes: 18 additions & 0 deletions scripts/gen-wire-fixtures.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -368,6 +368,23 @@ async function integrityChain() {
return events;
}

// The board's stream header, which is the one message the browser sends that
// is not a data packet: LiveKit chunks the JPEG itself, and what the two sides
// have to agree on is the topic it arrives under and the attributes the agent
// reads off it. The sizes are the ones a real board produces.
function boardCases() {
return [
{ name: "first board", options: lib.boardStreamOptions(1, 3, 21_504) },
{
name: "an approach checkpoint",
options: lib.boardStreamOptions(17, 214, 96_318, "algorithm"),
},
// A board that was cleared: no strokes, and still a board, because the
// interviewer has to see that what they were asked about is gone.
{ name: "cleared board", options: lib.boardStreamOptions(18, 0, 4_096) },
];
}

// Imported rather than restated: a hardcoded list here would be a third place
// to disagree with. The constant, not `languagesFor`, because which tabs a
// given judge offers is a UX choice and this is the whole set src/agent.rs has
Expand All @@ -381,6 +398,7 @@ const files = {
"control.json": { topic: lib.topics.control, cases: controlCases() },
"test-results.json": { topic: lib.topics.tests, cases: testResultsCases() },
"integrity-chain.json": await integrityChain(),
"board-stream.json": { topic: lib.topics.board, cases: boardCases() },
};

const check = process.argv.includes("--check");
Expand Down
Loading