Conversation
|
Need some time for testing and self-review; will mark this PR as ready for review once it's set. |
|
Perhaps we will need to record the state of the whiteboard at each stage so that we can provide more information to the user at the end of the interview. |
|
Additionally, pseudocode should be graded separately rather than being combined with optimization like it is currently. |
|
On the concern from #65 about whether Gemini can read a drawn board reliably: one option worth considering is Excalidraw as the board layer. The drawing tools are not the point; the data model is. An Excalidraw scene is a list of typed elements, and boxes, arrows (bound to the boxes they connect) and text keep their actual strings. So next to the JPEG, the agent can get a plain-text summary of the board, e.g. I checked whether it fits the vendoring rules, following the
The costs:
Only the board layer changes. The byte stream, |
|
Hi, @alanhc Thanks for the discussion! Excalidraw was also brought up in the discussions under the Facebook post back then. After looking into it, I think it seems like a solid option to consider, but I also have a few thoughts. If our goal is to simulate an actual interview environment as closely as possible, would Excalidraw make things too convenient for the user? After all, in a real interview, you only have a marker and a physical whiteboard. That raw experience is precisely what we want to deliver in this mode—allowing users to practice explaining their problem-solving approach and thought process purely through drawing on a blank board. Currently, a tester, @Eason0729 , has tested this feature (the non-Excalidraw version) and provided a lot of feedback. Here is a brief summary of the key concerns raised from the testing:
For now, I'll focus on addressing the points raised in the feedback first. As for Excalidraw, we can discuss it further as a potential future improvement. Overall, I actually think this is quite a promising direction. However, my primary concern is that users should have an experience that feels as close to reality as possible; we probably shouldn't compromise on that just to make recognition easier for the model. WDYH |
88b2724 to
e497263
Compare
A whiteboard interview takes the same problem bank and the same six steps and swaps the editor and the test runner for a board. Nothing runs, so correctness is what the candidate can defend by tracing an example across their own drawing. The board travels as a LiveKit byte stream rather than on a data topic, because one board is tens of kilobytes and a data packet carries fifteen, and it reaches Gemini the way a camera frame does, as a realtime image. A tool response is JSON and cannot carry a picture, so read_board asks for the board to be sent again and answers with what an image cannot say: how much is on it and how long ago it was drawn. Bundle 21 is the live prompt following the surface. A whiteboard session is told it has no editor and no test runner, is offered read_board in place of read_editor, and records board_snapshot where the other records an editor snapshot or a test event; neither may record the other's source. The phases about written work are gated on strokes instead of on characters, and Test on the cases named against the drawing instead of on a run, or an interview with no editor would be refused the second half of its own flow. The report still reads the transcript alone and the replay still carries no board.
The whiteboard was live but nowhere else: the report was written from a transcript and an editor nobody opened, the recording kept no drawing, and the checklist beside the timer called step four Coding while the interviewer was asking the candidate to trace. The board now rides the report request as an inline image, ahead of the brief and on every repair, and the brief says what it is looking at: nothing ran, so correctness is the trace the candidate walked, and the phases keep their names with Coding meaning that trace, Test the cases named against the drawing, and Optimizations the complexity they confirmed. The system instruction is left as it is, so every report call still shares its prefix, and the brief tells a whiteboard reviewer how to read the rules that speak of code. That is report prompt 14, inside the same bundle 21, and the reviewer is told when no board arrived rather than sent looking for an attachment that is not there. The recording keeps the drawing as the operations that made it. One board as a JPEG is past the per-event ceiling on its own and would spend the whole per-interview budget on a handful of frames; the same board as strokes is a few kilobytes, so the replay page and the recording template redraw any moment of the interview with the module the candidate drew on, rather than the few moments a photograph could afford.
`init()` runs at the top of interview.js and calls `initWhiteboard`, which reached for a `const` declared eight hundred lines further down: still inside its temporal dead zone, so it threw. The editor panel was already gone and the board already shown, which is why the page looked right, but `init()` stopped there and never reached the media preflight it runs next. The gate sat on "Starting camera and microphone..." for ever, with the browser never having asked for either device, and no whiteboard interview could be started at all. The state moves up with the rest of the module's, where web/app.js keeps its own for this exact reason. Nothing reading source text could have caught it, because the source was right and the order it ran in was not: the browser lane grows a whiteboard flow that clears the media gate, and the page errors it already collected get a list of their own to assert against.
A final board cannot represent work cleared between phases. Keep JPEG checkpoints for grading and vector markers for replay. That lets review survive a clear without putting images into the replay budget.
CI runs commentflow in addition to rustfmt. Keep the checkpoint queue comment in the form that the complete formatting gate requires.
Clear stored strokes only on the redo stack, so Undo stayed disabled. Preserve the cleared board as one action. Drive Clear from whether any strokes remain.
The checkpoint store matches an incoming board to its phase by label, and no test told a second phase apart from the first: cargo mutants found that turning the match around, so a new phase overwrote an earlier one and a repeated phase was stored twice, left every test passing.
36bb6ba to
f30c0bd
Compare
Summary
A whiteboard interview uses the same problem bank and six REACTO phase ids, but replaces the editor and test runner with a vector drawing surface. Nothing runs: the candidate demonstrates correctness by drawing an approach, tracing an example, and defending edge cases and complexity.
end_interview, so report generation cannot freeze an older image.Closes #65.
How an image reaches Gemini Live
The browser keeps strokes as vectors for editing and replay, then exports the full canvas as a JPEG after drawing settles or a phase completes. JPEGs use a LiveKit byte stream because they are larger than a single data packet. The agent drains that stream on its own task, keeping audio and room events responsive, and forwards the latest image to Gemini Live as realtime visual input.
flowchart LR Canvas["Canvas<br/>vector strokes"] -->|"toBlob: JPEG"| JPEG["Full-board JPEG"] JPEG -->|"LiveKit byte stream<br/>topic: board_image"| Drain["Agent drain task"] Drain --> Latest["Latest board in memory"] Latest -->|"realtimeInput.video"| Live["Gemini Live interviewer"] Tool["Gemini calls read_board"] --> Resend["Agent resends latest image"] Latest --> Resend Resend -->|"realtimeInput.video"| Liveread_boarditself returns JSON with the stroke count, board age, and timer. The image is resent separately because a tool response cannot carry the picture.Grading
Each
framework_stateupdate is reduced to the known REACTO ids, in canonical order. When a new phase appears, the browser records a replay marker and captures the canvas immediately. The agent validates the id again, assigns its own label, and retains one JPEG per phase. A changed final board is appended after the checkpoints; an identical final board is omitted.sequenceDiagram participant B as Browser participant A as Interview agent participant G as Gemini report model B->>A: JPEG + validated phase id A->>A: Keep one labeled image per REACTO phase B->>A: Final JPEG B->>A: end_interview after upload completes A->>A: Freeze transcript, evidence, and board images A->>G: generateContent(labels + inline JPEGs + prompt) Note over A,G: Retries and schema repairs resend every image G-->>A: Structured JSON report A->>A: Validate and stamp the report contract A-->>B: Candidate reportThe labels are server-owned:
Repeat,Example,Approach,Trace,Edge cases, andComplexity. A candidate-supplied string cannot become an instruction beside an image. Because earlier checkpoints remain in memory, clearing the live board does not erase evidence the report reviewer needs.Replay and report review
Replay stores no JPEGs. It stores the bounded operation journal (
stroke,undo,redo, andclear) plus an optional checkpoint id. This stays within the replay quotas and lets the review page reconstruct the exact board at every phase.flowchart LR Edit["stroke / undo / redo / clear"] --> Model["Vector board model"] Model -->|"drawing settles"| Event["Board event<br/>bounded operation batch"] Phase["REACTO phase completes"] --> Checkpoint["Board event<br/>ops + checkpoint id"] Model --> Checkpoint Event --> API["Replay event API"] Checkpoint --> API API --> Store["Ordered replay events"] Store --> Prefix["Events through selected moment"] Prefix -->|"createBoard + applyOp"| Rebuild["Reconstructed canvas"] Rebuild --> Review["Named phase checkpoint<br/>beside the report"]The clear operation remains in the journal, so replay shows the board becoming empty at the correct moment. Earlier named checkpoints still reconstruct the drawing that existed before the clear.
Clear behavior
clearreplay operation.Verification
./scripts/test.sh./scripts/browser-check.shWhiteboard