Skip to content

Add the whiteboard interview mode - #72

Draft
ColtenOuO wants to merge 7 commits into
sysprog21:mainfrom
ColtenOuO:feat/whiteboard-mode
Draft

ColtenOuO wants to merge 7 commits into
sysprog21:mainfrom
ColtenOuO:feat/whiteboard-mode

Conversation

@ColtenOuO

@ColtenOuO ColtenOuO commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A whiteboard interview uses the same problem bank and six REACTO phase ids, but replaces the editor and test runner with a vector drawing surface. Nothing runs: the candidate demonstrates correctness by drawing an approach, tracing an example, and defending edge cases and complexity.

  • Add a clearly labeled Clear board action. Clear is one undoable operation, so Undo restores the whole board.
  • Capture the board when each REACTO phase is completed. Grading receives labeled JPEG checkpoints; replay receives compact vector markers at the same moments.
  • Send the final board before end_interview, so report generation cannot freeze an older image.
  • Keep the existing startup fix that lets whiteboard interviews pass the media preflight.

Closes #65.

How an image reaches Gemini Live

The browser keeps strokes as vectors for editing and replay, then exports the full canvas as a JPEG after drawing settles or a phase completes. JPEGs use a LiveKit byte stream because they are larger than a single data packet. The agent drains that stream on its own task, keeping audio and room events responsive, and forwards the latest image to Gemini Live as realtime visual input.

flowchart LR
    Canvas["Canvas<br/>vector strokes"] -->|"toBlob: JPEG"| JPEG["Full-board JPEG"]
    JPEG -->|"LiveKit byte stream<br/>topic: board_image"| Drain["Agent drain task"]
    Drain --> Latest["Latest board in memory"]
    Latest -->|"realtimeInput.video"| Live["Gemini Live interviewer"]
    Tool["Gemini calls read_board"] --> Resend["Agent resends latest image"]
    Latest --> Resend
    Resend -->|"realtimeInput.video"| Live
Loading

read_board itself returns JSON with the stroke count, board age, and timer. The image is resent separately because a tool response cannot carry the picture.

Grading

Each framework_state update is reduced to the known REACTO ids, in canonical order. When a new phase appears, the browser records a replay marker and captures the canvas immediately. The agent validates the id again, assigns its own label, and retains one JPEG per phase. A changed final board is appended after the checkpoints; an identical final board is omitted.

sequenceDiagram
    participant B as Browser
    participant A as Interview agent
    participant G as Gemini report model

    B->>A: JPEG + validated phase id
    A->>A: Keep one labeled image per REACTO phase
    B->>A: Final JPEG
    B->>A: end_interview after upload completes
    A->>A: Freeze transcript, evidence, and board images
    A->>G: generateContent(labels + inline JPEGs + prompt)
    Note over A,G: Retries and schema repairs resend every image
    G-->>A: Structured JSON report
    A->>A: Validate and stamp the report contract
    A-->>B: Candidate report
Loading

The labels are server-owned: Repeat, Example, Approach, Trace, Edge cases, and Complexity. A candidate-supplied string cannot become an instruction beside an image. Because earlier checkpoints remain in memory, clearing the live board does not erase evidence the report reviewer needs.

Replay and report review

Replay stores no JPEGs. It stores the bounded operation journal (stroke, undo, redo, and clear) plus an optional checkpoint id. This stays within the replay quotas and lets the review page reconstruct the exact board at every phase.

flowchart LR
    Edit["stroke / undo / redo / clear"] --> Model["Vector board model"]
    Model -->|"drawing settles"| Event["Board event<br/>bounded operation batch"]
    Phase["REACTO phase completes"] --> Checkpoint["Board event<br/>ops + checkpoint id"]
    Model --> Checkpoint
    Event --> API["Replay event API"]
    Checkpoint --> API
    API --> Store["Ordered replay events"]
    Store --> Prefix["Events through selected moment"]
    Prefix -->|"createBoard + applyOp"| Rebuild["Reconstructed canvas"]
    Rebuild --> Review["Named phase checkpoint<br/>beside the report"]
Loading

The clear operation remains in the journal, so replay shows the board becoming empty at the correct moment. Earlier named checkpoints still reconstruct the drawing that existed before the clear.

Clear behavior

  • Clear board is disabled while the board is empty.
  • Clearing synchronizes the empty state to Jim and writes a clear replay operation.
  • Undo restores the cleared board as one action.
  • Phase checkpoints already captured for grading and replay remain available.

Verification

  • ./scripts/test.sh
  • ./scripts/browser-check.sh

Whiteboard

Whiteboard interview

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Need some time for testing and self-review; will mark this PR as ready for review once it's set.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Perhaps we will need to record the state of the whiteboard at each stage so that we can provide more information to the user at the end of the interview.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Additionally, pseudocode should be graded separately rather than being combined with optimization like it is currently.

@alanhc

alanhc commented Sep 24, 2026

Copy link
Copy Markdown
Collaborator

On the concern from #65 about whether Gemini can read a drawn board reliably: one option worth considering is Excalidraw as the board layer.

The drawing tools are not the point; the data model is. An Excalidraw scene is a list of typed elements, and boxes, arrows (bound to the boxes they connect) and text keep their actual strings. So next to the JPEG, the agent can get a plain-text summary of the board, e.g. box "left" -> box "mid", and typed pseudo-code arrives as exact text rather than something the model has to read off pixels. The report prompt can quote it too.

I checked whether it fits the vendoring rules, following the three-vrm.js precedent (one esbuild bundle, committed, with a reproduce recipe):

  • @excalidraw/excalidraw@0.18.1 + React 19 bundle into a single ES module: 4.9 MB (1.6 MB gzip), plus 145 KB of CSS. That is with @excalidraw/mermaid-to-excalidraw aliased to a stub; without it, 8.5 MB.
  • Fonts load from window.EXCALIDRAW_ASSET_PATH. The Latin ones are ~0.5 MB. The CJK font (Xiaolai) is 13 MB; I left it out, and Excalidraw then falls back to esm.sh for it. The current CSP refuses those loads (230 console violations in my run, drawing and export unaffected), but it needs either vendoring or silencing.
  • Served locally with every non-local request aborted, under the policy from src/web/policy.rs (script-src 'self', style-src 'self', nothing inline): it mounts in ~200 ms, freehand, shapes and text all work, and exportToBlob gives a ~20 KB JPEG. The only inline bit was setting EXCALIDRAW_ASSET_PATH, which moves to a same-origin script.
  • A 41-point freehand stroke is ~1.5 KB of JSON, so replay would still journal per-element changes from onChange rather than whole scenes.

The costs:

  • Far bigger than the 280-line whiteboard.js, and it brings React into the page.
  • The structured benefit exists only when the candidate uses shapes and the text tool. Freehand pseudo-code is still just points. Whether typed text belongs in a whiteboard interview is a product call: less realistic, much easier to grade.
  • Text input opens the integrity side: paste and library/file import would have to be disabled.

Only the board layer changes. The byte stream, read_board, the report attachment and the replay plumbing stay as they are. Happy to put together the vendored bundle and the scene-to-text summary, either on top of this branch or as a follow-up once it lands, whichever you prefer.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Hi, @alanhc

Thanks for the discussion!

Excalidraw was also brought up in the discussions under the Facebook post back then. After looking into it, I think it seems like a solid option to consider, but I also have a few thoughts.

If our goal is to simulate an actual interview environment as closely as possible, would Excalidraw make things too convenient for the user? After all, in a real interview, you only have a marker and a physical whiteboard. That raw experience is precisely what we want to deliver in this mode—allowing users to practice explaining their problem-solving approach and thought process purely through drawing on a blank board.

Currently, a tester, @Eason0729 , has tested this feature (the non-Excalidraw version) and provided a lot of feedback. Here is a brief summary of the key concerns raised from the testing:

  1. Should we support tablet touch input? Drawing with a mouse can significantly degrade the user's drawing experience and heavily impact their performance. Therefore, I would strongly prefer having this supported.

  2. Unreliable AI recognition will severely impact the quality of the questions asked.

For now, I'll focus on addressing the points raised in the feedback first. As for Excalidraw, we can discuss it further as a potential future improvement.

Overall, I actually think this is quite a promising direction. However, my primary concern is that users should have an experience that feels as close to reality as possible; we probably shouldn't compromise on that just to make recognition easier for the model.

WDYH

A whiteboard interview takes the same problem bank and the same six
steps and swaps the editor and the test runner for a board. Nothing
runs, so correctness is what the candidate can defend by tracing an
example across their own drawing.

The board travels as a LiveKit byte stream rather than on a data topic,
because one board is tens of kilobytes and a data packet carries
fifteen, and it reaches Gemini the way a camera frame does, as a
realtime image. A tool response is JSON and cannot carry a picture, so
read_board asks for the board to be sent again and answers with what an
image cannot say: how much is on it and how long ago it was drawn.

Bundle 21 is the live prompt following the surface. A whiteboard
session is told it has no editor and no test runner, is offered
read_board in place of read_editor, and records board_snapshot where
the other records an editor snapshot or a test event; neither may
record the other's source. The phases about written work are gated on
strokes instead of on characters, and Test on the cases named against
the drawing instead of on a run, or an interview with no editor would
be refused the second half of its own flow. The report still reads the
transcript alone and the replay still carries no board.
The whiteboard was live but nowhere else: the report was written from a
transcript and an editor nobody opened, the recording kept no drawing,
and the checklist beside the timer called step four Coding while the
interviewer was asking the candidate to trace.

The board now rides the report request as an inline image, ahead of the
brief and on every repair, and the brief says what it is looking at:
nothing ran, so correctness is the trace the candidate walked, and the
phases keep their names with Coding meaning that trace, Test the cases
named against the drawing, and Optimizations the complexity they
confirmed. The system instruction is left as it is, so every report
call still shares its prefix, and the brief tells a whiteboard reviewer
how to read the rules that speak of code. That is report prompt 14,
inside the same bundle 21, and the reviewer is told when no board
arrived rather than sent looking for an attachment that is not there.

The recording keeps the drawing as the operations that made it. One
board as a JPEG is past the per-event ceiling on its own and would
spend the whole per-interview budget on a handful of frames; the same
board as strokes is a few kilobytes, so the replay page and the
recording template redraw any moment of the interview with the module
the candidate drew on, rather than the few moments a photograph could
afford.
`init()` runs at the top of interview.js and calls `initWhiteboard`,
which reached for a `const` declared eight hundred lines further down:
still inside its temporal dead zone, so it threw. The editor panel was
already gone and the board already shown, which is why the page looked
right, but `init()` stopped there and never reached the media preflight
it runs next. The gate sat on "Starting camera and microphone..." for
ever, with the browser never having asked for either device, and no
whiteboard interview could be started at all.

The state moves up with the rest of the module's, where web/app.js keeps
its own for this exact reason. Nothing reading source text could have
caught it, because the source was right and the order it ran in was not:
the browser lane grows a whiteboard flow that clears the media gate, and
the page errors it already collected get a list of their own to assert
against.
A final board cannot represent work cleared between phases. Keep JPEG
checkpoints for grading and vector markers for replay. That lets review
survive a clear without putting images into the replay budget.
CI runs commentflow in addition to rustfmt. Keep the checkpoint queue
comment in the form that the complete formatting gate requires.
Clear stored strokes only on the redo stack, so Undo stayed disabled.
Preserve the cleared board as one action. Drive Clear from whether any
strokes remain.
The checkpoint store matches an incoming board to its phase by label,
and no test told a second phase apart from the first: cargo mutants
found that turning the match around, so a new phase overwrote an
earlier one and a repeated phase was stored twice, left every test
passing.
@ColtenOuO
ColtenOuO force-pushed the feat/whiteboard-mode branch from 36bb6ba to f30c0bd Compare October 1, 2026 12:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a whiteboard interview mode

2 participants