Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 15 additions & 1 deletion .cargo/mutants.toml
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# Functions the mutation gate cannot judge, because `cargo test` cannot reach
# them. Matched against the mutant names that `cargo mutants --list` prints.
#
# EXCLUSIONS: 54
# EXCLUSIONS: 57
#
# That number is checked by `scripts/test.sh`, so adding an entry means editing
# this line too. The point is not the count, it is that the list only ever grows
Expand Down Expand Up @@ -247,6 +247,17 @@
# what may trigger it on a tick is `RuntimeActivity::settle_stalls`; both are
# tested.
#
# `maybe_refresh_context`, `record_live_usage` and `drain_live_usage` are the
# room's half of the Live usage accounting and the compression checkpoint. The
# first sends the checkpoint on the Gemini socket and logs it for the room;
# the other two read the socket's usage ledger and log each observation for
# the room. All three take the live `GeminiEventContext`, so `()` looks the
# same from outside as the work, for the reason `send_model_text` does. What
# they decide is tested where it is decided: `RuntimeActivity::checkpoint_due`
# and `can_refresh_context` for when a checkpoint may go out,
# `account_live_usage` for what an observation adds and how its line reads,
# and the ledger's ordering against a local socket in tests/unit/gemini.rs.
#
# Keep this list short and each entry justified. An entry that is really "we
# never got around to testing this" belongs in a test, not here.
exclude_re = [
Expand Down Expand Up @@ -297,6 +308,9 @@ exclude_re = [
"generate_interim_review",
"end_through_control",
"spend_deferred_restart",
"maybe_refresh_context",
"record_live_usage",
"drain_live_usage",
"handle_media_event",
"attach_audio",
"next_audio_frame",
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -211,7 +211,7 @@ The common ones:
| `GEMINI_LIVE_MODEL` | `gemini-3.1-flash-live-preview` | Realtime interviewer model |
| `GEMINI_REPORT_MODEL` | `gemini-3.1-flash-lite` | Report model |
| `CODETRIAL_MAX_INTERIM_REVIEWS` | `6` | Quiet-pause report-model reviews per interview; `0` disables them and `72` is the maximum |
| `CODETRIAL_GEMINI_CANDIDATE_VIDEO_ENABLED` | `false` | Forward candidate video to Gemini |
| `CODETRIAL_GEMINI_CANDIDATE_VIDEO_ENABLED` | `false` | Forward candidate video to Gemini, one low-resolution frame in five seconds |
| `CODETRIAL_COMPILER_EXPLORER_ENABLED` | `true` | Enable remote C, C++, and Java runs |
| `CODETRIAL_MAX_CONCURRENT_INTERVIEWS` | `16` | Interviews one `web` process hosts agents for |

Expand Down
6 changes: 6 additions & 0 deletions config/codetrial.env.example
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,12 @@ CODETRIAL_WEB_DIR=web
CODETRIAL_WEB_ADDR=127.0.0.1:3000
CODETRIAL_COMPILER_EXPLORER_ENABLED=true
CODETRIAL_GEMINI_CANDIDATE_VIDEO_ENABLED=false
# Optional pair for Live context compression experiments. Unset keeps the
# provider defaults; smaller windows discard older dialogue. See
# docs/provider-cost-and-degradation.md before choosing production values, and
# run `codetrial check-gemini` to confirm the model accepts them.
# GEMINI_CONTEXT_TRIGGER_TOKENS=20000
# GEMINI_CONTEXT_TARGET_TOKENS=8000
# How long the candidate has to stay quiet before Gemini takes a turn, and how
# eagerly it starts one. Capped at 30000; START_SENSITIVITY_HIGH interrupts more.
# GEMINI_SILENCE_MS=1000
Expand Down
3 changes: 2 additions & 1 deletion docs/interview-contract-versions.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,10 +8,11 @@ can select it.

## The active bundle

Bundle 23: live prompt 15, report prompt 15, rubric 1, report schema 2.
Bundle 24: live prompt 16, report prompt 15, rubric 1, report schema 2.

| Bundle | Introduced |
|---|---|
| 24 | The Live main instructions drop repeated explanations and illustrative examples and keep every timer, round, evidence-source and hint restriction. The greeting answers only the platform's startup request, and missing history, a compression or a tool result is not a new interview. `end_interview` is called silently, before any acknowledgment or goodbye, and the platform supplies the closing. A cut `read_editor` page or a checkpoint excerpt does not show the whole buffer, so an implementation or technique is not called absent before the named lines are read. The `read_editor` description asks for only the code the current question needs that nothing has shown, from a known relevant line rather than a refill of the whole editor. The greeting no longer repeats the exercise's title and brief, which THE EXERCISE already carries and the greeting now points at; the framework headers drop a scoring premise the disclosure rule already covers; test-run reactions and the earlier-steps reminder state their rule once, more briefly; and the `end_interview` description no longer restates the instruction it sits beside. With a configured compression window, a silent checkpoint rebuilt from local state follows a detected cut: the chosen language, the current round, the evidence, a bounded transcript that keeps a long behavioral round's opening, a bounded test report and, in the coding round, the editor's opening and ending. Its next step applies to the next candidate input, not to the checkpoint itself. Omission alone does not close a behavioral round, repeat its question or establish that its follow-up is unused, and a refusal or request to finish supplies no STAR evidence. Under the same window, editor, hint and evidence tool answers carry the latest unanswered candidate utterance as quoted historical data, never as a new turn. |
| 23 | A candidate who hides the worked examples in the preflight sends `hideExamples` with the token request, and the live prompt then says no examples are on their screen: the interviewer never points them at one, says a clarification or hint clue that mentions an example with a case they proposed or one of its own, and in the Example step asks for their ordinary and boundary cases before offering a small example once they have tried or are stuck. A session that does not hide them gets the live prompt unchanged. |
| 22 | Browser-reported failed judge cases include their bounded input in the live reaction, `read_editor`, and final report test summary, so the interviewer can connect an expected result or exception to the case that produced it. The live reaction still lists one failure, while `read_editor` and the report retain their existing fuller failure account. |
| 21 | English interview instructions treat unclear, unexpectedly non-English or unrelated speech as possible recognition failure and ask one neutral clarification without supplying an answer or recording evidence from the uncertain turn; typed code comments can clarify speech. Interim and final assessment share an evidence-reliability policy excluding uncertain speech and unsupported rolling observations. The final report additionally excludes them from credit, deductions and verdict reasoning, leaving unsupported phase scores null; interviewer agreement cannot prove an answer was correct, and clear technical mistakes remain assessable. When recognition leaves little reliable communication evidence, the report judges communication and the decision rule from what remains, says so in the summary, and never makes the gap an improvement. The server scan refuses a report that names the language a transcript came out in, as "in Japanese" or "a Japanese response", or judges English proficiency, and an improvement that asks the candidate to speak English, audibly or more clearly. The live instructions, which every reconnect sends again, say that a possibly misrecognized recovered line, or agreement with one, supports no missing evidence. Live setup sends the documented `inputAudioTranscription` hints: `languageCodes` for `en-US`, and `customVocabulary` with the scenario title, the names in the starter and a fixed list of terms every interview uses, such as "time complexity". Both bias only the transcript that notes, the report and recovery read, not what the interviewer hears, and neither locks recognition. Interim notes and the report read a fixed marker, which their prompts name, in place of any candidate turn written mostly in a non-Latin script, a lone symbol or two excepted; the live interviewer, replay and stored transcript keep the recognizer's text, and the server refuses speech evidence while the candidate's latest turn is one the marker hides. Recognition errors in Latin letters are not marked, and apart from the scan and the marker these safeguards are prompt instructions. None of it guarantees transcription accuracy. |
Expand Down
134 changes: 131 additions & 3 deletions docs/provider-cost-and-degradation.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,17 @@ editor, round and evidence state instead.

`GEMINI_RESTART_LIMIT` bounds a failing endpoint rather than a long interview.
It allows 8 opens in a row, and any socket that lived past a minute clears the
run.
run. A project that cannot pay is not an endpoint that may recover: a 402, or a
close reason saying the prepaid credit is depleted or billing is not enabled,
takes that key off both the Live and the report surface and is retried only on
another configured key. With none left the interview ends at once instead of
spending the remaining opens, and its summary line says `outcome=billing`.

Every Live turn is billed on the whole context it runs in, retained audio and
images included, so what stays in the context costs again on every later turn.
Candidate video, off by default, therefore sends one frame in five seconds and
asks for the low media resolution; the camera is there for presence, and the
code reaches the model as text.

Final reporting has its own hard budget of six Gemini HTTP calls: initial
generation plus one semantic repair, each generation allowing its first call and
Expand Down Expand Up @@ -81,10 +91,128 @@ presents canned feedback as an agent evaluation.

Watch `codetrial dispatch_refused ... reason=at_capacity`, `livekit quota:`
transitions, token HTTP 429 with `Retry-After`, `gemini report
transport_failed call=... retry=...`, and the bounded incomplete-report
categories.
transport_failed call=... retry=...`, `codetrial live_usage ... outcome=billing`,
and the bounded incomplete-report categories.

Raise concurrency only after checking provider minutes, Gemini limits, CPU and
audio capacity, and the token burst policy. The deterministic dispatcher,
rate-isolation, resume-limit, report-budget, request-isolation, and browser-state
tests all run under `./scripts/test.sh`.

## Measuring token consumption

Google's [Live API billing guidance](https://ai.google.dev/gemini-api/docs/live-api/best-practices#pricing-and-billing)
bills each turn on the entire active context, retained raw audio included, and
adds a text-output charge for enabled audio transcription. The logs below are
what the provider reported to this server, not an invoice. Reconcile them with
an isolated project's billing before quoting a monetary figure, and do not
infer an account balance from them.

- `codetrial live_turn_usage` is one usage observation: room, `session` (the
epoch millisecond the room loop started, so two runs sharing a room stay
apart),
interview clock, socket, event index, `cause`, and the provider's counters.
- `cause` names the last platform input that asked for a reply before the
observation: `watch`, `turn` (greeting, test reaction, wrap-up and similar
stage directions), `tool` or `recovery`, or `candidate` when the platform
asked for nothing since the previous completed turn. Silent context, such as
a compression checkpoint, asks for nothing and is billed inside whichever
turn follows; `codetrial context_refresh` lines count it. The label says where
turns come from, not a causal record: speech overlapping a platform input is
credited to the input.
- `codetrial live_usage` sums a session's observations across its sockets, with
the model, elapsed seconds, socket count and an `outcome` of `ok`, `error`,
`gemini_unreachable` or `billing`. It is written on every exit of the room
loop, and once with `phase=startup` and no socket count when the interview
failed before its first turn, a first open that was refused included.
- `gemini report` and `gemini interim` lines carry the room, call number and
retry count of each HTTP call. A failed call has no usage line; every one is
counted from `gemini report transport_failed` (with `final=true` when it was
not retried) and `interim review skipped`, and none is assumed free.

The counters keep prompt, response, cached, thought, tool-use prompt and total
counts, plus modality details for prompt, response, tool-use and cached tokens.
Response aliases are alternatives, not additive counters. A scalar the provider
omits is logged as zero. Each `*_detail_samples` counts valid detail arrays, so
zero samples means the breakdown is unknown and fewer than `usage_samples`
means it is partial; details may also sum to less than their scalar. The
analyzer reports a direction with a zero total and no breakdown, such as tool
use in a session without tools, as `none_reported`, never as complete: an
unused direction and one the provider left out look the same. Never add
cached tokens to the prompt total.

`turn_complete_samples` counts observations that shared a frame with
`turnComplete`. Usage is recorded when its frame is decoded, before that frame's
content is queued, so a room that stops reading does not lose what it had
received. If a session's `usage_samples` exceeds its `turn_complete_samples`,
some observations arrived on frames of their own, and a sum may include
periodic snapshots of the same turn; treat it as an upper bound until the
provider's event semantics are confirmed. A process kill or task cancellation
can still leave a session without its summary.

`codetrial model_input_bytes` measures what this server constructed, by kind.
It is not billed tokens: it cannot see the retained context a turn is billed on.

## Comparing captured logs

`python3 scripts/analyze-gemini-usage.py interview.log` prints the counters above
as JSON, grouped by room and by Live session, and ignores everything else in
the log, including prefixes a log collector adds. `--room ROOM` selects one
room; stdin is read when no file is given, and the exit status is 2 when no
usage record matches.

A summarized session is reported from its summary; a session with events but
no summary is reported from them and marked incomplete. For each session with
events it also reports the prompt count of each completed turn in order, its
growth per turn, and counts by `cause`. The output makes no price or quality
claim.

Compare one change at a time, on the same synthetic interview, model, speech,
editor events and planned turns, with repeated runs. Record prompt and modality
counts, usage coverage, completed turns, reconnects, first-audio latency,
completion, and whether the interview still recalled the evidence it needed. A
projection is not an observed saving, and a smaller context must preserve the
interview's evidence before it becomes a default.

## Configuring a compression experiment

`GEMINI_CONTEXT_TRIGGER_TOKENS` and `GEMINI_CONTEXT_TARGET_TOKENS` are an
optional pair. Both must be positive integers, target strictly below trigger,
or config loading fails. Unset, setup still sends `slidingWindow: {}` and
leaves the thresholds to the provider. The pair is sent on every socket,
resumed ones included. Only the provider knows the model's context limit, so
run `codetrial check-gemini`, which opens a session with the same setup, before
an interview does.

With a pair configured, the room watches for a cut context and then sends a
silent checkpoint rebuilt from local state. A turn is taken to have run on a cut
context when its prompt count, judged at completion against the largest count
since the previous completed turn, fell by half the trigger-target gap (limited
to 1 to 2048 tokens), or fell at all from a context that had reached the
trigger. Both are heuristics, not an API compression event.

The checkpoint waits until the candidate is not mid-turn by transcript or by
microphone level, no reply, tool continuation or queued audio is outstanding,
and the interview is not paused; the room retries after Gemini events, at the
playout boundary and on the watch tick. It carries the platform timer, the
chosen language, the current round, the evidenced phases, a 2,500-byte
transcript budget (a long behavioral round keeps up to 750 bytes of its opening
beside the recent dialogue), the test report up to 1,000 bytes, and in the
coding round up to 1,800 bytes of the editor's opening and ending. Omitted
lines are named as omitted and remain available through `read_editor`; omission
is not evidence that a follow-up was unused or that no refusal occurred. A
replacement socket drops a pending checkpoint, because its recovery already
carries local state, and socket recovery keeps its own larger budgets.

Only while a pair is configured do `read_editor`, `log_hint` and
`record_framework_evidence` answers also carry the latest unanswered candidate
utterance, as quoted data, so a reply owed across a cut is not lost. Without a
pair nothing leaves the context during a tool call, and the text would only be
billed again on every later turn.

A lower threshold reduces the history retained for later turns, and may remove
information a follow-up needs or add compression latency. Final reporting still
uses the full local transcript and editor. The credentialed probes that compare
arms and check recall are the ignored tests in `tests/unit/livekit/cost.rs`,
outside the credential-free gate; their file header lists the environment they
need.
Loading