Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 14 additions & 1 deletion .cargo/mutants.toml
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# Functions the mutation gate cannot judge, because `cargo test` cannot reach
# them. Matched against the mutant names that `cargo mutants --list` prints.
#
# EXCLUSIONS: 57
# EXCLUSIONS: 60
#
# That number is checked by `scripts/test.sh`, so adding an entry means editing
# this line too. The point is not the count, it is that the list only ever grows
Expand Down Expand Up @@ -258,6 +258,16 @@
# `account_live_usage` for what an observation adds and how its line reads,
# and the ledger's ordering against a local socket in tests/unit/gemini.rs.
#
# `publish_thinking_state`, `flush_thinking_notice` and
# `publish_opening_attributes` are writes to the room and nothing else, the
# same class as `set_agent_state` and `publish_interviewer_state`: each takes a
# `&Room`, and replacing its body with `Ok(())` or `()` sends nothing a test
# with no room can miss. What decides them is tested without one: the notice
# the state methods leave (`take_thinking_notice`), the message's shape
# (`thinking_state_message`), and the attributes (`turn_window_attributes`,
# `agent_state_attributes`). The live turn-taking run through
# scripts/browser-check.sh is what saw them arrive.
#
# Keep this list short and each entry justified. An entry that is really "we
# never got around to testing this" belongs in a test, not here.
exclude_re = [
Expand Down Expand Up @@ -318,4 +328,7 @@ exclude_re = [
"pump_video",
"encode_video_frame_jpeg_off_thread",
"replace percent_encode_component -> String with String::new\\(\\)",
"publish_thinking_state",
"flush_thinking_notice",
"publish_opening_attributes",
]
2 changes: 1 addition & 1 deletion config/codetrial.env.example
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ CODETRIAL_GEMINI_CANDIDATE_VIDEO_ENABLED=false
# GEMINI_CONTEXT_TARGET_TOKENS=8000
# How long the candidate has to stay quiet before Gemini takes a turn, and how
# eagerly it starts one. Capped at 30000; START_SENSITIVITY_HIGH interrupts more.
# GEMINI_SILENCE_MS=1000
# GEMINI_SILENCE_MS=3000
# GEMINI_START_SENSITIVITY=START_SENSITIVITY_LOW
# How readily it decides the candidate has finished. Unset keeps the API's own;
# HIGH ends turns sooner and cuts slow speakers off more.
Expand Down
11 changes: 9 additions & 2 deletions docs/integrity-response-window-check.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,8 +86,8 @@ All six must hold. Any failure is a fail row, and the row names which one broke.
- An ordinary answered question does not say it holds no transcript. This is the
delivery race `DONE.md` records a fix for, and a real interview is the only
thing that has ever exercised it.
- The window covering the pause from step 4 says "interview paused during this
window", and no other window does.
- The window covering the pause from step 4 says "interview paused or thinking
time requested during this window", and no other window does.

## The judgement

Expand Down Expand Up @@ -153,3 +153,10 @@ answer.
| 2. Closing turn rendered or dropped? | | |
| 3. Consecutive empty windows worth marking (a number) | | |
| 4. Pause mark enough, or its own entry? | | |

Explicit thinking time is a declared conversation hold. Its `thinking_started`
and `thinking_ended` lifecycle rows mark an overlapping response window the
same way a declared pause does; the panel names both possibilities. Neither
requests a candidate judgment or changes the deadline. A subsequent interviewer
speaking row bounds a missing end row, since the runtime suppresses speech
while the hold is active.
3 changes: 2 additions & 1 deletion docs/interview-contract-versions.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,10 +8,11 @@ can select it.

## The active bundle

Bundle 24: live prompt 16, report prompt 15, rubric 1, report schema 2.
Bundle 25: live prompt 17, report prompt 15, rubric 1, report schema 2.

| Bundle | Introduced |
|---|---|
| 25 | Candidates can keep the floor while thinking, reclaim it during a reply, and yield it early. Explicit spoken requests for thinking time in English, including one that follows an answer in the same sentence or is asked as a question, suppress generated replies and automatic nudges until the candidate speaks again or chooses to continue. A hold ends on its own at the five-minute warning, at the round transition, and after two silent minutes with one brief check-in; the interviewer is told that anything it said during the hold was not heard. A Continue within ten seconds of the last one releases the hold without a reply of its own. Thinking keeps editor, microphone and test evidence live, gives the interviewer test runs and edits as context it does not answer, and never extends the deadline. The default endpointing window is three seconds, and the page shows it filling while the candidate is silent; yielding ends the audio stream so the interviewer replies without waiting it out. |
| 24 | The Live main instructions drop repeated explanations and illustrative examples and keep every timer, round, evidence-source and hint restriction. The greeting answers only the platform's startup request, and missing history, a compression or a tool result is not a new interview. `end_interview` is called silently, before any acknowledgment or goodbye, and the platform supplies the closing. A cut `read_editor` page or a checkpoint excerpt does not show the whole buffer, so an implementation or technique is not called absent before the named lines are read. The `read_editor` description asks for only the code the current question needs that nothing has shown, from a known relevant line rather than a refill of the whole editor. The greeting no longer repeats the exercise's title and brief, which THE EXERCISE already carries and the greeting now points at; the framework headers drop a scoring premise the disclosure rule already covers; test-run reactions and the earlier-steps reminder state their rule once, more briefly; and the `end_interview` description no longer restates the instruction it sits beside. With a configured compression window, a silent checkpoint rebuilt from local state follows a detected cut: the chosen language, the current round, the evidence, a bounded transcript that keeps a long behavioral round's opening, a bounded test report and, in the coding round, the editor's opening and ending. Its next step applies to the next candidate input, not to the checkpoint itself. Omission alone does not close a behavioral round, repeat its question or establish that its follow-up is unused, and a refusal or request to finish supplies no STAR evidence. Under the same window, editor, hint and evidence tool answers carry the latest unanswered candidate utterance as quoted historical data, never as a new turn. |
| 23 | A candidate who hides the worked examples in the preflight sends `hideExamples` with the token request, and the live prompt then says no examples are on their screen: the interviewer never points them at one, says a clarification or hint clue that mentions an example with a case they proposed or one of its own, and in the Example step asks for their ordinary and boundary cases before offering a small example once they have tried or are stuck. A session that does not hide them gets the live prompt unchanged. |
| 22 | Browser-reported failed judge cases include their bounded input in the live reaction, `read_editor`, and final report test summary, so the interviewer can connect an expected result or exception to the case that produced it. The live reaction still lists one failure, while `read_editor` and the report retain their existing fuller failure account. |
Expand Down
3 changes: 3 additions & 0 deletions scripts/gen-wire-fixtures.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,9 @@ function codeUpdateCases(languages) {

function controlCases() {
return [
{ name: "thinking start", payload: lib.thinkingPayload(true) },
Comment thread
jserv marked this conversation as resolved.
{ name: "thinking end", payload: lib.thinkingPayload(false) },
{ name: "yield turn", payload: lib.yieldTurnPayload() },
{ name: "time warning five minutes", payload: lib.timeWarningPayload(300) },
{ name: "time warning one minute", payload: lib.timeWarningPayload(60) },
{
Expand Down
53 changes: 51 additions & 2 deletions src/agent.rs
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,15 @@
//! and `prompts` are the data this file used to hold.

mod events;

mod turn_taking;
#[cfg(test)]
pub(crate) use turn_taking::THINKING_REQUEST_SETTLE;
pub use turn_taking::ThinkingHold;
pub(crate) use turn_taking::{
owed_context, thinking_change, thinking_check_in, thinking_owes_context, thinking_resume,
with_thinking_debt,
};
mod evidence;
mod integrity;
mod problem_guides;
Expand Down Expand Up @@ -145,8 +154,19 @@ const ROUND_TRANSITION_SKEW: std::time::Duration = std::time::Duration::from_sec
/// `the_time_warning_threshold_is_the_same_number_on_both_sides`.
pub const TIME_WARNING_S: u64 = 300;

pub const INTERVIEW_CONTRACT_BUNDLE_VERSION: u32 = 24;
pub const LIVE_PROMPT_VERSION: u32 = 16;
/// How long a declared thinking hold runs in silence before the interviewer
/// checks in once. A hold with no end kept every nudge quiet for as long as the
/// candidate stayed silent, which could be the rest of the interview.
pub const THINKING_CHECK_IN_S: u64 = 120;

/// A Continue this soon after the last one releases the hold without asking
/// the interviewer to say anything. Each release is otherwise a generated
/// turn, and a candidate clicking Thinking on and off would buy one per click.
pub(crate) const THINKING_RELEASE_COOLDOWN: std::time::Duration =
std::time::Duration::from_secs(10);

pub const INTERVIEW_CONTRACT_BUNDLE_VERSION: u32 = 25;
pub const LIVE_PROMPT_VERSION: u32 = 17;
pub const REPORT_PROMPT_VERSION: u32 = 15;
pub const RUBRIC_VERSION: u32 = 1;
pub const REPORT_SCHEMA_VERSION: u32 = 2;
Expand Down Expand Up @@ -721,6 +741,18 @@ pub struct RuntimeState {
/// transcript line. Recovery must include that turn if its text changed.
pub behavioral_round_prior_turn: Option<(usize, String)>,
pub paused: bool,
/// Thinking keeps media and editor evidence live, but suppresses replies.
/// Changed only through the methods in `turn_taking`.
pub thinking_hold: ThinkingHold,
/// When the Thinking button last ended a hold, so toggling it cannot make
/// the interviewer answer every click; see `THINKING_RELEASE_COOLDOWN`.
pub thinking_released_at: Option<std::time::Instant>,
/// The declared hold state the page has not been told yet; see
/// `take_thinking_notice`.
pub thinking_notice: Option<bool>,
/// A reply the hold dropped is still in the model's history, and the
/// release has to say so; see `thinking_resume`.
pub thinking_unheard_reply: bool,
pub framework_evidence: Vec<FrameworkEvidence>,
/// Deterministic, bounded facts derived from the live session. The ledger
/// deliberately holds no editor text or runner diagnostics: those remain
Expand Down Expand Up @@ -898,6 +930,10 @@ impl Default for RuntimeState {
behavioral_round_transcript_start: 0,
behavioral_round_prior_turn: None,
paused: false,
thinking_hold: ThinkingHold::Off,
thinking_released_at: None,
thinking_notice: None,
thinking_unheard_reply: false,
framework_evidence: Vec::new(),
evidence_ledger: EvidenceLedger::default(),
code: String::new(),
Expand Down Expand Up @@ -1914,6 +1950,19 @@ pub struct DataEventResult {
/// Some only when the pause state genuinely changed, so a browser asking
/// twice for what it already has publishes nothing.
pub pause_changed: Option<bool>,
pub thinking_changed: Option<bool>,
pub yield_turn: bool,
/// `generate_reply` carries what a hold left owed (see `thinking_resume`),
/// so delivering it pays that debt.
pub carries_thinking_debt: bool,
/// The clock outranks a thinking hold: the reply ends it rather than being
/// suppressed by it. Held until the candidate spoke again, the five-minute
/// warning reached a candidate thinking in silence only at the deadline,
/// and a new round cannot wait on thinking about the last one.
pub preempts_hold: bool,
/// A reply the thinking hold suppressed, to be given to the model as
/// context that asks for no answer.
pub held_context: Option<String>,
/// Agent-owned round transition result: `started` or `skipped`.
pub round_changed: Option<&'static str>,
/// How a test result was judged, for the room log; see `TestRunNote`.
Expand Down
Loading