Skip to content

Measure and trim what a Live interview spends - #203

Open
jserv wants to merge 2 commits into
mainfrom
token-reduction
Open

jserv wants to merge 2 commits into
mainfrom
token-reduction

Conversation

@jserv

@jserv jserv commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

An interview spends Gemini credit faster than users expect, and the server could not say where: it logged one usage sum per interview, could lose the last turn's usage when a socket was replaced or the room shut down, and treated a project out of prepaid credit like any other outage. This records usage as each frame is decoded and logs it per turn with the session and the input that asked for the generation, logs a per-session summary with how it ended (billing included), and adds scripts/analyze-gemini-usage.py to summarize those lines offline. A 402 or a depleted-credit close is now retried only on another key. The Live instructions drop repeated explanation (live prompt 16, bundle 24), candidate video sends one low-resolution frame in five seconds, and an optional GEMINI_CONTEXT_TRIGGER_TOKENS/GEMINI_CONTEXT_TARGET_TOKENS pair enables silent local-state checkpoints for compression experiments. Defaults stay with the provider until a measured comparison justifies a change; docs/provider-cost-and-degradation.md describes the logs and the experiment.

Verification:

  • ./scripts/test.sh passes; the lanes skipped locally are Playwright Chromium, cargo-audit and shellcheck.
  • codetrial check-gemini with gemini-3.1-flash-live-preview: setup accepted with the new setup message.
  • The credentialed probes in tests/unit/livekit/cost.rs against a real key: usage arrives once per turn, on the turnComplete frame (usage_samples=1, turn_complete_samples=1), so the per-turn sums do not double count; the language, declined-behavioral and positive-behavioral checkpoint probes pass; the editor probe passes about one run in three, as it did before this change.
  • cargo mutants --in-diff: the survivors of a full local run are covered by new tests, and a loop that hung under mutation was replaced with a bounded one.

Not included: a pre-session readiness check and browser messaging for quota and billing failures, a pre-session cost warning, a session length cap, a text-only mode, and any change to the default compression window.


Summary by cubic

Measures and trims what a Live interview spends on Gemini, so the server can say where credit goes and stops spending restart budget on accounts that cannot pay.

  • Records usage as each frame is decoded and logs it per turn with the session and the input that asked for the generation, plus a per-session summary naming how it ended (billing included). Adds scripts/analyze-gemini-usage.py to summarize those lines offline.
  • A 402 or depleted-credit close is retried only on another configured key; with none left the interview ends immediately.
  • Shortens the Live instructions (bundle 24, live prompt 16) and cuts candidate video to one low-resolution frame in five seconds.
  • Adds an optional GEMINI_CONTEXT_TRIGGER_TOKENS/GEMINI_CONTEXT_TARGET_TOKENS pair that silently checkpoints local state for compression experiments, leaving provider defaults unchanged until a measured comparison justifies switching. The comparison's replay harness now settles a tool turn the model never follows up, so later checkpoints are not refused.
  • codetrial web now treats a bad agent config entry as a hard error; only absent keys fall back to serving the web half alone.

Migration

  • docs/provider-cost-and-degradation.md documents the new logs and how to run the compression experiment; codetrial check-gemini confirms the model accepts the new setup message and compression pair.

Written for commit 80201fe. Summary will update on new commits.

Review in cubic

cubic-dev-ai[bot]

This comment was marked as resolved.

cubic-dev-ai[bot]

This comment was marked as resolved.

jserv added 2 commits October 1, 2026 12:09
An interview spends Gemini credit faster than users expect, and the
server could not say where: it logged one usage sum, could lose the last
turn's usage on shutdown, and treated a depleted prepaid account like
any outage, spending its restart budget on it. Usage is now recorded as
each frame is decoded and logged per turn, with its session and the
input that asked for the generation, and per session with how it ended,
billing included; scripts/analyze-gemini-usage.py summarizes those lines
offline. A billing failure is retried only on another key.

The Live instructions drop repeated explanation, and every generation is
billed on them again. Sent to gemini-3.1-flash-live-preview with one
identical candidate turn, alternating with the setup on main, the first
generation cost 5,152 prompt tokens against 5,650 in every run without a
tool call: 498 fewer on each generation, 8.8% of that turn, and about
1,000 fewer when a tool call bills the context a second time. Thought
tokens stayed at zero under the minimal level the family-keyed thinking
setting now picks. Candidate video, when enabled, sends a fifth of the
frames at the low media resolution. An optional compression pair adds
silent local-state checkpoints for measured experiments and leaves the
provider's defaults in place until a comparison says otherwise.
The web command read any agent config error as a missing Gemini key and
served the web side alone, so a single host with a mistyped optional
entry, such as an inverted compression pair, left every interview
waiting for an interviewer that would never join. Only missing keys now
mean the web half of a split deployment; an invalid entry stops the
server with the error.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant