Skip to content

fix(subtitles): keep cold client setup off the audio event loop - #180

Merged
Lucas1479 merged 2 commits into
mainfrom
codex/subtitle-client-cold-start
Oct 10, 2026
Merged

Lucas1479 merged 2 commits into
mainfrom
codex/subtitle-client-cold-start

Conversation

@Lucas1479

@Lucas1479 Lucas1479 commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

What and why

The first presentation subtitle builds its synchronous SDK/HTTP client on the shared event loop before offloading the provider request. Even though translation is scheduled in the background, cold client setup can delay an already-ready TTS task.

Move client acquisition into the existing request worker. Protect the existing client cache during construction so simultaneous first subtitles reuse one client; provider requests remain concurrent outside that lock. Subtitle content, provider options, playback ordering and speech queueing are unchanged.

Change class

  • Routine implementation/performance fix

Owning layer: presentation subtitle translator's synchronous client boundary.

User-visible effect: first subtitle initialization no longer blocks unrelated audio/event-loop work.

Compatibility or migration impact: none; no schema, settings, protocol or dependency changes. This PR is independent of #179 and does not alter foreground/background narration ordering.

Evidence

  • New regression tests on unmodified main: 2 failed, 1 passed. After the fix they pass for both OpenAI-compatible providers, concurrent client reuse, concurrent requests and failed initialization recovery.
  • python -X utf8 -m pytest -q tests/test_wallpaper_subtitle_translator.py tests/test_chat_translation_subtitle.py tests/test_pre_translation_runtime.py tests/test_playback_mouth_timing.py tests/test_chat_role_delivery.py tests/test_tts_threading.py — 143 passed.
  • ruff check ., both architecture/documentation generation checks, and git diff --check — passed.
  • Local Windows probe with real cold SDK/HTTP construction and stubbed provider responses, three alternating pairs: ready event-loop callback delay median 517.996 ms → 0.033 ms. Fixed samples were 0.028–0.320 ms. No network/model request or audio playback was involved; this measures loop blocking, not end-to-end acoustic latency.
  • All 14 CI checks passed on b7ebc12b4248fd038a3b08b2519544e44a24eb31, including all three Windows full-suite shards, model-less validation, desktop smoke and platform/voice profiles. Windows CI run.

Final check

  • Two-file fix for one implementation defect
  • Existing client reuse and narration ordering preserved
  • No private logs, credentials, conversation content or runtime state included

@Lucas1479
Lucas1479 marked this pull request as ready for review October 10, 2026 20:04
@Lucas1479
Lucas1479 merged commit 7753c22 into main Oct 10, 2026
14 checks passed
@Lucas1479
Lucas1479 deleted the codex/subtitle-client-cold-start branch October 10, 2026 20:38
Lucas1479 added a commit that referenced this pull request Oct 10, 2026
## What and why

Both VN translation entrypoints construct the synchronous SDK/HTTP
client on the shared event loop before offloading the provider request.
Cold client setup can therefore stall already-ready audio work in
Japanese-to-Chinese subtitles and Chinese-to-Japanese streaming speech
translation.

Move client acquisition into each existing request worker and protect
the shared cache during construction. Provider calls and synchronous
stream iteration remain outside the lock, so subtitle and speech
requests can still run concurrently.

Independent follow-up to #180, based on current main
`5351dd710d3ea68f533422f098556fa4b078cc21`. The defect is present in
unmodified upstream at that revision; no local customization is
included. This PR does not depend on #180 merging.

## Change class

- [x] Routine implementation/performance fix

Owning layer: VN Bridge's synchronous translation-client boundary.

User-visible effect: cold VN translation client setup no longer occupies
the audio event loop.

Compatibility or migration impact: none. No provider options, prompts,
settings, dependencies, speech queueing or playback order changes.

## Evidence

- New regression tests on unmodified main: **3 failed, 3 passed**; fixed
version: **6 passed**. Covers both OpenAI-compatible providers, shared
cold/warm clients across subtitle and streaming paths, concurrent
requests, initialization failure recovery, incremental ordered pieces,
and cancellation during initialization/iteration without late delivery.
- Test ordering uses explicit events/barriers, with 15-second deadlock
watchdogs rather than short latency thresholds.
- `python -X utf8 -m pytest -q tests/test_vn_translation_clients.py
tests/test_vn_subtitle_grouping.py
tests/test_vn_overlay_subtitle_order.py tests/test_vn_speech_segments.py
tests/test_vn_tts_enqueue_receipt.py
tests/test_character_prompt_snapshots.py
tests/test_pre_translation_runtime.py
tests/test_chat_translation_subtitle.py tests/test_tts_threading.py
tests/test_playback_mouth_timing.py` — **139 passed**.
- `python -m ruff check .`, `python tools/architecture/generate_views.py
--check`, `python tools/architecture/generate_readme_overview.py
--check`, and `git diff --check` — passed.
- Offline Windows x64 / CPython 3.12.10 / OpenAI 2.38.0 / HTTPX 0.28.1
probe: real SDK/HTTP construction with synthetic input/key and stubbed
completion responses; fresh subprocesses, warmed executor, three
alternating baseline/fixed pairs per path. Median ready event-loop
callback delay:
  - Subtitle: **462.876 ms → 0.093 ms** (fixed range 0.085–0.095 ms).
- Streaming speech translation: **479.785 ms → 0.167 ms** (fixed range
0.159–0.179 ms).

The probe measures event-loop blocking, not network first-token latency
or end-to-end audible first sound. No network/model request or audio
playback was performed. Client construction still takes time in its
worker. Cancellation stops the async consumer and discards late results;
it cannot forcibly interrupt synchronous work already running in
`to_thread`. Existing stream lifecycle and API-key cache identity are
unchanged.

Full remote CI is pending.

## Final check

- [x] Two-file fix for one implementation defect
- [x] Existing client reuse, translation options and speech ordering
preserved
- [x] No speculative shared client factory or unrelated cleanup
- [x] No credentials, private logs, conversations or runtime state
included
Lucas1479 added a commit that referenced this pull request Oct 10, 2026
Prepare `v0.16.1-alpha.0` as the next source Alpha after
`v0.16.0-alpha.0`, covering merged #177–#181: automatic NVIDIA TTS
acceleration, complete terminal Work context, ASR-overlapped application
preparation, and cold translation-client setup outside the audio event
loop.

Synchronize Python/Electron product versions and lock metadata, update
both README release links, include the new release/acceptance documents
in the source ZIP, and test native settings upgrade from 0.16.0. The
native upgrade probe also resolves history from the repository root and
preserves the previous tag's settings/shared-module layout: its old
single-file extraction failed on 0.16.0's catalog imports. No dependency
selection or production behavior changes are added by this preparation
PR.

Validation completed locally: 128 release/voice/entry regression tests;
`uv lock --check --offline`; version/lock consistency; third-party
provenance release gate; `git diff --check`; clean source-archive gate
(3,891 selected files, zero errors/warnings). All 3,891 ZIP file hashes
and its embedded manifest were verified. Local Electron build and native
encrypted-settings upgrade from both 0.16.0 and 0.15.2 passed; the
actual extracted ZIP was separately installed/built and passed the
0.16.0 native upgrade; the initial module-extraction failure and repair
are recorded in acceptance. Full final-head CI and extracted-source
install/build/package/startup are the merge gates.

The acceptance record contains sanitized real microphone observations
from #179/#180 and explicitly separates them from a controlled A/B,
acoustic measurement, ASR speedup, physical barge-in acceptance or #181
VN acceptance. Existing dependency findings and experimental-platform
limits remain. User configuration, raw logs and external media are
excluded.

After checks pass, publication uses the clean merged commit, the
repository's source-Alpha Latest convention, a verified tag, final
source ZIP, per-file manifest and SHA-256 checksums. Related: #176.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant