Skip to content

feat(voice): F011 on-device TTS spoken replies - #12

Merged
nimat-dev merged 1 commit into
mainfrom
feat/F011
Oct 8, 2026
Merged

nimat-dev merged 1 commit into
mainfrom
feat/F011

Conversation

@nimat-dev

Copy link
Copy Markdown
Owner

Summary

This PR implements F011: On-device TTS spoken replies (toggle), completing Phase 04 (Voice — thin) and bringing the entire ShellMind roadmap to 100% completion (12 / 12 features complete)!

Changes Included:

  1. On-Device TTS Provider Architecture (@shellmind/mobile/src/voice):
    • tts-types.ts: Defines ITextToSpeechProvider interface and TTSOptions (rate?, pitch?, language?, callbacks: onStart?, onDone?, onError?):
      • isAvailable(): Promise<boolean>
      • speak(text: string, options?: TTSOptions): Promise<void>
      • stop(): Promise<void>
      • isSpeaking(): boolean
    • summary.ts: Implements extractSpokenSummary(text, maxChars = 300):
      • Removes markdown formatting (fenced code blocks, inline code ticks, images, links, bold/italic, header/bullet syntax).
      • Cleans leaked tool JSON objects.
      • Respects sentence boundary capping (. ! ? within maxChars limit) and word boundary fallbacks, strictly guaranteeing length <= maxChars.
    • mock-tts.ts: MockTextToSpeechProvider providing test simulation with speech history tracking, autocomplete delays, and error/availability injection.
    • native-tts.ts: NativeTextToSpeechProvider safely bridging iOS AVSpeechSynthesizer / Expo Speech / browser speech synthesis with safe non-throwing fallbacks.
    • registry.ts: Provider registry accessors (getTextToSpeechProvider, setTextToSpeechProvider, resetTextToSpeechProvider).
    • Re-exported cleanly via packages/mobile/src/voice/index.ts and packages/mobile/src/index.ts.
  2. Chat UI Integration (@shellmind/mobile/src/components/ChatScreen.tsx):
    • Persistent spoken replies toggle (testID="tts-toggle") with visual icons and status text (🔊 Voice On / 🔇 Voice Off).
    • Active speaking indicator banner (testID="speaking-indicator") with a pulse dot and interrupt/mute button (testID="tts-stop-button").
    • Turn completion auto-speak: when an assistant turn finishes (event.type === "done"), extracts the concise summary and calls speak().
    • Instant interruption: user prompt submission, voice recording start, turn abort, or manual mute button immediately stops ongoing speech.
    • Rapid turns safety: previous utterance is halted before new speech commences (zero audio overlap).
  3. Tests & Evidence:
    • Comprehensive unit and integration tests in packages/mobile/src/mobile.test.ts verifying provider lifecycle, stop/interruption, autocomplete, summary cleaning, length capping, rapid turn handling, and ChatScreen options.
    • 130/130 tests passing monorepo-wide (29 protocol, 52 agent, 49 mobile).
    • Clean architecture verified with dependency-cruiser (79 modules, 229 dependencies cruised, 0 violations).
    • Maestro E2E flow in .maestro/voice_tts_flow.yaml.
    • Architecture summary, test summary, and E2E trace recorded in .harness/evidence/F011/.

@nimat-dev

Copy link
Copy Markdown
Owner Author

Maker-Checker Review: F011 (On-device TTS spoken replies)

1. Acceptance Criteria Verification

  • On-device TTS provider abstraction ITextToSpeechProvider defined in packages/mobile/src/voice/tts-types.ts.
  • Mock provider MockTextToSpeechProvider supports deterministic test fixtures, history tracking, autocomplete, and error/availability simulation.
  • Native provider NativeTextToSpeechProvider bridges platform synthesizers with graceful fallback.
  • Summary extractor extractSpokenSummary cleans markdown, code fences, and tool JSON, capping to concise sentences <= maxChars.
  • Provider registry in packages/mobile/src/voice/registry.ts supports runtime swapping (getTextToSpeechProvider, setTextToSpeechProvider, resetTextToSpeechProvider).
  • Spoken replies toggle (testID="tts-toggle"), active speaking indicator (testID="speaking-indicator"), and mute button (testID="tts-stop-button") implemented in ChatScreen.tsx.
  • Turn completion triggers spoken reply when toggle is active; remains completely silent when toggle is inactive.
  • Active speech is immediately interrupted upon prompt submission, voice recording start, abort, or manual stop.
  • Rapid turns do not overlap (prior utterance terminated before subsequent utterance starts).
  • 130/130 tests pass across all packages (41 mobile tests).
  • Dependency cruiser reports 0 violations across 79 modules.
  • Pure core invariant preserved: voice TTS is mobile-only; protocol and agent remain pure and audio-agnostic.
  • Maestro E2E specification in .maestro/voice_tts_flow.yaml.
  • Harness docs and evidence logged in .harness/evidence/F011/.

2. Evaluation Scores

  • Acceptance: 5/5
  • Correctness: 5/5
  • Boundaries: 5/5
  • Modularity: 5/5
  • Evidence: 5/5
  • Average: 5.0 (PASS)

3. Decision

APPROVE. Ready for squash merge to main.

@nimat-dev
nimat-dev merged commit 192354f into main Oct 8, 2026
2 checks passed
@nimat-dev
nimat-dev deleted the feat/F011 branch October 8, 2026 12:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant