Cross-platform TTS (Text-to-Speech) wrapper with C ABI. Mirrors js-tts-wrapper and swift-tts-wrapper.
| Engine | Type | Credentials | Streaming | Voice List | Word Boundaries | Speech Markdown |
|---|---|---|---|---|---|---|
| System (speech-dispatcher) | Local | None | — (daemon plays) | — | Estimated | — |
| AVSpeech (macOS) | Local | None | — (system plays) | — | Estimated | — |
| SAPI (Windows) | Local | None | — (system plays) | — | Estimated (SAPI events) | — |
| Sherpa-ONNX | Local (1300+ models) | None | Sentence batches | Speakers | Estimated | — |
| Qwen3-TTS local | Local (qwen3-tts.cpp, GGML) | Model dir (+ lib via QWEN3_TTS_LIB) | After generation | 10 languages + any reference audio | Estimated | Stripped |
| floravox | Local (piper/MMS/Matcha/Kokoro ONNX — 1,100+ voices across ~1,100 languages) | Model dir | Streamed (per segment) | Filesystem scan | Measured (patched) / student / estimated | Native SSML |
| Pocket (Kyutai) | Local (phoneme ONNX bundle + any donor wav) | None | Frame-by-frame | Reference wav | Measured (attention tap) | Native SSML + IPA phonemes |
| Azure | Cloud | Key + Region | Real-time (WS) / Streamed (REST) | API | Real (WS) | Platform-aware |
| Microsoft Edge (Read Aloud) | Cloud | None (free) | Real-time (WS) | API | Real (WS) | Platform-aware |
| Google Cloud | Cloud | API Key | After response (JSON) | API | Real (v1beta1 timepoints) | Platform-aware |
| Google Gemini (3.8 TTS) | Cloud | API Key | After response (JSON) | API | Estimated (scaled) | Platform-aware |
| Qwen (Alibaba Cloud Model Studio / DashScope) | Cloud | API Key (+ optional Model/Region/Instruction) | Real-time (WS) | Static (14 system voices) | Real (word timestamps, voice-dependent) | Stripped |
| OpenAI | Cloud | API Key | Streamed | — | Estimated | Platform-aware |
| ElevenLabs | Cloud | API Key (+ optional Model/Voice/Language) | Streamed (JSON w/ timestamps) | API | Estimated | Platform-aware (v4 default: audio tags + inline IPA) |
| Cartesia | Cloud | API Key | Streamed | API | Estimated | Platform-aware |
| Deepgram | Cloud | API Key | Streamed | — | Estimated | Platform-aware |
| PlayHT | Cloud | API Key + User ID | Streamed | — | Estimated | Platform-aware |
| Fish Audio | Cloud | API Key | Streamed | — | Estimated | Platform-aware |
| Hume AI | Cloud | API Key | Streamed | — | Estimated | Platform-aware |
| Mistral | Cloud | API Key | Streamed | — | Estimated | Platform-aware |
| Murf | Cloud | API Key | Streamed | — | Estimated | Platform-aware |
| Resemble AI | Cloud | API Key | Streamed | — | Estimated | Platform-aware |
| Unreal Speech | Cloud | API Key | Streamed | — | Estimated | Platform-aware |
| UpliftAI | Cloud | API Key | Streamed | — | Estimated | Platform-aware |
| Amazon Polly | Cloud | Key + Secret + Region | Streamed | — | Estimated | Platform-aware |
| IBM Watson | Cloud | Key + Region + Instance | Chunked | — | Estimated | Platform-aware |
| Wit.ai | Cloud | Token | Chunked | — | Estimated | Platform-aware |
| xAI | Cloud | API Key | Chunked | — | Estimated | Platform-aware |
| ModelsLab | Cloud | API Key | Chunked | — | Estimated | Platform-aware |
- Streaming: Audio is delivered through the
on_audiocallback in chunks, as it becomes available. REST engines stream the response body as bytes arrive over the network (MP3 decoded to PCM16 mono incrementally on a background reader thread; raw-PCM providers pass straight through); Azure and Edge deliver real-time over WebSockets; Sherpa-ONNX delivers each sentence batch as it is synthesised (via the generate progress callback — a single-sentence utterance still completes before delivery). Exceptions: Google and ElevenLabswith-timestampsreturn one JSON document with base64 audio, so they can only deliver after the response completes (an API limitation, not buffering). Estimated word boundaries (engines without API timing data) fire progressively during streaming, anchored to delivered audio, rather than all at once when the response completes.
Third offline engine — qwen3-tts.cpp (MIT; models Apache-2.0), and the only local engine with zero-shot voice cloning. The C++ library is not vendored: build it once, then point the build at it.
scripts/build-qwen3-local.sh # clones + builds GGML into ~/spikes
cd ~/spikes/qwen3-tts.cpp && python scripts/setup_pipeline_models.py # one-time GGUF conversion
QWEN3_TTS_LIB=~/spikes/qwen3-tts.cpp cargo build --features qwen3-local// Speak with any reference audio — voice = path to a WAV:
engine.speak("Hello", Some("/path/to/reference.wav"), 1.0, 1.0, 1.0, Some(&mut cb), None, None)?;
// Or a language voice ("en", "fr", … — the 10 Qwen locales), "default",
// or an emb:<base64> speaker-embedding handle from the qwen3-local
// cloner (--features qwen3-local,cloning): embedding extracted once,
// reused per utterance. Fully offline — reference audio never leaves
// the machine.Behaviour: PCM16 mono 24 kHz output; SSML-in is stripped; volume is real PCM gain; rate/pitch have no upstream control; word boundaries are duration-scaled estimates; engine calls are serialized (upstream is not concurrent-safe); CPU synthesis is batch-grade (~0.1× realtime on 6 cores — use Metal/CUDA builds for speed). Cloning is timbre-grade: the upstream x-vector mode carries the speaker's timbre, while accent/prosody come from the model's language prior.
Bank a voice once, enroll it with every cloning-capable engine, speak it everywhere. Rust-only for now (no FFI); enable with --features cloning.
use rust_tts_wrapper::cloning::{create_cloner, VoiceCorpus};
// Import an Apple Personal Voice "Recordings" export (CAF/ALAC → PCM),
// or an LJSpeech corpus directory (metadata.csv + wav/).
let corpus = VoiceCorpus::from_personal_voice_zip("Will's Personal Voice 1 - Recordings.zip")?;
let identity = corpus.to_identity(Some("en"));
// Enroll: qwen + elevenlabs are instant; google/azure are consent-gated (below).
let cloner = create_cloner("qwen", r#"{"apiKey":"sk-..."}"#).ok_or("no cloner")?;
let handle = match cloner.clone_voice(&identity)? {
rust_tts_wrapper::cloning::CloneOutcome::Ready(h) => h,
_ => unreachable!("qwen cloning is instant"),
};
// Speak it — a cloned voice id is just a voice string; Qwen handles are
// model-bound, so pass the recorded model via modelId.
let engine = create_engine("qwen", r#"{"apiKey":"sk-...","modelId":"qwen-audio-3.0-tts-flash"}"#)?;
engine.speak("Hello", Some(&handle.voice_id), 1.0, 1.0, 1.0, Some(&mut audio_cb), None, None)?;
cloner.delete_cloned(&handle)?; // quota hygiene- Capture is source-agnostic:
VoiceIdentityBuilderrecords live microphone PCM clip-by-clip (start_clip/push_pcm/finish_clip),VoiceIdentity::from_transcribed_pairstakes audio files + transcripts,AudioClip::from_audio_fileloads wav/mp3/m4a/flac,VoiceCorpus::from_audio_dirimports a folder. Personal Voice and LJSpeech are just two importers. CloneRegistrypersists identity → engine handles (~/.rust-tts-wrapper/clones.json):
use rust_tts_wrapper::cloning::{default_registry_path, CloneRegistry};
let mut registry = CloneRegistry::load(default_registry_path().unwrap())?;
registry.add(&identity.name, handle.clone());
registry.save()?; // atomic write
let handles: &[rust_tts_wrapper::cloning::CloneHandle] = registry.handles(&identity.name);- Consent-gated engines (Google ICV, Azure) need a recording of their fixed script first — discoverable before you ask the user to record:
let cloner = create_cloner("google", creds)?;
if let Some(spec) = cloner.consent_spec() {
// record spec.script verbatim, then attach:
identity.consent.push(ConsentRecording {
engine: "google".into(),
pcm, // PCM16 LE mono
sample_rate: 24_000,
metadata: [("language_code".to_string(), "en-US".to_string())].into(),
});
}
// azure additionally requires voiceTalentName/companyName/locale metadata
// and is the one Job-mode engine: clone_voice may return
// CloneOutcome::Pending { job_id } — resolve it with cloner.poll_clone(&job_id).- Personal Voice exports are audio-only (no transcripts); attach them via
VoiceCorpus::phrasesor ASR when an engine needs them (Qwen doesn't). - Consent-gated providers are implemented behind their real-world gates: Google Chirp 3 ICV (
generateVoiceCloningKey, allow-listed projects —consent_spec()returns the exact script to record) and Azure Personal Voice (consent → personal voice → long-running-operationpoll_clone; intake-gated at aka.ms/customneural; requires aprojectIdCustom Voice project). Azure cloned voices speak via the"{base_model}/{speakerProfileId}"voice-string convention (mstts:ttsembedding). Neither is live-testable without vendor approval; request shapes are doc-verified and unit-tested. - qwen3-local cloner (with
--features qwen3-local,cloning): zero-shot and fully offline — extracts an ECAPA speaker embedding from the identity's best clip; the resultingemb:handle is reused per utterance without re-encoding. No upload, no consent gate, no quota (live-verified). - Job-based providers (Murf, Resemble) are not implemented yet.
- Cloning someone's voice requires their permission; banked-voice programs' licensed synthetic voices must not be re-cloned.
The whole banked-voice-in-iOS pipeline works with no external tools:
- Export on device. The user exports their Personal Voice from the
Accessibility → Personal Voice settings (a user-initiated on-device
export — apps cannot trigger it). The artifact is
"{Voice Name} - Recordings.zip". - What's inside. A
TrainingData/folder of{session5}_{NN}.cafclips — Core Audio Format / Apple Lossless, 48 kHz mono. An iOS 17-era bank is ~150 clips (~12 min) across several recording sessions; iOS 26 banks ~10 prompts. Both are far more than every instant cloner needs (10–20 s) — engines pick the best clips. - Import (pure Rust).
from_personal_voice_zipdecodes CAF/ALAC via symphonia (no ffmpeg), canonicalizes to PCM16 mono 24 kHz, skips any corrupt frames, and takes the voice name from the zip title:
use rust_tts_wrapper::cloning::{create_cloner, CloneOutcome, VoiceCorpus};
let corpus = VoiceCorpus::from_personal_voice_zip("Dad's Voice - Recordings.zip")?;
// corpus.name == "Dad's Voice"; ~150 clips; 24 kHz mono PCM
let identity = corpus.to_identity(Some("en"));
let cloner = create_cloner("qwen", r#"{"apiKey":"sk-..."}"#).ok_or("no cloner")?;
let CloneOutcome::Ready(handle) = cloner.clone_voice(&identity)? else { unreachable!() };
// speak with create_engine("qwen", ...) + Some(&handle.voice_id) as above- Transcripts are not in the zip (Apple includes audio only). If
an engine needs them, attach via
VoiceCorpus::phrases(prompt-list mapping by the{NN}filename index) or ASR; Qwen and ElevenLabs ignore transcripts entirely.
examples/voice-clone.rs runs this exact flow end-to-end
(--zip "… - Recordings.zip"), and tests/cloning_live.rs exercises
import → clone → list → speak → delete against a real export.
End-to-end demo: examples/voice-clone.rs. Live test: tests/cloning_live.rs (QWEN_API_KEY + QWEN_PV_ZIP, --ignored).
pocket-timing feature — Kyutai pocket-tts as a fully local engine: clone a
voice from any donor wav and speak IPA phonemes with attention-measured
word boundaries (estimated=false). Enable with
--no-default-features --features pocket-timing.
cargo run --release --no-default-features --features pocket-timing \
--example pocket-engine-demo -- <bundle_dir> <reference.wav> \
"w ˈɛ ɹ| ɪ z| ð ə| b æ θ ɹ u m"Two bundle layouts are auto-detected:
- Sherpa 6L (
lm_main*.onnx+vocab.json): orthographic text, Viterbi tokenizer - Phoneme 24L (
flow_lm_main*.onnx+bundle.json+tokenizer4.json): our published phoneme fine-tune (willwade/pocket-tts-en-phonemes-24l, CC-BY-4.0), WordLevel whitespace tokenizer,bos_before_voice.npyprepended to voice conditioning
Phoneme input uses | to mark word boundaries (the model tokenizes single
phones; the separators carry word grouping for the timing output):
"w ˈɛ ɹ| ɪ z| ð ə| b æ θ ɹ u m" → "Where is the bathroom"
SSML/SpeechMarkdown is compiled via floravox-ssml: <break> pauses become
real silence between segments, <mark> events fire at exact positions.
Honest caveats (measured, 2026-10):
- Zero-shot cloning is timbre-approximate: donor style dominates — clean narrator-style prompts read far more intelligibly than casual/filler-style donors. Validate clones for intelligibility, not just similarity.
- For a true personal voice, a
600-step adaptation fine-tune on a 15-min Personal Voice corpus embeds the speaker's timbre in the weights ($0.50 GPU; seefloravox/docs/PHONEME-POCKET.mdfor the recipe and the frozen-latent-stats requirement). - English (en-US phoneme inventory); first word of an utterance is occasionally fragile; 1-step flow decode is intended.
# Format and lint (required before commit)
cargo fmt --all && cargo clippy --all-targets --all-features -- -D warnings
# Run tests
cargo test --all-featuresCI requires: rustfmt check, clippy clean, and tests pass.
Active development. Engine constructors, the C ABI, and the offline test suite run in CI on Linux, macOS, and Windows. Live cloud API calls are not exercised in CI — see tests/live_cloud.rs.template (copy to tests/live_cloud.rs, gitignored) and .env.example for running them locally with your own credentials. Live SherpaOnnx synthesis IS exercised in CI by the sherpaonnx-live.yml workflow (downloads small VITS/Matcha/Kokoro models and runs tests/sherpaonnx_live.rs), triggered on PRs touching src/sherpaonnx_engine.rs and available as a manual workflow_dispatch.
- Voice List: Engines with "API" can enumerate voices from the provider's API.
- Word Boundaries: Google returns real timing via v1beta1 timepoints with SSML marks. All others use word-length-adjusted estimation (150 WPM baseline, configurable).
- Speech Markdown: Auto-detected and converted to platform-specific SSML via speechmarkdown-rust. Azure gets Microsoft SSML, Google gets Assistant SSML, ElevenLabs gets model-matched prompt markup (see below), others get Alexa SSML.
- Playback timeline (
timelinemodule): turn word boundaries into a(playback_time → byte_offset)timeline for reader-style highlighting —PlaybackTimeline(built fromsynth_with_boundariesoutput or boundary-callback tuples; keyed in playback time so a tempo/speed factor is compensated) plusPlaybackClock(pause/resume/seek wall-clock). Binary-search lookups hold the last position across gaps. Seeexamples/timeline-demo.rsfor the full pattern. - ElevenLabs markup: ElevenLabs parses no SSML documents. The default model is
eleven_v3, so SpeechMarkdown renders as audio tags ([whispers],[pause],[long pause], native"/IPA/"); set themodelIdcredential to a pre-v3 model (eleven_multilingual_v2,flash_v2_5,flash_v2) to get<break time>prompt markup (≤3s, clamped) instead. The dialect is chosen from the model because v3 reads stray XML aloud and pre-v3 models read audio tags aloud. Therateparameter maps to the deterministicvoice_settings.speedAPI setting (0.7–1.2).
pub trait TtsEngine: Send + Sync + Debug {
// Speaking
fn speak(&self, text: &str, voice: Option<&str>, rate: f32, pitch: f32, volume: f32,
on_audio: Option<OnAudioCallback>, on_boundary: Option<OnBoundaryCallback>) -> TtsResult<()>;
fn speak_with_options(&self, text: &str, options: Option<&SpeakOptions>,
on_audio: Option<OnAudioCallback>, on_boundary: Option<OnBoundaryCallback>) -> TtsResult<()>;
fn speak_sync(&self, text: &str, voice: Option<&str>, rate: f32, pitch: f32, volume: f32,
on_audio: Option<OnAudioCallback>, on_boundary: Option<OnBoundaryCallback>) -> TtsResult<()>;
// Synthesis (no playback)
fn synth_to_bytes(&self, text: &str, voice: Option<&str>, rate: f32, pitch: f32, volume: f32) -> TtsResult<Vec<u8>>;
fn synth_to_bytes_with_options(&self, text: &str, options: Option<&SpeakOptions>) -> TtsResult<Vec<u8>>;
fn synth_with_boundaries(&self, text: &str, voice: Option<&str>, rate: f32, pitch: f32, volume: f32) -> TtsResult<(Vec<u8>, Vec<WordBoundary>)>;
// Control
fn stop(&self) -> TtsResult<()>;
fn pause(&self) -> TtsResult<()>;
fn resume(&self) -> TtsResult<()>;
// Introspection
fn get_voices(&self) -> TtsResult<Vec<Voice>>;
fn engine_id(&self) -> &'static str;
fn check_credentials(&self) -> TtsResult<bool>;
}pub type OnAudioCallback<'a> = &'a mut dyn FnMut(&[u8]);
pub type OnBoundaryCallback<'a> = &'a mut dyn FnMut(&str, f32, f32); // word, start_s, end_s
pub type OnStartCallback<'a> = &'a mut dyn FnMut();
pub type OnEndCallback<'a> = &'a mut dyn FnMut();
pub type OnErrorCallback<'a> = &'a mut dyn FnMut(&str);pub struct Voice {
pub id: String,
pub name: String,
pub gender: Gender, // Male | Female | Unknown
pub provider: String,
pub language_codes: Vec<LanguageCode>,
}
pub struct LanguageCode {
pub bcp47: String, // "en-US"
pub iso639_3: String, // "eng"
pub display: String, // "English (United States)"
}
pub struct WordBoundary {
pub text: String,
pub offset: u64, // milliseconds
pub duration: u64, // milliseconds
}
pub struct SpeakOptions {
pub rate: Option<f32>,
pub speech_rate: Option<SpeechRate>, // XSlow | Slow | Medium | Fast | XFast
pub pitch: Option<f32>,
pub speech_pitch: Option<SpeechPitch>, // XLow | Low | Medium | High | XHigh
pub volume: Option<f32>,
pub voice: Option<String>,
pub format: Option<AudioFormat>, // Mp3 | Wav | Ogg | Opus | Aac | Flac | Pcm
pub use_speech_markdown: bool,
pub use_word_boundary: bool,
pub raw_ssml: bool,
pub extra: HashMap<String, String>,
}
pub enum Gender { Male, Female, Unknown }
pub enum AudioFormat { Mp3, Wav, Ogg, Opus, Aac, Flac, Pcm }
pub enum SpeechRate { XSlow, Slow, Medium, Fast, XFast }
pub enum SpeechPitch { XLow, Low, Medium, High, XHigh }// Word boundary estimation (matches Swift WordTimingEstimator)
pub fn estimate_word_boundaries(text: &str) -> Vec<WordBoundary>;
pub fn estimate_word_boundaries_with_wpm(text: &str, words_per_minute: f64) -> Vec<WordBoundary>;
// Speech Markdown preprocessing
pub fn preprocess_speech_markdown(text: &str, platform: &str) -> (String, bool);
// Gender normalization
pub fn normalize_gender(value: &str) -> Gender;pub fn create_engine(engine_id: &str, credentials_json: &str) -> Option<Box<dyn TtsEngine>>;
pub fn engine_count() -> usize;
pub fn engine_list() -> Vec<EngineDescriptor>;All functions are extern "C", #[no_mangle]:
| Function | Description |
|---|---|
tts_create(engine_id, credentials_json) |
Create engine, returns opaque tts_ctx* |
tts_destroy(ctx) |
Free engine context |
tts_speak(ctx, text) |
Speak (returns 0/-1) |
tts_speak_ssml(ctx, ssml) |
Speak pre-built SSML, bypassing rate/pitch/volume wrapping. Engines that parse SSML get it directly; ElevenLabs gets it translated via SpeechMarkdown into its model-matched dialect (it parses no SSML) |
tts_speak_sync(ctx, text) |
Speak (blocking) |
tts_stop(ctx) |
Stop speech |
tts_pause(ctx) |
Pause in-progress speech |
tts_resume(ctx) |
Resume paused speech |
tts_synth_to_bytes(ctx, text, out_bytes, out_len) |
Synth to buffer (returns 0/-1) |
tts_free_bytes(bytes, len) |
Free buffer from tts_synth_to_bytes |
tts_get_voices(ctx, out_voices, out_count) |
Get voice list |
tts_free_voices(voices, count) |
Free voice array |
tts_set_voice(ctx, voice_id) |
Set voice |
tts_set_rate(ctx, rate) |
Set rate (1.0 = normal) |
tts_set_pitch(ctx, pitch) |
Set pitch (1.0 = normal) |
tts_set_volume(ctx, volume) |
Set volume (1.0 = normal) |
tts_set_on_audio(ctx, cb, userdata) |
Set streaming audio callback |
tts_set_on_boundary(ctx, cb, userdata) |
Set word boundary callback: cb(word, byte_offset, byte_len, start_s, end_s, estimated, userdata). Offsets/lengths are bytes into the spoken text; an unlocatable word holds the last known offset with length -1 |
tts_set_on_viseme(ctx, cb, userdata) |
Set viseme callback for lip-sync |
tts_set_on_start(ctx, cb, userdata) |
Set speech-started callback |
tts_set_on_end(ctx, cb, userdata) |
Set speech-completed callback |
tts_set_on_error(ctx, cb, userdata) |
Set error callback |
tts_get_engine_count() |
Count registered engines |
tts_get_engines(out_engines, out_count) |
Get engine descriptors |
tts_free_engines(engines, count) |
Free engine info array |
tts_get_last_error(ctx) |
Get last error message |
#include "tts_wrapper.h"
#include <stdio.h>
void on_audio(const uint8_t* chunk, uintptr_t size, void* userdata) {
printf("Audio chunk: %zu bytes\n", size);
}
void on_boundary(const char* word, int32_t offset, int32_t len,
float start, float end, int32_t estimated, void* userdata) {
printf("Word '%s' %d+%d %.3f-%.3f %s\n", word, offset, len, start, end,
estimated ? "(estimated)" : "(measured)");
}
int main() {
tts_ctx* ctx = tts_create("openai", "{\"apiKey\":\"your-key\"}");
tts_set_on_audio(ctx, on_audio, NULL);
tts_set_on_boundary(ctx, on_boundary, NULL);
tts_set_voice(ctx, "alloy");
tts_speak_sync(ctx, "Hello world");
tts_destroy(ctx);
}use rust_tts_wrapper::{factory, types::SpeakOptions};
let engine = factory::create_engine("openai", r#"{"apiKey":"key"}"#).unwrap();
// Simple speak
engine.speak("Hello", Some("alloy"), 1.0, 1.0, 1.0, None, None).unwrap();
// With callbacks
let mut audio_cb = |chunk: &[u8]| println!("{} bytes", chunk.len());
let mut boundary_cb = |word: &str, s: f32, e: f32| println!("{}: {:.3}-{:.3}", word, s, e);
engine.speak_sync("Hello world", Some("alloy"), 1.0, 1.0, 1.0,
Some(&mut audio_cb), Some(&mut boundary_cb)).unwrap();
// With SpeakOptions
let opts = SpeakOptions { voice: Some("alloy".into()), ..Default::default() };
engine.speak_with_options("Hello", Some(&opts), None, None).unwrap();
// Synth to bytes
let audio = engine.synth_to_bytes("Hello", Some("alloy"), 1.0, 1.0, 1.0).unwrap();
// Get voices
for v in engine.get_voices().unwrap() {
println!("{} ({}) - {}", v.name, v.gender, v.primary_language());
}
// Check credentials
assert!(engine.check_credentials().unwrap());cargo build --all-featuressystem— speech-dispatcher (Linux system TTS)avsynth— AVSpeechSynthesizer (macOS system TTS)sapi— SAPI (Windows system TTS)cloud— all 20 cloud engines via HTTP + speechmarkdown-rust + base64sherpaonnx— Sherpa-ONNX offline TTS (1300+ models)floravox— floravox offline TTS (piper/MMS/Matcha/Kokoro ONNX voices, ~1,100 languages; native SSML —<break>/<prosody>/<mark>/<phoneme>/<sub>— with three-tier word timings: measured on patched voices, student sidecar, proportional. See the floravox Voices section)floravox-lexicons— adds the lexicon+Phonetisaurus G2P chain and thelangcredential that auto-fetches published bundles (non-English phoneme voices)pocket-timing— PocketTtsEngine: Kyutai pocket-tts ONNX bundle with real attention-measured word boundaries, voice cloning from any donor wav, and IPA-phoneme input (see the Pocket TTS section)cloning— voice banking: import an Apple Personal Voice zip or LJSpeech corpus, enroll with cloud cloning engines (see Voice cloning)
cargo fmt --all -- --check
cargo clippy --all-features -- -D warnings
cargo test --all-featuresEvery binding wraps the flat C ABI in include/tts_wrapper.h; see
bindings/README.md for the full guide (loading
conventions, test matrix, which package to use). All five suites — Rust
ABI conformance, a C harness compiled with -Wall -Wextra -Werror, Node,
.NET and Swift — run in CI on every push (.github/workflows/bindings.yml).
from tts_wrapper import TTSClient
client = TTSClient("openai", {"apiKey": "your-key"})
client.on_audio(lambda chunk: print(f"{len(chunk)} bytes"))
# word, byte_offset, byte_len, start_s, end_s, estimated
client.on_boundary(lambda w, off, ln, s, e, est: print(f"{w}: {s:.3f}-{e:.3f}{'~' if est else ''}"))
client.set_voice("alloy")
client.speak_sync("Hello world")
client.stop()using RustTtsWrapper;
using var client = new TtsClient("openai", new() { ["apiKey"] = "your-key" });
client.SetOnBoundary((word, offset, len, start, end, estimated) =>
Console.WriteLine($"{word}: {start:F3}-{end:F3} {(estimated ? "estimated" : "measured")}"));
client.SetVoice("alloy");
client.SpeakSync("Hello world");NuGet contents per RID (since 0.5.3): win-x64 and win-x86 bundle
one DLL with sapi + cloud + sherpaonnx (the sherpa builds carry
lexicon/G2P bundles, measured boundaries, SSML marks).
let client = try TtsClient(engineId: "openai", credentials: ["apiKey": "your-key"])
client.setOnBoundary { word, offset, len, start, end, estimated in
print("\(word): \(start)-\(end) \(estimated ? "estimated" : "measured")")
}
client.setVoice("alloy")
try client.speakSync("Hello world")const { TtsClient } = require("@aactools/tts-wrapper");
const client = new TtsClient({ engineId: "openai", credentials: { apiKey: "your-key" } });
client.on("boundary", ({ word, startSec, endSec, estimated }) =>
console.log(`${word}: ${startSec}-${endSec} ${estimated ? "estimated" : "measured"}`));
client.setVoice("alloy");
client.speakSync("Hello world");
client.close();bindings/c/tts_abi_harness.c exercises the whole ABI against the
cdylib; make -C bindings/c test builds, compiles the header with
-Wall -Wextra -Werror and runs it.
TtsEngine (trait)
|
+----------+----------+----------+----------+
| | | | |
SystemEngine CloudEngine SherpaOnnx Floravox PocketTts
(speech- (20 cloud (1300+ (student (cloning +
dispatcher) providers) models) fleet) phonemes)
Cloud engines use provider-specific CloudConfig:
- Azure: SSML XML body with prosody tags, XML escaping
- Google: JSON body with base64 audio, v1beta1 timepoint support
- All others: Standard JSON bodies
1300+ models from the sherpa-onnx-models registry crate (canonical: AACTools/sherpa-onnx-tts-models; updates are dependency bumps). Models are loaded from ~/.rust-tts-wrapper/sherpaonnx/.
The registry ships as the
sherpa-onnx-models crate,
published from AACTools/sherpa-onnx-tts-models
(tag crate-v*). Refresh it like any dependency:
cargo update -p sherpa-onnx-models # within the pinned 0.x line
# or bump the version in Cargo.toml for a new lineThe registry's enriched fields (license, sha256, voice_names,
min_sherpa_onnx_version, deprecated, …) are currently ignored by
parse_model but carried through for future opt-in — the sync is
backwards-compatible.
floravox (Apache-2.0 OR MIT, pure
Rust, no Python, no GPL) is an offline synthesis engine over four ONNX
voice families. Enable with --features floravox.
| Family | Files on disk | Languages (typical) | Notes |
|---|---|---|---|
| piper VITS | X.onnx + X.onnx.json |
~30 (per-voice) | the original target |
| MMS VITS | X.onnx + tokens.txt (+ config.json) |
~1,100 (per-voice, one language each) | the 1,138-voice patched collection ships on Hugging Face |
| Matcha | acoustic *.onnx + tokens.txt + vocoder (hifigan*/vocos*) |
per-voice | audio comes from the vocoder |
| Kokoro | model.onnx + tokens.txt + voices.bin |
en + zh (11 voices in en-v0.19) | multi-speaker via the speaker credential |
Every boundary carries an honest estimated flag saying which tier
produced it — the engine never re-estimates or upgrades a tier:
- Measured (
estimated: false) — duration-patched voices report word timings from the model's own duration tensor, sample-accurate.<break>lands at a real silence edge;<mark>fires at a measured position. - Student (
estimated: true) — a sibling<voice-stem>.studentfile engages the timing student automatically (~340 languages trained; median error 59 ms against the teacher). - Proportional (
estimated: true) — 150-wpm estimate, same as the other local engines.
<break>, <prosody rate>, <mark>, <phoneme>, <sub>, <say-as>
are parsed locally by floravox-ssml with byte-exact spans. <mark>
events surface through the on_mark callback and as zero-duration
measured boundaries. SpeechMarkdown input expands through the standard
pipeline into the SSML dialect floravox parses natively.
| Key | Meaning |
|---|---|
modelsDir |
voice directory (defaults to ~/.rust-tts-wrapper/floravox); a voice is a directory or flat pair holding X.onnx (+ .onnx.json / tokens.txt) |
modelId |
voice to load (bare stem, directory, or .onnx path); also selectable per call via voice |
misaki |
"us" (default) / "gb" — English document pre-pass (heteronyms, numbers); "off" disables |
chars |
character frontend for MMS-style voices: "true" lowercases through the voice's own table; any other value is an ISO 639-3 uroman code (e.g. "hin") |
speaker |
speaker id for multi-speaker voices (kokoro style slots, piper sid) |
- English: misaki (the phonemizer Kokoro voices were trained with —
heteronyms and numbers come out right). Dialect
us/gb. - MMS voices (1,100+ languages): character frontend, auto-detected; non-Latin scripts romanized with uroman.
- Non-English phoneme voices (German/French/… piper): the
lexicon+Phonetisaurus chain, enabled by
--features floravox-lexicons. Pointlexiconat a compiled lexicon stem (stem.fst+stem.pho, gruut-derived bundles from voicegarden-lexicons),phonetisaurusat a WFST for unseen words, or just passlang(e.g."de") and the published bundle is fetched automatically. - ByT5 (opt-in): set
byt5Encoder+byt5Decoderto the byt5-g2p-multilingual ONNX pair (~18.5 MB int8) and unseen words in ~130 languages resolve neurally instead of letter-spelling. Chain order: lexicon → Phonetisaurus → ByT5 → letter spelling.
wasm32 is available for the floravox offline engine via the published
floravox-wasm crate (ort-web
backend), consumed by the floravox-web demo; the JavaScript/Node package
in js/ (unified speak() across floravox + cloud engines) is built and
awaiting npm publish.
MIT