Native iOS implementation of Kyutai Pocket TTS using Rust/Candle.
This crate provides on-device text-to-speech for iOS using the Kyutai Pocket TTS model. It uses the Candle ML framework for inference and UniFFI for Swift bindings.
Pre-built XCFrameworks are available on the Releases page.
Each release includes:
PocketTTS.xcframework - iOS static library (device + simulator)See docs/INTEGRATION.md for detailed integration instructions.
A full-featured iOS demo app is included for testing and validation.
The demo app includes:
See tests/ios-harness/README.md for setup instructions.
┌─────────────────────────────────────────────────┐
│ Swift/SwiftUI App │
├─────────────────────────────────────────────────┤
│ Generated Swift Bindings (UniFFI) │
├─────────────────────────────────────────────────┤
│ PocketTTSEngine │
├─────────────────────────────────────────────────┤
│ FlowLM │ MLPSampler │ MimiDecoder │
│ (70M) │ (10M) │ (20M) │
└─────────────────────────────────────────────────┘
Rust toolchain with iOS targets:
rustup target add aarch64-apple-ios
rustup target add aarch64-apple-ios-sim
Xcode with iOS SDK
./scripts/build-ios.sh
This creates:
target/xcframework/PocketTTS.xcframework - Static librarytarget/xcframework/pocket_tts_ios.swift - Swift bindingsPocketTTS.xcframework into your Xcode projectpocket_tts_ios.swift to your Swift sourcesimport Foundation
// Initialize engine with model path
let modelPath = Bundle.main.path(forResource: "kyutai-pocket-ios", ofType: nil)!
let engine = try PocketTTSEngine(modelPath: modelPath)
// Configure
let config = TTSConfig(
voiceIndex: 0, // Alba
temperature: 0.7,
topP: 0.9,
speed: 1.0,
consistencySteps: 2,
useFixedSeed: false,
seed: 42
)
try engine.configure(config: config)
// Synthesize
let result = try engine.synthesize(text: "Hello, world!")
// result.audioData contains WAV bytes
The model files should be placed in:
kyutai-pocket-ios/
├── model.safetensors # Main model weights (225MB)
├── tokenizer.model # SentencePiece tokenizer (60KB)
└── voices/ # Voice embeddings (4.2MB)
├── alba.safetensors
├── marius.safetensors
├── javert.safetensors
├── jean.safetensors
├── fantine.safetensors
├── cosette.safetensors
├── eponine.safetensors
└── azelma.safetensors
Run latency tests to validate performance:
./scripts/run-latency-bench.sh --streaming # Measure TTFA
./scripts/run-latency-bench.sh --all # Test both modes
See docs/LATENCY_TESTING.md for detailed benchmarking instructions.
Why: When optimizing a complex ML pipeline like TTS, it's easy to introduce regressions—small changes that degrade speech quality in subtle ways. Without objective measurements, you might only notice quality degradation after it's too late, or worse, ship degraded audio to users.
The Challenge: Getting the last few percentage points of quality requires rigorous validation:
Our Solution: Comprehensive audio quality metrics with automated regression detection.
We measure five key aspects of TTS output quality:
| Metric | What It Measures | Target |
|---|---|---|
| WER (Word Error Rate) | Intelligibility via Whisper ASR | <5% excellent |
| MCD (Mel-Cepstral Distortion) | Spectral similarity to reference | <6 dB good |
| SNR (Signal-to-Noise Ratio) | Signal health and cleanliness | >25 dB excellent |
| THD (Total Harmonic Distortion) | Audio distortion level | <40% acceptable |
| Spectral (Centroid, Rolloff, Flatness) | Frequency characteristics | Tracked |
Every code change is validated automatically:
# Run quality check locally
cd validation
python quality_metrics.py \
--audio output.wav \
--text "Hello, this is a test." \
--whisper-model base \
--output-json quality_results.json
# Compare to baseline
python baseline_tracker.py \
--check-regression \
--baseline baselines/baseline_v0.4.1.json \
--metrics quality_results.json
Quality metrics run automatically in GitHub Actions:
See validation/README.md for detailed usage.
Before trusting quality metrics, we validate them against known cases:
Only after all validation runs pass do we establish the quality baseline.
Docs:
This system enables us to:
The last few percentage points of quality matter—they're the difference between "good enough" and "production ready."
An autoresearch-style optimization loop that autonomously improves TTS audio quality. An AI agent iteratively modifies parameters or code, evaluates against a composite quality score, keeps improvements, and discards regressions — looping indefinitely toward a perfect score.
How it works: Each iteration follows: REMEMBER → ANALYZE → CHECK (dead ends) → MODIFY one thing → EVALUATE → COMPARE → DECIDE (commit or reset) → RECORD → REPEAT. A persistent memory system tracks dead ends, promising leads, and learned rules across sessions so the agent never repeats mistakes.
# Phase 1: Establish baseline
python autotuning/autotune.py --phase baseline --model-dir kyutai-pocket-ios
# Phase 2: Sweep individual parameters
python autotuning/autotune.py --phase sweep --param temperature --model-dir kyutai-pocket-ios
# Phase 3: Joint optimization
python autotuning/autotune.py --phase optimize --iterations 100 --model-dir kyutai-pocket-ios
# Phase 4: Autonomous AI agent loop (start a fresh Claude Code session, paste autotuning/program.md)
The composite score combines intelligibility (40%, WER), acoustic similarity (25%, MCD), signal quality (15%, SNR), waveform correlation (10%), and low distortion (10%, THD) into a single 0-1 scalar.
See autotuning/README.md for details and docs/research/autoresearch-tts-adaptation.md for the full design document.
This project uses comprehensive development infrastructure:
See docs/quality/QUALITY_PLAN.md for details.
This implementation builds upon excellent work from:
See ATTRIBUTION.md for detailed attribution information.
MIT (code), CC-BY-4.0 (model weights)
161 followers · starred Feb 2026
Rust
45.9%
Python
36.9%
Swift
10.0%
Shell
6.7%
Native iOS implementation of Kyutai Pocket TTS using Rust/Candle.
This crate provides on-device text-to-speech for iOS using the Kyutai Pocket TTS model. It uses the Candle ML framework for inference and UniFFI for Swift bindings.
Pre-built XCFrameworks are available on the Releases page.
Each release includes:
PocketTTS.xcframework - iOS static library (device + simulator)See docs/INTEGRATION.md for detailed integration instructions.
A full-featured iOS demo app is included for testing and validation.
The demo app includes:
See tests/ios-harness/README.md for setup instructions.
┌─────────────────────────────────────────────────┐
│ Swift/SwiftUI App │
├─────────────────────────────────────────────────┤
│ Generated Swift Bindings (UniFFI) │
├─────────────────────────────────────────────────┤
│ PocketTTSEngine │
├─────────────────────────────────────────────────┤
│ FlowLM │ MLPSampler │ MimiDecoder │
│ (70M) │ (10M) │ (20M) │
└─────────────────────────────────────────────────┘
Rust toolchain with iOS targets:
rustup target add aarch64-apple-ios
rustup target add aarch64-apple-ios-sim
Xcode with iOS SDK
./scripts/build-ios.sh
This creates:
target/xcframework/PocketTTS.xcframework - Static librarytarget/xcframework/pocket_tts_ios.swift - Swift bindingsPocketTTS.xcframework into your Xcode projectpocket_tts_ios.swift to your Swift sourcesimport Foundation
// Initialize engine with model path
let modelPath = Bundle.main.path(forResource: "kyutai-pocket-ios", ofType: nil)!
let engine = try PocketTTSEngine(modelPath: modelPath)
// Configure
let config = TTSConfig(
voiceIndex: 0, // Alba
temperature: 0.7,
topP: 0.9,
speed: 1.0,
consistencySteps: 2,
useFixedSeed: false,
seed: 42
)
try engine.configure(config: config)
// Synthesize
let result = try engine.synthesize(text: "Hello, world!")
// result.audioData contains WAV bytes
The model files should be placed in:
kyutai-pocket-ios/
├── model.safetensors # Main model weights (225MB)
├── tokenizer.model # SentencePiece tokenizer (60KB)
└── voices/ # Voice embeddings (4.2MB)
├── alba.safetensors
├── marius.safetensors
├── javert.safetensors
├── jean.safetensors
├── fantine.safetensors
├── cosette.safetensors
├── eponine.safetensors
└── azelma.safetensors
Run latency tests to validate performance:
./scripts/run-latency-bench.sh --streaming # Measure TTFA
./scripts/run-latency-bench.sh --all # Test both modes
See docs/LATENCY_TESTING.md for detailed benchmarking instructions.
Why: When optimizing a complex ML pipeline like TTS, it's easy to introduce regressions—small changes that degrade speech quality in subtle ways. Without objective measurements, you might only notice quality degradation after it's too late, or worse, ship degraded audio to users.
The Challenge: Getting the last few percentage points of quality requires rigorous validation:
Our Solution: Comprehensive audio quality metrics with automated regression detection.
We measure five key aspects of TTS output quality:
| Metric | What It Measures | Target |
|---|---|---|
| WER (Word Error Rate) | Intelligibility via Whisper ASR | <5% excellent |
| MCD (Mel-Cepstral Distortion) | Spectral similarity to reference | <6 dB good |
| SNR (Signal-to-Noise Ratio) | Signal health and cleanliness | >25 dB excellent |
| THD (Total Harmonic Distortion) | Audio distortion level | <40% acceptable |
| Spectral (Centroid, Rolloff, Flatness) | Frequency characteristics | Tracked |
Every code change is validated automatically:
# Run quality check locally
cd validation
python quality_metrics.py \
--audio output.wav \
--text "Hello, this is a test." \
--whisper-model base \
--output-json quality_results.json
# Compare to baseline
python baseline_tracker.py \
--check-regression \
--baseline baselines/baseline_v0.4.1.json \
--metrics quality_results.json
Quality metrics run automatically in GitHub Actions:
See validation/README.md for detailed usage.
Before trusting quality metrics, we validate them against known cases:
Only after all validation runs pass do we establish the quality baseline.
Docs:
This system enables us to:
The last few percentage points of quality matter—they're the difference between "good enough" and "production ready."
An autoresearch-style optimization loop that autonomously improves TTS audio quality. An AI agent iteratively modifies parameters or code, evaluates against a composite quality score, keeps improvements, and discards regressions — looping indefinitely toward a perfect score.
How it works: Each iteration follows: REMEMBER → ANALYZE → CHECK (dead ends) → MODIFY one thing → EVALUATE → COMPARE → DECIDE (commit or reset) → RECORD → REPEAT. A persistent memory system tracks dead ends, promising leads, and learned rules across sessions so the agent never repeats mistakes.
# Phase 1: Establish baseline
python autotuning/autotune.py --phase baseline --model-dir kyutai-pocket-ios
# Phase 2: Sweep individual parameters
python autotuning/autotune.py --phase sweep --param temperature --model-dir kyutai-pocket-ios
# Phase 3: Joint optimization
python autotuning/autotune.py --phase optimize --iterations 100 --model-dir kyutai-pocket-ios
# Phase 4: Autonomous AI agent loop (start a fresh Claude Code session, paste autotuning/program.md)
The composite score combines intelligibility (40%, WER), acoustic similarity (25%, MCD), signal quality (15%, SNR), waveform correlation (10%), and low distortion (10%, THD) into a single 0-1 scalar.
See autotuning/README.md for details and docs/research/autoresearch-tts-adaptation.md for the full design document.
This project uses comprehensive development infrastructure:
See docs/quality/QUALITY_PLAN.md for details.
This implementation builds upon excellent work from:
See ATTRIBUTION.md for detailed attribution information.
MIT (code), CC-BY-4.0 (model weights)
161 followers · starred Feb 2026
Rust
45.9%
Python
36.9%
Swift
10.0%
Shell
6.7%