aufklarer/Kokoro-82M-CoreML
53.3k
Kokoro-82M CoreML
End-to-end CoreML export of hexgrad/Kokoro-82M at FP16, optimized for Apple Neural Engine. Requires iOS 18+ / macOS 15+.
A single kokoro_5s.mlmodelc runs the full pipeline (BERT → duration prediction → fixed-shape alignment → prosody → decoder) in one CoreML call. G2P (grapheme-to-phoneme) is a separate pair of CoreML models.
Looking for a smaller variant? See `aufklarer/Kokoro-82M-CoreML-INT8` — INT8 k-means palettized, 83 MB vs 325 MB here, with log-spec distance 0.42 vs this FP16 reference on a validation utterance.
Model
Files
Voices
54 preset voices across 10 languages: English (US/UK), Spanish, French, Hindi, Italian, Japanese, Korean, Portuguese, Chinese.
Usage
Add speech-swift to Package.swift:
.package(url: "https://github.com/soniqo/speech-swift", branch: "main")Then synthesize:
import KokoroTTS
let tts = try await KokoroTTSModel.fromPretrained(
modelId: "aufklarer/Kokoro-82M-CoreML"
)
let audio = try await tts.synthesize(
"Hello world, this is a Kokoro test.",
voice: "af_heart"
)CLI:
swift run audio kokoro "Hello world" --voice af_heart --output out.wavSource
- Base model: hexgrad/Kokoro-82M (Apache-2.0)
- Dictionaries and G2P: Apache-2.0
License
- Model weights: Apache-2.0
- CoreML conversion: Apache-2.0
Links
- speech-swift — Apple SDK
- soniqo.audio — website
- MLX vs CoreML on Apple Silicon — a practical guide — related blog post
- soniqo.audio/blog — blog
