aufklarer/Omnilingual-ASR-CTC-1B-MLX-8bit
Omnilingual ASR — CTC 1B (MLX 8-bit)
MLX-compatible 8-bit quantization of Meta's Omnilingual ASR CTC-1B model for on-device inference on Apple Silicon (M1/M2/M3/M4). Prefer this variant when you need the smallest possible WER regression from fp32 and can afford an extra ~460 MB compared to the 4-bit build.
Omnilingual ASR is a wav2vec 2.0-style encoder-only model with a linear CTC head, trained by Meta for speech recognition across 1,600+ languages. The CTC variant is language-agnostic at inference time.
Model
Files
Architecture
Wav2Vec2FeatureExtractor (7-layer CNN, 320× downsample) → Linear 512→1280 → conv position encoder → 48× pre-norm Transformer encoder (dim 1280, 20 heads, ffn 5120) → LayerNorm → Linear CTC head (→ 10,288 tokens).
Performance
See the 4-bit variant for architecture notes and the 300M reference for FLEURS WER across en/fr/de/ar/hi. The 1B model is ~3× the encoder capacity and delivers correspondingly lower WER on low-resource languages.
Source
- Upstream model: facebook/omniASR-CTC-1B
- Paper: *Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages*
- Meta blog: Omnilingual ASR announcement
Links
- speech-swift — Apple SDK
- soniqo.audio — website
- blog
License
Apache 2.0 (inherited from upstream).
- Guide: soniqo.audio/guides/omnilingual
- Docs: soniqo.audio
- GitHub: soniqo/speech-swift
