aufklarer/Omnilingual-ASR-CTC-300M-MLX-4bit
Omnilingual ASR — CTC 300M (MLX 4-bit)
MLX-compatible 4-bit quantization of Meta's Omnilingual ASR CTC-300M model, targeting on-device inference on Apple Silicon (M1/M2/M3/M4).
Omnilingual ASR is a wav2vec 2.0–style encoder-only model with a linear CTC head, trained by Meta for speech recognition across 1,600+ languages. The CTC variant is language-agnostic at inference time (no language hint needed).
Model
Files
Architecture
Raw audio [1, samples]
→ Wav2Vec2FeatureExtractor (7-layer 1D conv, stride 320×)
→ Linear 512 → 1024
→ Wav2Vec2PositionEncoder (weight-normalized conv, kernel 128, groups 16)
→ 24 × StandardTransformerEncoderLayer (pre-norm, dim 1024, heads 16, ffn 4096)
→ LayerNorm
→ Linear 1024 → 10288 (CTC head)
→ logits [1, T/320, 10288]CTC greedy decoding with duplicate collapsing over the argmax path.
Performance
FLEURS test set, CTC-300M fp32 on CPU (Apple M-series), 30 utterances/language, aggregate WER via exact-edit-distance scorer (no external text normalization):
Aggregate CPU RTF ≈ 0.05; on M-series GPU via MLX, expect RTF < 0.02. (4-bit quantization typically adds <1% absolute WER on wav2vec2-class models; treat these as close upper bounds for the quantized variant.)
Usage
import mlx.core as mx
from mlx.utils import tree_unflatten
from safetensors import safe_open
weights = {}
with safe_open("model.safetensors", framework="mlx") as f:
for k in f.keys():
weights[k] = f.get_tensor(k)
# Your MLX wav2vec2 + CTC implementation consumes these keys.
# Expected input : float32 audio [1, samples] at 16 kHz, zero-mean unit-var
# Expected output: logits [1, T, 10288] then CTC greedy decode via the
# tokenizer in tokenizer.modelSwift inference is provided by speech-swift (see Sources/OmnilingualASR/).
Source
- Upstream model: facebook/omniASR-CTC-300M
- Paper: *Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages*
- Meta blog: Omnilingual ASR announcement
Links
- speech-swift — Apple SDK
- soniqo.audio — website
- blog
License
Apache 2.0 (inherited from upstream).
- Guide: soniqo.audio/guides/omnilingual
- Docs: soniqo.audio
- GitHub: soniqo/speech-swift
