janmbuys/omniASR-CTC-300m-v2-Zulu-Lwazi
isiZulu telephone ASR — CTC model + subword n-gram LM
A standalone Wav2Vec2ForCTC model: the `uctnlp/omniASR-CTC-300m-v2-Zulu-Baseline` checkpoint with a LoRA adaptation to 8 kHz telephone speech already merged in, plus a 6-gram subword language model and the lexicon needed to decode with it.
Loads with plain from_pretrained — no peft, no separate base download:
from transformers import Wav2Vec2ForCTC, Wav2Vec2CTCTokenizer
model = Wav2Vec2ForCTC.from_pretrained("<repo-id>")
tok = Wav2Vec2CTCTokenizer.from_pretrained("<repo-id>")The LoRA adapter is also kept under lora/ for provenance, should you want to re-merge it onto a different base.
Results
Lwazi isiZulu, scored against fragment-resolved references (lm_weight=0.75, word_score=-0.5, beam 100):
For reference, the same LM on 48 kHz studio speech (African Next Voices isiZulu dev, 3,063 utts) gives 21.16% → 17.10% WER. Lwazi is much harder: narrowband telephone audio, spontaneous speech, ~10% OOV.
Two things that will bite you
1. The space token
tokenizer_config.json in the base checkpoint declares word_delimiter_token: "▁", but `▁` is not in `vocab.json`. The real word delimiter is a literal space, id 4.
Decoding with the shipped setting happens to work — no ▁ is ever emitted, so the replacement is a no-op. But encoding maps every space to <unk> (id 3). Fine-tuning with that tokenizer trains the model to emit <unk> at every word boundary: the loss falls steadily while WER climbs past 100%.
This repo ships a corrected `tokenizer_config.json`, so Wav2Vec2CTCTokenizer.from_pretrained("<repo-id>") is already right:
tok = Wav2Vec2CTCTokenizer.from_pretrained("<repo-id>")
assert tok("a b").input_ids == [3859, 4, 3425] # 4 is the space, not <unk>If you point at the original base checkpoint instead, override it explicitly:
tok = Wav2Vec2CTCTokenizer("vocab.json", unk_token="<unk>", pad_token="<s>",
bos_token="<s>", eos_token="</s>",
word_delimiter_token=" ") # NOT "▁"2. The model is character-level, not subword
Despite the base model card's wording, this checkpoint emits characters: all 10,284 non-special entries in vocab.json are single characters. The BPE subwords here are language model units, not acoustic units. CTC blank is pad_token_id = 0.
How the LM integration works
audio -> CTC -> character posteriors -> flashlight LexiconDecoder -> text
| |
lexicon | | KenLM (6-gram)
subword -> characters over subwordsThe lexicon bridges the two. Word boundaries ride on a leading space:
▁ngi -> | n g i word-initial: consumes the preceding space
ya -> y a word-medial: does not
- -> -Consequences to respect:
- Prepend a space frame. Word-initial pieces spell a leading space, so the first word of an utterance is unreachable without one. Prepend a single frame with log-prob 0 on the space token and a large negative elsewhere.
- Set `sil` to blank, not to space. Boundaries already come from the leading-space spellings; also treating space as optional silence double-counts them and measurably hurts (with a positive
word_scoreit degenerates badly). - `lm_weight` ≈ 0.75 is optimal, well below the 1.5–2.5 typical of word-level lexicon decoding, because a subword LM fires several times more often per utterance. This value was optimal on both studio and telephone speech.
- Emissions are reduced. The model has 10,288 output classes but only 39 are reachable from the lexicon.
keep_ids.npyselects them; slice the posteriors and renormalise. This is ~250× smaller and much faster.
Contents
model.safetensors merged model (base + LoRA), 1.3 GB fp32
config.json Wav2Vec2ForCTC config
vocab.json 10,288 character classes
tokenizer_config.json CORRECTED: word_delimiter_token is " "
preprocessor_config.json 16 kHz, do_normalize
special_tokens_map.json
lora/ LoRA adapter, for provenance / re-merging
lm/lm6.bin KenLM 6-gram over BPE-32k subwords (binary trie)
lm/lexicon.txt subword -> character spellings
lm/tokens.txt 39-token reduced set ('#' = blank, '|' = space)
lm/keep_ids.npy column indices into the 10,288-class output
lm/bpe32000.model SentencePiece model (to re-encode text for LM training)
inference_example.pyUsage
Greedy decoding needs only transformers:
import torch, soundfile as sf
from transformers import Wav2Vec2ForCTC, Wav2Vec2CTCTokenizer
model = Wav2Vec2ForCTC.from_pretrained("<repo-id>").eval()
tok = Wav2Vec2CTCTokenizer.from_pretrained("<repo-id>")
wav, sr = sf.read("utt.wav", dtype="float32") # resample to 16 kHz first
wav = (wav - wav.mean()) / (wav.std() + 1e-7)
with torch.no_grad():
ids = model(torch.from_numpy(wav)[None]).logits.argmax(-1)[0]
print(tok.decode(ids.tolist()))LM decoding additionally needs flashlight-text built with KenLM; the PyPI wheel is not:
pip install torch torchaudio transformers soundfile numpy sentencepiece kenlm
USE_KENLM=1 CMAKE_POLICY_VERSION_MINIMUM=3.5 \
pip install --no-build-isolation --no-binary flashlight-text flashlight-textKenLM then lives at flashlight.lib.text.decoder.kenlm.KenLM — a submodule, not re-exported into flashlight.lib.text.decoder. On Anaconda you may also need conda install -c conda-forge "libstdcxx-ng>=13", since the bundled libstdc++ (3.4.29) is too old for the kenlm wheel.
# from a local clone of this repo
python3 inference_example.py --audio utt.wav # LM decoding
python3 inference_example.py --audio utt.wav --greedy # no LMTraining details
Batch size 1 with gradient accumulation 8, lr 1e-4 (OneCycle), 8 epochs, best epoch by validation WER. Batch 1 is deliberate: the feature extractor sets return_attention_mask=False, so the model was trained without an attention mask, and padding a batch would push unmasked silence through a stable-layer-norm encoder.
The adapter was trained on fragment-resolved transcriptions. Lwazi marks false starts, bound concord prefixes and truncations all with a trailing hyphen; naive cleaning turns elaw- elawini into elaw elawini and ku- Peter into ku peter. Training on resolved text instead was worth 3.9 WER points on test.
Limitations
- Tuned for 8 kHz telephone speech. On wideband audio the un-adapted base model may do better.
- Test sets are small (~3.5k reference words); differences under ~1.5 WER points are not meaningful.
- The LM is built from written web/news text, which mismatches spontaneous telephone speech; ~10% of Lwazi reference words are outside its vocabulary.
- No code-mixing evaluation was run in isolation.
Attribution
- Base model:
uctnlp/omniASR-CTC-300m-v2-Zulu-Baseline - Adaptation data: Lwazi ASR Corpus (CC BY 3.0) — Barnard, Davel & van Heerden, "ASR Corpus Design for Resource-Scarce Languages", Interspeech 2009
- LM data: MzansiText (Apache-2.0)
