CoolFace
Modelpublic

janmbuys/omniASR-CTC-300m-v2-Zulu-Lwazi

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes187downloads
Model Card

isiZulu telephone ASR — CTC model + subword n-gram LM

A standalone Wav2Vec2ForCTC model: the `uctnlp/omniASR-CTC-300m-v2-Zulu-Baseline` checkpoint with a LoRA adaptation to 8 kHz telephone speech already merged in, plus a 6-gram subword language model and the lexicon needed to decode with it.

Loads with plain from_pretrained — no peft, no separate base download:

python
from transformers import Wav2Vec2ForCTC, Wav2Vec2CTCTokenizer
model = Wav2Vec2ForCTC.from_pretrained("<repo-id>")
tok   = Wav2Vec2CTCTokenizer.from_pretrained("<repo-id>")

The LoRA adapter is also kept under lora/ for provenance, should you want to re-merge it onto a different base.

ModelWav2Vec2ForCTC, 326M params, character-level CTC (10,288 classes)
AdaptationLoRA r=16, α=32 on attention + FFN (7.1M params, 2.1%), merged
Adapted onLwazi isiZulu train split, 4,562 utterances (~7 h)
LM6-gram over SentencePiece BPE-32k subwords, KenLM
LM dataMzansiText isiZulu 71.6M words + English 5.9M (7.6%)

Results

Lwazi isiZulu, scored against fragment-resolved references (lm_weight=0.75, word_score=-0.5, beam 100):

systemvalidation WER / CER
base, greedy53.36 / 18.58
base + LM46.05 / 18.57
+ LoRA, greedy38.16 / 9.44
+ LoRA + LM31.58 / 8.78

For reference, the same LM on 48 kHz studio speech (African Next Voices isiZulu dev, 3,063 utts) gives 21.16% → 17.10% WER. Lwazi is much harder: narrowband telephone audio, spontaneous speech, ~10% OOV.

Two things that will bite you

1. The space token

tokenizer_config.json in the base checkpoint declares word_delimiter_token: "▁", but `▁` is not in `vocab.json`. The real word delimiter is a literal space, id 4.

Decoding with the shipped setting happens to work — no ▁ is ever emitted, so the replacement is a no-op. But encoding maps every space to <unk> (id 3). Fine-tuning with that tokenizer trains the model to emit <unk> at every word boundary: the loss falls steadily while WER climbs past 100%.

This repo ships a corrected `tokenizer_config.json`, so Wav2Vec2CTCTokenizer.from_pretrained("<repo-id>") is already right:

python
tok = Wav2Vec2CTCTokenizer.from_pretrained("<repo-id>")
assert tok("a b").input_ids == [3859, 4, 3425]      # 4 is the space, not <unk>

If you point at the original base checkpoint instead, override it explicitly:

python
tok = Wav2Vec2CTCTokenizer("vocab.json", unk_token="<unk>", pad_token="<s>",
                           bos_token="<s>", eos_token="</s>",
                           word_delimiter_token=" ")     # NOT "▁"

2. The model is character-level, not subword

Despite the base model card's wording, this checkpoint emits characters: all 10,284 non-special entries in vocab.json are single characters. The BPE subwords here are language model units, not acoustic units. CTC blank is pad_token_id = 0.

How the LM integration works

audio -> CTC -> character posteriors -> flashlight LexiconDecoder -> text
                                          |            |
                                  lexicon |            | KenLM (6-gram)
                     subword -> characters             over subwords

The lexicon bridges the two. Word boundaries ride on a leading space:

▁ngi    ->   | n g i        word-initial: consumes the preceding space
ya      ->   y a            word-medial: does not
-       ->   -

Consequences to respect:

  • —Prepend a space frame. Word-initial pieces spell a leading space, so the first word of an utterance is unreachable without one. Prepend a single frame with log-prob 0 on the space token and a large negative elsewhere.
  • —Set `sil` to blank, not to space. Boundaries already come from the leading-space spellings; also treating space as optional silence double-counts them and measurably hurts (with a positive word_score it degenerates badly).
  • —`lm_weight` ≈ 0.75 is optimal, well below the 1.5–2.5 typical of word-level lexicon decoding, because a subword LM fires several times more often per utterance. This value was optimal on both studio and telephone speech.
  • —Emissions are reduced. The model has 10,288 output classes but only 39 are reachable from the lexicon. keep_ids.npy selects them; slice the posteriors and renormalise. This is ~250× smaller and much faster.

Contents

model.safetensors        merged model (base + LoRA), 1.3 GB fp32
config.json              Wav2Vec2ForCTC config
vocab.json               10,288 character classes
tokenizer_config.json    CORRECTED: word_delimiter_token is " "
preprocessor_config.json 16 kHz, do_normalize
special_tokens_map.json
lora/                    LoRA adapter, for provenance / re-merging
lm/lm6.bin               KenLM 6-gram over BPE-32k subwords (binary trie)
lm/lexicon.txt           subword -> character spellings
lm/tokens.txt            39-token reduced set ('#' = blank, '|' = space)
lm/keep_ids.npy          column indices into the 10,288-class output
lm/bpe32000.model        SentencePiece model (to re-encode text for LM training)
inference_example.py

Usage

Greedy decoding needs only transformers:

python
import torch, soundfile as sf
from transformers import Wav2Vec2ForCTC, Wav2Vec2CTCTokenizer
model = Wav2Vec2ForCTC.from_pretrained("<repo-id>").eval()
tok = Wav2Vec2CTCTokenizer.from_pretrained("<repo-id>")
wav, sr = sf.read("utt.wav", dtype="float32")      # resample to 16 kHz first
wav = (wav - wav.mean()) / (wav.std() + 1e-7)
with torch.no_grad():
    ids = model(torch.from_numpy(wav)[None]).logits.argmax(-1)[0]
print(tok.decode(ids.tolist()))

LM decoding additionally needs flashlight-text built with KenLM; the PyPI wheel is not:

bash
pip install torch torchaudio transformers soundfile numpy sentencepiece kenlm
USE_KENLM=1 CMAKE_POLICY_VERSION_MINIMUM=3.5 \
  pip install --no-build-isolation --no-binary flashlight-text flashlight-text

KenLM then lives at flashlight.lib.text.decoder.kenlm.KenLM — a submodule, not re-exported into flashlight.lib.text.decoder. On Anaconda you may also need conda install -c conda-forge "libstdcxx-ng>=13", since the bundled libstdc++ (3.4.29) is too old for the kenlm wheel.

bash
# from a local clone of this repo
python3 inference_example.py --audio utt.wav              # LM decoding
python3 inference_example.py --audio utt.wav --greedy     # no LM

Training details

Batch size 1 with gradient accumulation 8, lr 1e-4 (OneCycle), 8 epochs, best epoch by validation WER. Batch 1 is deliberate: the feature extractor sets return_attention_mask=False, so the model was trained without an attention mask, and padding a batch would push unmasked silence through a stable-layer-norm encoder.

The adapter was trained on fragment-resolved transcriptions. Lwazi marks false starts, bound concord prefixes and truncations all with a trailing hyphen; naive cleaning turns elaw- elawini into elaw elawini and ku- Peter into ku peter. Training on resolved text instead was worth 3.9 WER points on test.

Limitations

  • —Tuned for 8 kHz telephone speech. On wideband audio the un-adapted base model may do better.
  • —Test sets are small (~3.5k reference words); differences under ~1.5 WER points are not meaningful.
  • —The LM is built from written web/news text, which mismatches spontaneous telephone speech; ~10% of Lwazi reference words are outside its vocabulary.
  • —No code-mixing evaluation was run in isolation.

Attribution

  • —Base model: uctnlp/omniASR-CTC-300m-v2-Zulu-Baseline
  • —Adaptation data: Lwazi ASR Corpus (CC BY 3.0) — Barnard, Davel & van Heerden, "ASR Corpus Design for Resource-Scarce Languages", Interspeech 2009
  • —LM data: MzansiText (Apache-2.0)