CoolFace
Modelpublic

yehoshua01/waxal-whisper-turbo-lin-r2

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes10downloads
Model Card

Lingala ASR — whisper-large-v3-turbo fine-tune, round 2

Component of the 4th-place solution to the Google WAXAL ASR Challenge (Zindi, phase 2): 892 unseen clips, two African languages, no language metadata, scored 1 - (WER + CER) / 2 on raw text. Private leaderboard 0.771848284.

Code, full method and one-command verification: [yehoshua0/waxal-asr-phase2](https://github.com/yehoshua0/waxal-asr-phase2) The repository reproduces the submitted CSV byte for byte on a laptop in about a minute, and re-decodes every input from the audio on rented GPUs in about four hours.

Role in the system

Variant, NOT the one in the shipped chain. Round 2 scores better on dev than round 1 and is the wrong file for reproduction — the voter was decoded from r1. The two differ on 9.89% of words.

What it measured

Kept published precisely because the near-identical names are a trap: reproduction follows what ran, not what is best.

Usage

python
from transformers import WhisperForConditionalGeneration, WhisperProcessor
import torch, soundfile as sf, torchaudio.functional as AF

proc = WhisperProcessor.from_pretrained("yehoshua01/waxal-whisper-turbo-lin-r2")
model = WhisperForConditionalGeneration.from_pretrained("yehoshua01/waxal-whisper-turbo-lin-r2").eval().cuda()
tok = proc.tokenizer
# the language slot is a LEARNED decoder state -- use the token this checkpoint trained under
forced = [(1, tok.convert_tokens_to_ids("<|sw|>")),
          (2, tok.convert_tokens_to_ids("<|transcribe|>")),
          (3, tok.convert_tokens_to_ids("<|notimestamps|>"))]

wav, sr = sf.read("clip.wav", dtype="float32")
wav = AF.resample(torch.from_numpy(wav), sr, 16_000).numpy()[:30 * 16_000]
f = proc.feature_extractor(wav, sampling_rate=16_000, return_tensors="pt").input_features
out = model.generate(f.to("cuda", model.dtype), forced_decoder_ids=forced, max_new_tokens=220)
print(proc.batch_decode(out, skip_special_tokens=True)[0])

The rest of the system

code, method, verificationyehoshua0/waxal-asr-phase2
cached decodes and chain inputs`yehoshua01/waxal-phase2-chain-inputs`
all checkpoints`yehoshua01` on the Hub

Sibling checkpoints (primaries, voters and ablations of the same system): `waxal-mms-1b-lin-pl2-spk` · `waxal-sunbird51-sna-pl2-spk` · `waxal-whisper-turbo-lin-r1` · `waxal-qlora-largev3-lin` · `waxal-omni-ctc1b-lin` · `waxal-omni-ctc1b-sna` · `waxal-sunbird51-lin-ft-r2` · `waxal-sunbird51-lin-ft-light` · `waxal-mms-1b-lin-full` · `waxal-mms-1b-lin-fullmeta` · `waxal-ssa-hubert-lin`

Licence and intended use

apache-2.0. Training data is `google/WaxalNLP` (CC-BY-SA-4.0, share-alike), so derivatives carry that too.

These weights are not a general-purpose ASR model. Pseudo-labels were computed on the phase-2 test audio (transductive self-training, permitted for phase-2 training by the host), so the checkpoint is partly adapted to that specific set.