billingsmoore/tibetan-asr-nict-tib1-whisper-tiny
Whisper Tiny — Tibetan ASR (full fine-tune)
Fine-tuned `openai/whisper-tiny` for automatic speech recognition (ASR) on Lhasa Tibetan, released alongside:
J. Moore, S. Li and P. Lauren, "Evaluating Tibetan ASR With Segmented Word Error Rate: Beyond Character-Level Metrics," in IEEE Access, vol. 14, pp. 101790-101805, 2026, doi: 10.1109/ACCESS.2026.3709206.
This repo is one of 14 model checkpoints released with the paper, benchmarking 5 ASR architectures (Whisper Tiny/Base/Small, Wav2Vec 2.0 Base, HuBERT Base) under full fine-tuning, LoRA, and QLoRA (8-bit/4-bit) adaptation. See billingsmoore/tibetan-asr-nict-tib1-* for the full set.
Model description
Whisper is an encoder–decoder transformer, originally pre-trained by OpenAI on 680,000 hours of weakly-labeled multilingual speech-text data, fine-tuned here for sequence-to-sequence Tibetan transcription. This checkpoint uses full fine-tuning (all parameters updated, no quantization or adapters).
Training data
Fine-tuned on NICT-Tib1, a corpus of transcribed Lhasa Tibetan speech from 20 speakers (Soky, Gong & Li, 2022). The data was split by speaker (85/15, not by utterance) so that no speaker appears in both splits: 15,099 training utterances from 17 speakers, 1,547 test utterances from the remaining 3 speakers.
Training procedure
- Base model:
openai/whisper-tiny - Learning rate: 3.75e-5
- Max steps: 4000 (warmup 500)
- Effective batch size: 16 (per-device batch size 2, gradient accumulation 8)
- Mixed precision: fp16, with gradient checkpointing
- Generation max length: 225
Evaluation results
Evaluated on the 1,547-utterance NICT-Tib1 test set under five metrics: Character Error Rate (CER), Syllable Error Rate (SER), and three Segmented Word Error Rate (SWER) variants using different automatic word segmenters (BoTok-SWER, BERT-SWER, and Gem-SWER, the last computed on a 500-utterance subset for API cost reasons). All values are micro-averaged. See the paper for full bootstrap confidence intervals.
See the full architecture comparison table below for how this configuration ranks against the other models in the study.
Full model family comparison (standard fine-tuning)
How to use
from transformers import pipeline
pipe = pipeline("automatic-speech-recognition", model="billingsmoore/tibetan-asr-nict-tib1-whisper-tiny")
result = pipe("path/to/audio.wav")
print(result["text"])Limitations
- Trained and evaluated only on read-speech, modern Lhasa (Central) Tibetan news recordings; performance on Amdo or Kham dialects, Classical Tibetan, or spontaneous/conversational speech is unknown.
- The test set contains only 3 speakers (by design, to prevent speaker leakage), which limits how confidently results generalize across speaker, age, and prosodic variation.
- CER alone understates errors that matter at the word level for Tibetan; see the paper for why SER and SWER are necessary complements when judging output quality.
Citation
If you use this model, please cite the paper and the source NICT-Tib1 corpus:
@ARTICLE{11592371,
author={Moore, Jacob and Li, Sheng and Lauren, Paula},
journal={IEEE Access},
title={Evaluating Tibetan ASR With Segmented Word Error Rate: Beyond Character-Level Metrics},
year={2026},
volume={14},
number={},
pages={101790-101805},
keywords={Modeling;Automatic speech recognition;Error analysis;LoRa;Measurement;Ranking (statistics);Quantization (signal);Bit error rate;Standards;Training;Tibetan;automatic speech recognition;word error rate;low-resource language},
doi={10.1109/ACCESS.2026.3709206}
}
@inproceedings{soky2022nict,
title={Nict-tib1: A public speech corpus of lhasa dialect for benchmarking tibetan language speech recognition systems},
author={Soky, Kak and Gong, Zhuo and Li, Sheng},
booktitle={2022 25th Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA)},
pages={1--5},
year={2022},
organization={IEEE}
}License
Released under Apache License 2.0, inherited from the base model openai/whisper-tiny.
