CoolFace
Modelpublic

billingsmoore/tibetan-asr-nict-tib1-whisper-base-lora-8bit

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes18downloads
Model Card

Whisper Base — Tibetan ASR (QLoRA, 8-bit)

Fine-tuned `openai/whisper-base` for automatic speech recognition (ASR) on Lhasa Tibetan, released alongside:

J. Moore, S. Li and P. Lauren, "Evaluating Tibetan ASR With Segmented Word Error Rate: Beyond Character-Level Metrics," in IEEE Access, vol. 14, pp. 101790-101805, 2026, doi: 10.1109/ACCESS.2026.3709206.

This repo is one of 14 model checkpoints released with the paper, benchmarking 5 ASR architectures (Whisper Tiny/Base/Small, Wav2Vec 2.0 Base, HuBERT Base) under full fine-tuning, LoRA, and QLoRA (8-bit/4-bit) adaptation. See billingsmoore/tibetan-asr-nict-tib1-* for the full set.

Model description

Whisper is an encoder–decoder transformer, originally pre-trained by OpenAI on 680,000 hours of weakly-labeled multilingual speech-text data, fine-tuned here for sequence-to-sequence Tibetan transcription. This checkpoint uses QLoRA fine-tuning (LoRA adapters trained on top of an 8-bit quantized base model).

LoRA configuration: r=8, alpha=16, dropout=0.1, targeting the attention projections (k_proj, q_proj, v_proj, out_proj) and both feed-forward layers (fc1, fc2) of every Whisper transformer block.

Base model weights are loaded in 8-bit precision (symmetric rounding, double quantization) via bitsandbytes before LoRA adapters are applied.

Training data

Fine-tuned on NICT-Tib1, a corpus of transcribed Lhasa Tibetan speech from 20 speakers (Soky, Gong & Li, 2022). The data was split by speaker (85/15, not by utterance) so that no speaker appears in both splits: 15,099 training utterances from 17 speakers, 1,547 test utterances from the remaining 3 speakers.

Training procedure

  • —Base model: openai/whisper-base
  • —Learning rate: 2.5e-5
  • —Max steps: 4000 (warmup 500)
  • —Effective batch size: 16 (per-device batch size 2, gradient accumulation 8)
  • —Mixed precision: fp16, with gradient checkpointing
  • —Generation max length: 225

Evaluation results

Evaluated on the 1,547-utterance NICT-Tib1 test set under five metrics: Character Error Rate (CER), Syllable Error Rate (SER), and three Segmented Word Error Rate (SWER) variants using different automatic word segmenters (BoTok-SWER, BERT-SWER, and Gem-SWER, the last computed on a 500-utterance subset for API cost reasons). All values are micro-averaged. See the paper for full bootstrap confidence intervals.

MetricValue
CER (micro)0.8086
SER (micro)0.9717
BoTok-SWER (micro)1.0059
BERT-SWER (micro, full 1,547-utt. test set)0.574
Gem-SWER (micro, 500-utt. subset)1.1821

Full model family comparison (standard fine-tuning)

ModelRepoCERSERBoTok-SWERBERT-SWERGem-SWER
HuBERT Basetibetan-asr-nict-tib1-hubert-base0.13520.36900.44770.1610.9653
Wav2Vec 2.0 Basetibetan-asr-nict-tib1-wav2vec2-base0.07450.21520.27470.0970.7447
Whisper Tinytibetan-asr-nict-tib1-whisper-tiny0.15600.23510.29750.1180.6759
Whisper Basetibetan-asr-nict-tib1-whisper-base0.14170.20830.26000.1050.6314
Whisper Smalltibetan-asr-nict-tib1-whisper-small0.11850.16920.20420.0860.5337
[!WARNING] This is a LoRA/QLoRA fine-tuned checkpoint. In the paper's benchmark, all LoRA and QLoRA configurations showed catastrophic word-level degradation relative to full fine-tuning of the same base model, despite retaining partial character-level accuracy. This repo is published for reproducibility of that (negative) result, not as a recommended deployment artifact. If you need a usable Tibetan ASR model, use the standard fine-tuned checkpoint instead.

How to use

python
from transformers import WhisperForConditionalGeneration, WhisperProcessor
from transformers import BitsAndBytesConfig
from peft import PeftModel

bnb_config = BitsAndBytesConfig(load_in_8bit=True, llm_int8_skip_modules=None)
base_model = WhisperForConditionalGeneration.from_pretrained(
    "openai/whisper-base", quantization_config=bnb_config, device_map="auto"
)
model = PeftModel.from_pretrained(base_model, "billingsmoore/tibetan-asr-nict-tib1-whisper-base-lora-8bit")
processor = WhisperProcessor.from_pretrained("openai/whisper-base", language="bo", task="transcribe")

# generate as usual, e.g.:
# inputs = processor(audio, sampling_rate=16000, return_tensors="pt")
# predicted_ids = model.generate(**inputs)
# transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)

Limitations

  • —Trained and evaluated only on read-speech, modern Lhasa (Central) Tibetan news recordings; performance on Amdo or Kham dialects, Classical Tibetan, or spontaneous/conversational speech is unknown.
  • —The test set contains only 3 speakers (by design, to prevent speaker leakage), which limits how confidently results generalize across speaker, age, and prosodic variation.
  • —CER alone understates errors that matter at the word level for Tibetan; see the paper for why SER and SWER are necessary complements when judging output quality.
  • —LoRA/QLoRA adapters showed severe sequence-level collapse in this study (SWER often exceeding 1.0, i.e. worse than a random-length hypothesis) despite retaining moderate CER — treat this checkpoint as a research artifact documenting that failure mode, not as a usable transcription model.

Citation

If you use this model, please cite the paper and the source NICT-Tib1 corpus:

bibtex
@ARTICLE{11592371,
  author={Moore, Jacob and Li, Sheng and Lauren, Paula},
  journal={IEEE Access},
  title={Evaluating Tibetan ASR With Segmented Word Error Rate: Beyond Character-Level Metrics},
  year={2026},
  volume={14},
  number={},
  pages={101790-101805},
  keywords={Modeling;Automatic speech recognition;Error analysis;LoRa;Measurement;Ranking (statistics);Quantization (signal);Bit error rate;Standards;Training;Tibetan;automatic speech recognition;word error rate;low-resource language},
  doi={10.1109/ACCESS.2026.3709206}
}

@inproceedings{soky2022nict,
  title={Nict-tib1: A public speech corpus of lhasa dialect for benchmarking tibetan language speech recognition systems},
  author={Soky, Kak and Gong, Zhuo and Li, Sheng},
  booktitle={2022 25th Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA)},
  pages={1--5},
  year={2022},
  organization={IEEE}
}

License

Released under Apache License 2.0, inherited from the base model openai/whisper-base.