Nasimbahar/pashto-ghag-whisper-medium-asr
Pashto Ghag | پښتو غږ
Pashto Ghag is an open-source Pashto automatic speech recognition model built by fine-tuning openai/whisper-medium on 54.4 hours of validated Pashto speech.
The purpose of this release is simple: publish a strong, usable Pashto STT baseline that the community can test, improve, and build on.
I trained multiple checkpoints, then rescored them on three held-out healthy evaluation splits:
dev_cleandev_realtest_frozen
This release uses `checkpoint-16000`, because it was the strongest overall checkpoint in that rescore and achieved a 14.63% balanced mean WER.
The name Pashto Ghag means Pashto Voice. It keeps the repo searchable on Hugging Face while giving the project a clear Pashto identity.
Model Summary
- Base model:
openai/whisper-medium - Task: automatic speech recognition
- Language: Pashto (
ps) - Release checkpoint:
checkpoint-16000 - Checkpoint selection metric: balanced mean WER across three healthy-eval splits
- Training data scale:
73,355train clips /54.391hours - Release goal: a practical open-source Pashto STT baseline
Why This Checkpoint
The training run produced multiple checkpoints. Instead of publishing the last checkpoint by default, I rescored the stronger checkpoints on a cleaner held-out evaluation suite and selected the one with the best overall balance.
checkpoint-16000 ranked first by balanced mean WER, so it is the public release checkpoint.
Training Data
This checkpoint was trained on Pashto Common Voice validated speech with normalized Pashto transcripts.
In concrete terms, the work behind this release was:
- fine-tune
openai/whisper-mediumfor Pashto transcription - train on 54.4 hours of Pashto speech
- normalize and prepare Pashto transcripts for training
- evaluate multiple checkpoints
- release the strongest checkpoint as an open-source Pashto STT model
This is the current public baseline, not the final stage of the project. The next step is further data cleaning and stronger future Pashto STT releases.
Training split sizes from the local manifests:
Training configuration highlights:
- max clip length: 20 seconds
- train batch size: 2
- gradient accumulation: 8
- learning rate:
1e-5 - epochs configured: 10
- model selection metric during training:
eval_wer
Evaluation
Release checkpoint checkpoint-16000 was rescored on the healthy-eval suite. Lower is better.
Internal validation metrics from the training run were slightly different because checkpoint selection and later healthy-eval rescoring used different validation layouts. For the public release, the healthy-eval table above is the headline result.
Intended Use
This model is intended for:
- Pashto speech transcription
- subtitle generation
- search/indexing over Pashto speech archives
- research and downstream ASR benchmarking for Pashto
Limitations
- Whisper is natively a short-form model. Multi-minute audio should be transcribed with chunking rather than naive manual window concatenation.
- Performance will vary across dialects, microphones, room acoustics, and code-switched speech.
- Noisy far-field audio and overlapped speech remain hard cases.
- Like other Whisper models, decoding can degrade on long audio if the inference pipeline handles chunk merging poorly.
Quick Start
The standard Hugging Face transformers pipeline is the simplest way to use this model:
import torch
from transformers import pipeline
model_id = "Nasimbahar/pashto-ghag-whisper-medium-asr"
device = 0 if torch.cuda.is_available() else -1
dtype = torch.float16 if torch.cuda.is_available() else torch.float32
pipe = pipeline(
"automatic-speech-recognition",
model=model_id,
torch_dtype=dtype,
device=device,
chunk_length_s=30,
)
result = pipe(
"sample.wav",
generate_kwargs={"language": "pashto", "task": "transcribe"},
)
print(result["text"])Detailed usage notes, long-audio guidance, and lower-level examples are in USAGE.md.
Recommended Inference Practice
- For short clips: direct inference is fine.
- For long recordings: use
transformers.pipeline(..., chunk_length_s=30)or another timestamp-aware chunking method. - If you are building a web app, treat long audio as an asynchronous job instead of a blocking request.
Acknowledgements
- Base model: openai/whisper-medium
- Model card format and metadata conventions: Hugging Face model cards
- Upload workflow reference: Hugging Face upload guide
Citation
If you use this model in a paper, report, or product benchmark, cite:
- OpenAI Whisper
- Mozilla Common Voice
- this Hugging Face model repo
