CoolFace
Modelpublic

Nasimbahar/pashto-ghag-whisper-medium-asr

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes226downloads
Model Card

Pashto Ghag | پښتو غږ

Pashto Ghag is an open-source Pashto automatic speech recognition model built by fine-tuning openai/whisper-medium on 54.4 hours of validated Pashto speech.

The purpose of this release is simple: publish a strong, usable Pashto STT baseline that the community can test, improve, and build on.

I trained multiple checkpoints, then rescored them on three held-out healthy evaluation splits:

  • —dev_clean
  • —dev_real
  • —test_frozen

This release uses `checkpoint-16000`, because it was the strongest overall checkpoint in that rescore and achieved a 14.63% balanced mean WER.

The name Pashto Ghag means Pashto Voice. It keeps the repo searchable on Hugging Face while giving the project a clear Pashto identity.

Model Summary

  • —Base model: openai/whisper-medium
  • —Task: automatic speech recognition
  • —Language: Pashto (ps)
  • —Release checkpoint: checkpoint-16000
  • —Checkpoint selection metric: balanced mean WER across three healthy-eval splits
  • —Training data scale: 73,355 train clips / 54.391 hours
  • —Release goal: a practical open-source Pashto STT baseline

Why This Checkpoint

The training run produced multiple checkpoints. Instead of publishing the last checkpoint by default, I rescored the stronger checkpoints on a cleaner held-out evaluation suite and selected the one with the best overall balance.

checkpoint-16000 ranked first by balanced mean WER, so it is the public release checkpoint.

Training Data

This checkpoint was trained on Pashto Common Voice validated speech with normalized Pashto transcripts.

In concrete terms, the work behind this release was:

  1. 1.fine-tune openai/whisper-medium for Pashto transcription
  2. 2.train on 54.4 hours of Pashto speech
  3. 3.normalize and prepare Pashto transcripts for training
  4. 4.evaluate multiple checkpoints
  5. 5.release the strongest checkpoint as an open-source Pashto STT model

This is the current public baseline, not the final stage of the project. The next step is further data cleaning and stronger future Pashto STT releases.

Training split sizes from the local manifests:

SplitSamplesHours
train73,35554.391
validation4,0753.041
test4,0763.041

Training configuration highlights:

  • —max clip length: 20 seconds
  • —train batch size: 2
  • —gradient accumulation: 8
  • —learning rate: 1e-5
  • —epochs configured: 10
  • —model selection metric during training: eval_wer

Evaluation

Release checkpoint checkpoint-16000 was rescored on the healthy-eval suite. Lower is better.

SplitSamplesWERCER
dev_clean79012.47%3.57%
dev_real1,20015.42%5.23%
test_frozen1,57016.01%5.17%
balanced mean3,56014.63%4.66%

Internal validation metrics from the training run were slightly different because checkpoint selection and later healthy-eval rescoring used different validation layouts. For the public release, the healthy-eval table above is the headline result.

Intended Use

This model is intended for:

  • —Pashto speech transcription
  • —subtitle generation
  • —search/indexing over Pashto speech archives
  • —research and downstream ASR benchmarking for Pashto

Limitations

  • —Whisper is natively a short-form model. Multi-minute audio should be transcribed with chunking rather than naive manual window concatenation.
  • —Performance will vary across dialects, microphones, room acoustics, and code-switched speech.
  • —Noisy far-field audio and overlapped speech remain hard cases.
  • —Like other Whisper models, decoding can degrade on long audio if the inference pipeline handles chunk merging poorly.

Quick Start

The standard Hugging Face transformers pipeline is the simplest way to use this model:

python
import torch
from transformers import pipeline

model_id = "Nasimbahar/pashto-ghag-whisper-medium-asr"
device = 0 if torch.cuda.is_available() else -1
dtype = torch.float16 if torch.cuda.is_available() else torch.float32

pipe = pipeline(
    "automatic-speech-recognition",
    model=model_id,
    torch_dtype=dtype,
    device=device,
    chunk_length_s=30,
)

result = pipe(
    "sample.wav",
    generate_kwargs={"language": "pashto", "task": "transcribe"},
)

print(result["text"])

Detailed usage notes, long-audio guidance, and lower-level examples are in USAGE.md.

Recommended Inference Practice

  • —For short clips: direct inference is fine.
  • —For long recordings: use transformers.pipeline(..., chunk_length_s=30) or another timestamp-aware chunking method.
  • —If you are building a web app, treat long audio as an asynchronous job instead of a blocking request.

Acknowledgements

Citation

If you use this model in a paper, report, or product benchmark, cite:

  • —OpenAI Whisper
  • —Mozilla Common Voice
  • —this Hugging Face model repo