CoolFace
Modelpublic

Sparkplugx1904/whisper-tiny-id

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes7downloads
Model Card

Whisper Tiny Model – Indonesian ASR

Model Description

This model is a fine-tuned version of openai/whisper-tiny for Automatic Speech Recognition (ASR) in Indonesian (id). It supports transcription of Indonesian speech into text across various audio conditions, with performance and resource usage depending on the selected model size.

Intended Use

  • Indonesian speech-to-text transcription
  • Research and experimentation
  • Educational and academic purposes
  • Application development and benchmarking

Model variants (tiny, base, small, medium, large) differ in accuracy, speed, and hardware requirements. Users should select the size that best matches their constraints and objectives.

Limitations

  • Transcription quality depends on audio clarity, speaker accent, and background noise
  • Smaller variants may produce higher error rates on long or complex audio
  • Larger variants require significantly more compute and memory
  • Outputs should be reviewed before use in critical or high-risk applications

Training Data

This model was fine-tuned using Mozilla Common Voice v23.0 (Indonesian). Common Voice is a publicly available, community-driven speech dataset released by Mozilla under a permissive license. Dataset characteristics such as speaker diversity, recording quality, and utterance length may influence model behavior.

Evaluation

The model is typically evaluated using Word Error Rate (WER). Evaluation results may vary depending on dataset, domain, audio conditions, and model size.

Training results

StepTraining Loss
1001.282900
2000.682300
3000.568900
4000.487500
5000.372700
6000.375500
7000.276200
8000.226000
9000.223800
10000.188600
11000.164300
12000.151400
13000.130000
14000.133900
15000.119700
15500.117300