shujaAK/whisper-medium-hindi-hinglish-asr-fine-tuned
Whisper Medium Hindi-Hinglish ASR (Fine-Tuned)
A fine-tuned version of OpenAI Whisper Medium for Hindi and Hinglish Automatic Speech Recognition (ASR).
This model was developed to improve Whisper's performance on conversational Hindi and code-mixed Hindi-English speech using a carefully curated synthetic speech corpus.
Model Details
Motivation
OpenAI Whisper provides excellent multilingual speech recognition, but conversational Hindi and Hinglish often remain challenging due to code-mixing, pronunciation variations, and domain-specific vocabulary.
This project fine-tunes Whisper Medium using a curated synthetic dataset to significantly improve recognition quality while preserving Whisper's multilingual capabilities.
Dataset
The training dataset was generated using a multi-stage synthetic data generation pipeline consisting of:
- LLM-generated conversational Hindi/Hinglish text
- High-quality speech synthesis
- Transcript refinement and normalization
- Quality filtering
- Curated speech-text pairs for Whisper fine-tuning
The focus was on maximizing transcript quality while maintaining linguistic diversity.
Evaluation
Evaluation was performed on an unseen held-out test set.
Overall Performance
Performance by Difficulty
Performance by Language
Performance by Error Category
Usage
Load the model using the Hugging Face Transformers library.
import torch from transformers import WhisperProcessor, WhisperForConditionalGeneration
processor = WhisperProcessor.from_pretrained( "shujaAK/whisper-medium-hindi-hinglish-asr-fine-tuned" )
model = WhisperForConditionalGeneration.from_pretrained( "shujaAK/whisper-medium-hindi-hinglish-asr-fine-tuned" )
device = "cuda" if torch.cuda.is_available() else "cpu"
model = model.to(device) model.eval()
Intended Uses
This model is intended for:
- Hindi Automatic Speech Recognition
- Hinglish Automatic Speech Recognition
- Academic Research
- Speech Recognition Experiments
- Fine-tuning Research
Limitations
- The model is primarily trained on synthetic speech.
- Performance may decrease on noisy recordings.
- Performance may vary across unseen accents and speaking styles.
- Not evaluated for streaming ASR.
Training Framework
- Hugging Face Transformers
- PyTorch
Results Summary
Acknowledgements
This project builds upon:
- OpenAI Whisper
- Hugging Face Transformers
Citation
If you use this model in your research or applications, please cite this Hugging Face repository.
@misc{whisper_medium_hindi_hinglish_asr,
title={Whisper Medium Hindi-Hinglish ASR (Fine-Tuned)},
author={Suja Akhter},
year={2026},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/shujaAK/whisper-medium-hindi-hinglish-asr-fine-tuned}}
}