transitionGap/ASR_Marathi_Sentences
Marathi Sentence-Level ASR Dataset 📌 Overview This dataset contains sentence-level Marathi speech segments aligned with transcripts. The dataset was created by extracting subtitle timestamps (SRV3 format) from Marathi YouTube content and segmenting the corresponding audio using precise time alignment. Each sample contains: A WAV audio file (sentence-level) The corresponding Marathi transcript text This dataset is suitable for: Whisper fine-tuning Wav2Vec2 CTC… See the full description on the dataset page: https://huggingface.co/datasets/transitionGap/ASR_Marathi_Sentences.
038
