Harshkmr/omniscribe_corpus
OmniScribe Corpus A multilingual speech transcription corpus designed for fine-tuning ASR models on Indian medical and general-domain speech. It covers Hindi, Marathi, and Indian English, with a focus on clinical and healthcare contexts. Overview Split Rows (after oversampling) Approx. Duration train ~30750 ~230 hrs benchmark ~4,089 ~25 hrs Audio samples average 20–30 seconds each. All samples are at least 5 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Harshkmr/omniscribe_corpus.
07
Add dataset README
Add dataset README
Upload dataset
initial commit
