CoolFace
Datasetpublic

Harshkmr/omniscribe_corpus

OmniScribe Corpus A multilingual speech transcription corpus designed for fine-tuning ASR models on Indian medical and general-domain speech. It covers Hindi, Marathi, and Indian English, with a focus on clinical and healthcare contexts. Overview Split Rows (after oversampling) Approx. Duration train ~30750 ~230 hrs benchmark ~4,089 ~25 hrs Audio samples average 20–30 seconds each. All samples are at least 5 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Harshkmr/omniscribe_corpus.

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes7downloads
4 commits on main
3feba3b5mo ago

Add dataset README

Harshkmr
f1657385mo ago

Add dataset README

Harshkmr
2b100d95mo ago

Upload dataset

Harshkmr
89e1e845mo ago

initial commit

Harshkmr