datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-audio-reserve-r1
Nepali long-form speech, restored
~548.1 h of Nepali speech at 24 kHz. One speaker per clip, average 32 s, up to several minutes.
A backup, not a release: the transcripts are machine-generated, and none of this
audio passed the quality gate that produced our training corpus.
Derived from AI4Bharat IndicVoices-R (CC-BY-4.0) and restored with sidon-v0.1. Attribution is required by that licence, so it is given here. These are the long-form files our quality gate rejected; the 74 h… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r1.medimind-r11-train
MediMind R11 — ASR training data
Unified manifest + packed audio for fine-tuning Whisper-large-v3 on Norwegian
clinical and conversational speech.
Training manifest: r11_manifest.jsonl — 11,022 packs
Held-out eval set: r11_heldout_eval.jsonl — 291 packs (NEVER train on these)
~see manifest audit packs total
11 sources: lege_*, podcasts (motiv/podk/stet), nb_samtale, nb_tale_m3, tts_drugs
Schema
See r11_manifest.jsonl (one JSON object per line) and… See the full description on the dataset page: https://huggingface.co/datasets/gallip0li/medimind-r11-train.
