datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.medimind-r11-train
MediMind R11 — ASR training data
Unified manifest + packed audio for fine-tuning Whisper-large-v3 on Norwegian
clinical and conversational speech.
Training manifest: r11_manifest.jsonl — 11,022 packs
Held-out eval set: r11_heldout_eval.jsonl — 291 packs (NEVER train on these)
~see manifest audit packs total
11 sources: lege_*, podcasts (motiv/podk/stet), nb_samtale, nb_tale_m3, tts_drugs
Schema
See r11_manifest.jsonl (one JSON object per line) and… See the full description on the dataset page: https://huggingface.co/datasets/gallip0li/medimind-r11-train.
