datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.f1-team-radio
F1 Team Radio Dataset
A comprehensive dataset of Formula 1 team radio communications with transcriptions.
Dataset Description
This dataset contains team radio audio clips from Formula 1 races along with their text transcriptions. Team radio communications are the real-time messages exchanged between F1 drivers and their pit wall engineers during race weekends.
Dataset Statistics
Metric
Value
Total audio clips
14,681
Grand Prix events
149
Unique… See the full description on the dataset page: https://huggingface.co/datasets/fluffypotatoes/f1-team-radio.fluers-mn
fleurs-mn
Mongolian speech recognition dataset , recombined and split into a 90% train and 10% test set.
Dataset Statistics
Total samples: 4,428Total duration: 15h 33m 10s (15.55 h)
Per-split breakdown
Split
Samples
Total Duration
Avg Duration
train
3,985
13h 56m 57s (13.95 h)
12.60 s
test
443
1h 36m 13s (1.60 h)
13.03 s
fluid-2-sft-asr
Fluid 2 — synthetic dictation cleanup
Fluid 2 is an English supervised-fine-tuning corpus for models that turn noisy automatic-speech-recognition output into the written insertion a user intended. It contains 354,549 rows in official document-grouped 96/2/2 splits, 861.3 hours of processed 16 kHz speech, and 8.48M target-side loss tokens in 355 Parquet shards (49.25 GiB).
This is not an ordinary transcription dataset. The model sees document context plus an ASR hypothesis and… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/fluid-2-sft-asr.
