datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.fluid-2-sft-asr
Fluid 2 — synthetic dictation cleanup
Fluid 2 is an English supervised-fine-tuning corpus for models that turn noisy automatic-speech-recognition output into the written insertion a user intended. It contains 354,549 rows in official document-grouped 96/2/2 splits, 861.3 hours of processed 16 kHz speech, and 8.48M target-side loss tokens in 355 Parquet shards (49.25 GiB).
This is not an ordinary transcription dataset. The model sees document context plus an ASR hypothesis and… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/fluid-2-sft-asr.
