datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-voice-17-regmix-webdataset
Common Voice 17 RegMix WebDataset
Public, training-oriented WebDataset conversion of fsicoli/common_voice_17_0, pinned to source revision 8262c16bf297c87a9cd88c51997c4758ed7a8ba2.
Layout
Each language is a RegMix cluster at data/<language>/*.tar. Every sample is a pair with the same key:
<key>.opus: mono Ogg Opus audio at 24 kbps, preserving the source sample rate (normally 48 kHz)
<key>.json: UTF-8 training metadata and the complete original TSV row
The source… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/common-voice-17-regmix-webdataset.fleurs-regmix-webdataset
FLEURS RegMix WebDataset
Public, training-oriented WebDataset conversion of google/fleurs, pinned to source revision 70bb2e84b976b7e960aa89f1c648e09c59f894dd.
Layout
Each language is a RegMix cluster at data/<language>/*.tar. Shard names keep the source split, data/<language>/<language>-<split>-<index>.tar, so a training
mixture can be assembled without pulling the FLEURS evaluation splits into it. Every sample is a pair with the same key:
<key>.opus: mono Ogg… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/fleurs-regmix-webdataset.asrtts_packed_webdataset
ASR+TTS Repacked Data (3.56M samples, mp3)
This dataset is a WebDataset repack prepared for FlexiSLM training (ASR+TTS tasks).
Paper: https://arxiv.org/abs/2606.31247
Demo page: https://flexislm.github.io/
Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset
for training FlexiSLM, a spoken language model.
This repository contains the paired prompt-and-response audio portion of the release in… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/asrtts_packed_webdataset.
