datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
commonvoice22_sidon
CV22-Sidon
Overview
This dataset hosts a release of Mozilla Common Voice 22 restored with the Sidon speech restoration model.
Source: Mozilla Common Voice 22.0
Processing: Sidon denoising (sarulab-speech/sidon-v0.1) with 21 s chunks and 48 kHz reconstruction
Format: WebDataset shards (.tar.gz)
Manifest: paths.yaml enumerates every shard path for Hugging Face–style loading
License: Original Common Voice license (CC0 1.0)
Languages
137 language folders are… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon.commonvoice22-sidon-dacvae
CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents
Source
sarulab-speech/commonvoice22_sidon
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.common-voice-subset-for-clapcommon_voice_21_ru
Dataset Description
Набор данных validated.tsv отфильтрованный по down_votes = 0
📊 Статистика датасета
Информация по сплитам
🔹 Тренировочный набор (train)
Метрика
Значение
Количество семплов
93,531
Общая продолжительность
132.25 часов (476,089.70 секунд)
Средняя продолжительность семпла
5.09 секунд
🔹 Валидационный набор (validate)
Метрика
Значение
Количество семплов
38,836
Общая продолжительность
55.21… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/common_voice_21_ru.common_voice_21_0_yuecantonese only
common_voice_englishcommon_voice_16_1_fr_smallmozilla-common-voice-spontaneous-speech-asr-shared-task
Mozilla Common Voice Spontaneous Speech ASR Shared Task
This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task
train/dev and test archives in one place.
Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk,
cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc,
rwm, sco, tob, top, ttj, ukv, ush.
Split package
Mozilla Data Collective dataset ID
Hub archive
Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.common_voice_16_1common-voice-mn-24
🇲🇳 Common Voice Mongolian 24.0 Dataset
This repository hosts the latest release (v24.0) of the Mozilla Common Voice Scripted Speech dataset for Mongolian (mn). This dataset is a vital resource for training robust Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems for the Mongolian language.
📊 Dataset Statistics
Metric
Value
Total Clips
96,308
Total Duration
140.56 Hours
Validated Duration
49.19 Hours
Total Speakers
606
Format… See the full description on the dataset page: https://huggingface.co/datasets/onlysainaa/common-voice-mn-24.common-voice-scripted-speech-24.0-mongoliancommonvoice22_sidon_be_rawcommon_voice_14UrduSpeech-CommonVoice22-SIDONCommon_Voice_Delta_Segment_11.0Common_Voice_Corpus_21.0common_voice_sq_20_localCapSpeech-CommonVoice
CapSpeech-CommonVoice Audio
DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
Overview
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech-CommonVoice.
