datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.xtreme_sXTREME-S covers four task families: speech recognition, classification, speech-to-text translation and retrieval. Covering 102
languages from 10+ language families, 3 different domains and 4
task families, XTREME-S aims to simplify multilingual speech
representation evaluation, as well as catalyze research in “universal” speech representation learning.audiobooks-xxlcommonvoice-12.0-arabic-voice-converted
Dataset Card for Voice Converted Arabic Common Voice 12.0
This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.Redmond-Sentence-Recall
Dataset Summary
The Redmond Sentence Recall (RSR) measures a child’s ability to repeat sentences that contain regular past tense forms and past participle forms (e.g., “He kicked” vs. “He was kicked”). This task helps identify language impairments, with each child repeating 16 sentences heard through headphones. The dataset includes anonymized audio recordings of these repetitions.
What makes the RSR dataset uniquely valuable is its focus on sentence recall using both regular past… See the full description on the dataset page: https://huggingface.co/datasets/xlab-ub/Redmond-Sentence-Recall.DeepDialogue-xtts
DeepDialogue-xtts
DeepDialogue-xtts is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions.
This repository contains the XTTS-v2 variant of the dataset, where speech is generated using XTTS-v2 with explicit emotional conditioning.
🚨 Important
This dataset is large (~180GB) due to the inclusion of high-quality audio files. When cloning the… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-xtts.ePark_ju_xing_pian_gao_zhong_sentence_patterns_senior_high
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_ju_xing_pian_gao_zhong_sentence_patterns_senior_high
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_ju_xing_pian_gao_zhong_sentence_patterns_senior_high.Hausa-Synthetic-ASR-Dataset-XTTSSynthetic Hausa ASR dataset generated using a fine-tuned version of the XTTS-v2 model.
Sample rate: 24kHz.
Total duration: 574 hours.
xbmu_amdo31
Dataset Card for [XBMU-AMDO31]
Dataset Summary
XBMU-AMDO31 dataset is a speech recognition corpus of Amdo Tibetan dialect. The open source corpus contains 31 hours of speech data and resources related to build speech recognition systems, including transcribed texts and a Tibetan pronunciation dictionary.
Supported Tasks and Leaderboards
automatic-speech-recognition: The dataset can be used to train a model for Amdo Tibetan Automatic Speech Recognition (ASR). It… See the full description on the dataset page: https://huggingface.co/datasets/syzym/xbmu_amdo31.wjbmattingly_xhosa_merged_audio
Xhosa Merged Audio
This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below).
The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/ilyes25/wjbmattingly_xhosa_merged_audio.english-x-code-switching
Synthetic English Code-Switching Evaluation Set
This dataset contains synthetic long-form English code-switching audio samples built from ML-SUPERB hybrid data.
Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.
Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching.english-en-x-code-switching-main-lang
English EN-X Code-Switching Main-Language
This dataset contains synthetic English-plus-one-language code-switching samples built from FLEURS.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang.XIMA-2122
DOVEXAI Nigerian Languages Speech & Translation Dataset (XIMA-2122)
Curated by DOVEXAI LTD (Nigeria) · dovexai.africa · dovexai.io
A provenance-tracked dataset of Nigerian-language voice recordings paired with
human transcriptions and translations, built for evaluation and supervised fine-tuning (SFT) of
speech and translation models. Every recording is cryptographically receipted,
fully anonymized, and quality-gated before release.
Dataset snapshot
Feature… See the full description on the dataset page: https://huggingface.co/datasets/DOVEXAI/XIMA-2122.ePark_yue_du_shu_xie_pian_reading_writing
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_yue_du_shu_xie_pian_reading_writing
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_yue_du_shu_xie_pian_reading_writing.librivox-tracksA dataset of all audio files uploaded to LibriVox before 26th September 2023.
Forked from https://huggingface.co/datasets/pykeio/librivox-tracks
Changes:
Used archive.org metadata API to annotate rows with "duration" column
pinga-fogo-chico-xavier
🎙️ Pinga-Fogo com Chico Xavier — TV Tupi, 1971
As duas entrevistas históricas do médium Chico Xavier, transmitidas ao vivo pela
TV Tupi em 1971, transcritas e estruturadas em turnos de fala com timestamp.
345 turnos (115 deles respostas do próprio Chico Xavier), a partir de
6 horas de áudio — o registro mais extenso do médium falando de improviso,
sem edição, diante de um painel de jornalistas.
Arquivos
Arquivo
Programa
Turnos
Respostas do Chico… See the full description on the dataset page: https://huggingface.co/datasets/ia-espirita/pinga-fogo-chico-xavier.fleurs_xho
FLEURS -- isiXhosa (xh_za)
Re-mirrored from google/fleurs, config
xh_za. n-way parallel read speech built on FLORES-101 text -- an
evaluation-sized corpus (~19 h), not training scale, but the
de facto African-language ASR/TTS benchmark (used in Whisper, MMS, SeamlessM4T,
USM papers).
Licence
CC BY 4.0 -- inherited unchanged from the source.
What changed from the source
audio peak-normalized per clip (see below); sample rate and encoding otherwise… See the full description on the dataset page: https://huggingface.co/datasets/simpra/fleurs_xho.xhosa_merged_audio
Xhosa Merged Audio
This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below).
The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/wjbmattingly/xhosa_merged_audio.malagasy-nwt-bible
Malagasy NWT Dataset
Dataset from the Malagasy New World Translation Bible (JW.org 2021).
Format
audio — 16 kHz mono WAV
text — clean Malagasy transcript, digits converted to Malagasy words
Speaker Information
This dataset contains multiple speakers — several readers who each narrate
different books or chapters of the Bible. Automatic speaker diarization and
clustering was attempted using
pyannote/speaker-diarization-3.1,
but reliable speaker identity… See the full description on the dataset page: https://huggingface.co/datasets/XedriX/malagasy-nwt-bible.X-Voice-TestsetX-Voice Multilingual Test Set
High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages
This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model.
Dataset Summary
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.xtreme_sXTREME-S covers four task families: speech recognition, classification, speech-to-text translation and retrieval. Covering 102
languages from 10+ language families, 3 different domains and 4
task families, XTREME-S aims to simplify multilingual speech
representation evaluation, as well as catalyze research in “universal” speech representation learning.xhosa_merged_audio
Xhosa Merged Audio
This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below).
The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/xhosa_merged_audio.mozilla-common-voice-20-eu
Mozilla Common Voice Basque Dataset v20.0
This is the Basque portion of the Mozilla Common Voice dataset version 20.0.
ScreenTalk-XS
🎬 ScreenTalk-XS: Sample Speech Dataset from Screen Content 🖥️
📢 What is ScreenTalk-XS?
ScreenTalk-XS is a high-quality transcribed speech dataset containing 10k speech samples from diverse screen content.It is designed for automatic speech recognition (ASR), natural language processing (NLP), and conversational AI research.
✅ This dataset is freely available for research and educational use.🔹 If you need a larger dataset with more diverse speech samples… See the full description on the dataset page: https://huggingface.co/datasets/Itbanque/ScreenTalk-XS.english-x-code-switching-samples
Synthetic English Code-Switching Evaluation Set Samples
This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset.
Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.
Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.nchlt_speech_xho
NCHLT Speech Corpus -- isiXhosa
This is the isiXhosa language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): xho
URI: https://hdl.handle.net/20.500.12185/279
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_xho.omnievalkit-data-test
OmniEvalKit Evaluation Datasets
Evaluation datasets for OmniEvalKit,
a comprehensive evaluation framework for omni-modal (audio + video + image + text) models.
Overview
Total subsets: 89
Total samples: 353,610
Total size: 352.3 GB (Parquet with embedded audio/image, no video)
Subsets requiring video download: 42
Note: Video files are NOT embedded in the Parquet files due to size constraints.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/xiaofff/omnievalkit-data-test.ePark_ju_xing_pian_guo_zhong_sentence_patterns_junior_high
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_ju_xing_pian_guo_zhong_sentence_patterns_junior_high
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_ju_xing_pian_guo_zhong_sentence_patterns_junior_high.english-en-x-code-switching-main-lang-samples-merged
English EN-X Code-Switching Main-Language Merged Samples
This dataset contains contiguous same-language segments from the paired mixed dataset.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples-merged.ePark_xue_xi_ci_biao_learning_vocabulary
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary.
