datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EchoX-Dialogues-Plus
EchoX-Dialogues-Plus: Training Data Plus for EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
🐈⬛ Github | 📃 Paper | 🚀 Space
🧠 EchoX-8B | 🧠 EchoX-3B | 📦 EchoX-Dialogues (base)
EchoX-Dialogues-Plus
EchoX-Dialogues-Plus extends KurtDu/EchoX-Dialogues with large-scale Speech-to-Speech (S2S) and Speech-to-Text (S2T) dialogues.
All assistant/output speech is synthetic (single, consistent timbre for S2S). Texts are from… See the full description on the dataset page: https://huggingface.co/datasets/KurtDu/EchoX-Dialogues-Plus.EchoX-Dialougues
EchoX-Dialogues: Training Data for EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
🐈⬛ Github | 📃 Paper | 🚀 Space
🧠 EchoX-8B | 🧠 EchoX-3B | 📦 EchoX-Dialogues-Plus
EchoX-Dialogues provides the primary speech dialogue data used to train EchoX, restricted to S2T (speech → text) in this repository.
All input speech is synthetic; text is derived from public sources with multi-stage cleaning and rewriting. Most turns include asr /… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/EchoX-Dialougues.Echoes-Platos-CaveEchoes in Plato's Cave — Controlled Speech–Text Corpus
Controlled corpus of 14,400 synthetic English utterances in which the same 600 sentences are
rendered by 6 speakers × 4 emotions, so that speaker identity and prosody vary
while linguistic content is held fixed. It was built for the paper: Echoes in Plato's Cave: Measuring Global and Local Alignment Between Speech and Language Representations, accepted as an oral presentation at the Speech and Audio Language… See the full description on the dataset page: https://huggingface.co/datasets/alefiury/Echoes-Platos-Cave.common_voice_11_0
Dataset Card for Common Voice Corpus 11.0
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 24210 recorded hours in the dataset also include demographic metadata like age, sex, and accent
that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 16413 validated hours in 100 languages, but more voices and languages are always added.
Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/echodict/common_voice_11_0.EchoLens
EchoLens
A Human-Speech Dataset for Auditing Demographic Sensitivity in Audio-Language Models
📄 Paper (EMNLP 2026 Findings) ·
💻 Code
Voice interfaces are increasingly moving away from transcription pipelines toward end-to-end systems that directly respond to audio inputs. This development in turn requires a shift in evaluation methodology away from transcription accuracy and towards more substantive markers such as response validity. We introduce EchoLens, a demographically… See the full description on the dataset page: https://huggingface.co/datasets/alexsdl/EchoLens.echo
Echo
Echo is a speech dataset for Romanian language crowd-sourced from the community.
The dataset contains over 300 hours of speech data from 300 speakers and is
available for non-commercial research purposes only. The dataset is collected
using the Echo platform.
Zambezi_ECHO_v1
Zambezi ECHO v1: Shona-English Code-Switched Maternal Health Queries
Dataset Description
Zambezi ECHO v1 SESB (Shona-English Speech Benchmark) is a dataset of short,
simulated patient voice queries in Shona (Zimbabwe), code-switched with
English, covering common maternal and child health concerns — pregnancy
symptoms, danger signs, child illness, and general health questions asked
the way patients actually phrase them in the field, mixing Shona with
English… See the full description on the dataset page: https://huggingface.co/datasets/dawahealth/Zambezi_ECHO_v1.Zambezi_ECHO_v1
Zambezi ECHO v1: Shona-English Code-Switched Maternal Health Queries
Dataset Description
Zambezi ECHO v1 SESB (Shona-English Speech Benchmark) is a dataset of short,
simulated patient voice queries in Shona (Zimbabwe), code-switched with
English, covering common maternal and child health concerns — pregnancy
symptoms, danger signs, child illness, and general health questions asked
the way patients actually phrase them in the field, mixing Shona with
English… See the full description on the dataset page: https://huggingface.co/datasets/tarirozw/Zambezi_ECHO_v1.
