datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genshin-voice
Genshin Voice
Genshin Voice is a dataset of voice lines from the popular game Genshin Impact.
Hugging Face 🤗 Genshin-Voice
ModelScope Genshin-Voice
Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index.
Last update at 2026-08-13
654252 wavs
7291 without speaker (1%)
52693 without transcription (8%)
1088 without inGameFilename (0%)
Dataset Details
Dataset Description
The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.zenless-voice
Zenless Voice
Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero.
Hugging Face 🤗 Zenless-Voice
ModelScope Zenless-Voice
Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index.
Last update at 2026-09-17, game version 3.2.0
406720 wavs
78785 without speaker (19%)
123429 without transcription (30%)
83509 without inGameFilename (21%)
Speaker archives contain 327,935 WAVs in 4,322 ZIPs. The 78,785 rows… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/zenless-voice.starrail-voice
StarRail Voice
StarRail Voice is a dataset of voice lines from the popular game Honkai: Star Rail.
Hugging Face 🤗 StarRail-Voice
ModelScope StarRail-Voice
Last update at 2026-07-16, game version 4.4.0
403437 wavs
60164 without speaker (15%)
61375 without transcription (15%)
57869 without inGameFilename (14%)
Dataset Details
Dataset Description
The dataset contains voice lines from the game's characters in multiple languages, including Chinese… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/starrail-voice.SimbaBench_dataset
SibmaBench Data Release & Benchmarking
To evaluate your model on SimbaBench across all supported tasks (ASR, TTS, and SLID), simply load the corresponding configuration for the task and language you wish to benchmark.
Each task is organized by configuration name (e.g., asr_test_afr, tts_test_wol, slid_61_test). Loading a configuration provides the standardized evaluation split for that specific benchmark.Example:
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/SimbaBench_dataset.simchoir-parquet
FastMSS synthetic multi-speaker meetings - parquet edition
Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring.
Subsets and splits
debug — splits: train — 1 mixtures, 1.6 min total, 6 unique speakers… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/simchoir-parquet.fleurs_xho
FLEURS -- isiXhosa (xh_za)
Re-mirrored from google/fleurs, config
xh_za. n-way parallel read speech built on FLORES-101 text -- an
evaluation-sized corpus (~19 h), not training scale, but the
de facto African-language ASR/TTS benchmark (used in Whisper, MMS, SeamlessM4T,
USM papers).
Licence
CC BY 4.0 -- inherited unchanged from the source.
What changed from the source
audio peak-normalized per clip (see below); sample rate and encoding otherwise… See the full description on the dataset page: https://huggingface.co/datasets/simpra/fleurs_xho.simple-escwa
🗣️ Simple-ESCWA: A simpler version of ESCWA-CS Corpus
The ESCWA-CS Corpus was collected over two days of meetings of the United Nations Economic and Social Commission for Western Asia (ESCWA) held in 2019.It contains intra-sentential code-switching between Arabic and English, with some speakers—particularly from Algeria, Tunisia, and Morocco—alternating between Arabic and French.
The dataset spans approximately 2.8 hours of speech, featuring dialectal Arabic and a Code Mixing Index… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/simple-escwa.simulated_rirs_dataset
Simulated Rirs Dataset
Dataset Description
This dataset contains 400 samples organized across multiple splits and 4 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
original: 100 samples
train: 100 samples
largeroom: 100 samples
train: 100 samples
mediumroom: 100 samples
train: 100 samples
smallroom: 100 samples
train: 100 samples
Usage
Load specific subset and… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/simulated_rirs_dataset.dia-simsamu-test-fr
Simsamu — French diarization (mirror)
Mirror byte-exact du dataset diarizers-community/simsamu (HF Inria / Eole),
réutilisé tel quel comme dataset de bench DIA français pour STTSTAGE.
Contenu
61 enregistrements audio
3h 15 min au total (~3 min 11 s par enregistrement)
16 kHz mono
2 locuteurs par enregistrement (médecin régulateur + appelant simulé)
Langue : français (fr), domaine : régulation médicale / téléphonique simulée
Licence : MIT (héritée upstream)… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-simsamu-test-fr.speech-simulated-medical-exams
Speech Simulated Medical Exams
Simulated patient-physician medical exam conversations with rich speech metadata annotations. Built for training single-step ASR models that transcribe and annotate multiple concepts simultaneously, including speaker changes, emotions, intents, and roles.
Dataset Details
Property
Value
Examples
25,706
Language
English
Audio
16 kHz WAV
Source
Simulated medical interviews (respiratory focus)
Features… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/speech-simulated-medical-exams.room_sim_reverb_dataset
Room Sim Reverb Dataset
Dataset Description
This dataset contains 100 samples organized across multiple splits and 1 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
room_simulator: 100 samples
train: 100 samples
Usage
Load specific subset and split:
from datasets import load_dataset
# Load specific subset and split
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/room_sim_reverb_dataset.
