datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genshin-voice
Genshin Voice
Genshin Voice is a dataset of voice lines from the popular game Genshin Impact.
Hugging Face 🤗 Genshin-Voice
ModelScope Genshin-Voice
Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index.
Last update at 2026-08-13
654252 wavs
7291 without speaker (1%)
52693 without transcription (8%)
1088 without inGameFilename (0%)
Dataset Details
Dataset Description
The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.zenless-voice
Zenless Voice
Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero.
Hugging Face 🤗 Zenless-Voice
ModelScope Zenless-Voice
Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index.
Last update at 2026-09-17, game version 3.2.0
406720 wavs
78785 without speaker (19%)
123429 without transcription (30%)
83509 without inGameFilename (21%)
Speaker archives contain 327,935 WAVs in 4,322 ZIPs. The 78,785 rows… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/zenless-voice.starrail-voice
StarRail Voice
StarRail Voice is a dataset of voice lines from the popular game Honkai: Star Rail.
Hugging Face 🤗 StarRail-Voice
ModelScope StarRail-Voice
Last update at 2026-07-16, game version 4.4.0
403437 wavs
60164 without speaker (15%)
61375 without transcription (15%)
57869 without inGameFilename (14%)
Dataset Details
Dataset Description
The dataset contains voice lines from the game's characters in multiple languages, including Chinese… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/starrail-voice.SimbaBench_dataset
SibmaBench Data Release & Benchmarking
To evaluate your model on SimbaBench across all supported tasks (ASR, TTS, and SLID), simply load the corresponding configuration for the task and language you wish to benchmark.
Each task is organized by configuration name (e.g., asr_test_afr, tts_test_wol, slid_61_test). Loading a configuration provides the standardized evaluation split for that specific benchmark.Example:
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/SimbaBench_dataset.simsamu
Simsamu dataset
This repository contains recordings of simulated medical dispatch dialogs in the
french language, annotated for diarization and transcription. It is published
under the MIT license.
These dialogs were recorded as part of the training of emergency medicine
interns, which consisted in simulating a medical dispatch call where the interns
took turns playing the caller and the regulating doctor.
Each situation was decided randomly in advance, blind to who was playing the… See the full description on the dataset page: https://huggingface.co/datasets/medkit/simsamu.xh-tts-samples
isiXhosa TTS — reference audio and training samples
Two very different kinds of audio live here. Check the folder before judging
anything.
folder
what it is
source
speakers/
REAL HUMAN speech — 20 s excerpts of ViXSD readers, for choosing a voice
ViXSD recordings
samples_vixsd/
MODEL OUTPUT — what the VITS model generates at a given training step
generated
speakers/ — ground truth
male_xho_reader_008_20s.wav etc. Excerpts taken from the middle of… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-samples.SIMAXxh-tts-vixsd
isiXhosa TTS clips (ViXSD, segmented)
3,861 clips, 22,050 Hz mono, 8 speakers, cut from
long-form recordings by CTC forced alignment.
Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI /
Way With Words, under the Esethu License — see
https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD
Pipeline
vixsd_extract.py — parquet to mono 22,050 Hz. Source is heterogeneous:
rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-vixsd.musifyer_datasetsimchoir-parquet
FastMSS synthetic multi-speaker meetings - parquet edition
Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring.
Subsets and splits
debug — splits: train — 1 mixtures, 1.6 min total, 6 unique speakers… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/simchoir-parquet.simplified_google_speech_commands_wav2vec2_960hcommonvoice_13_0_pt_48kHz_simplificado_augmented_white_noise
Dataset Card for "commonvoice_13_0_pt_48kHz_simplificado_augmented_white_noise"
More Information needed
xh-tts-vixsd-norm
isiXhosa TTS clips (ViXSD, segmented)
3,861 clips, 22,050 Hz mono, 8 speakers, cut from
long-form recordings by CTC forced alignment.
Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI /
Way With Words, under the Esethu License — see
https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD
Pipeline
vixsd_extract.py — parquet to mono 22,050 Hz. Source is heterogeneous:
rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-vixsd-norm.xh-tts-slr32
isiXhosa TTS — SLR32 prepared for VITS
Multi-speaker isiXhosa speech, resampled and text-normalised for VITS training.
Each audio file is paired with its transcript in metadata.csv.
Attribution (required by the licence)
Derived from OpenSLR SLR32, "High quality TTS data for four South African
languages (af, st, tn, xh)", created by North West University and Google
(2017), released under CC BY-SA 4.0.
Source: https://openslr.org/32/
This derivative is likewise CC… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-slr32.MC_proton_simulation_DoseRAD2026
Water-Phantom Proton Monte Carlo (DoseRAD2026)
Full 3D Monte-Carlo energy-deposit distributions for 85 proton energies in a water
phantom, 10⁹ primaries each. This is the reference data behind the machine look-up table
of team DoseHappens's DoseRAD2026 Grand Challenge
entry (submitted algorithm codename MAALGO): the analytic pencil-beam engine's depth-dose and lateral-spread curves are fitted to
these volumes, and then refined by backpropagating the dose engine against them.… See the full description on the dataset page: https://huggingface.co/datasets/zimmeryWo/MC_proton_simulation_DoseRAD2026.xh-tts-check
ViXSD segmentation check
A stratified sample of 40 clips from 3,861 produced by vixsd_segment.py.
Listen to each and confirm the audio says exactly the text. A misaligned pair teaches the model a wrong sound-to-letter mapping, and it is invisible to every automated check.
worst rows are the lowest-scoring clips in the whole run — if those are right, the rest almost certainly are.
clip
why
sec
score
ends
transcription
cds_xho_079_1_0025.wav
worst
1.7
0.41
sentence… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-check.xh-tts-slr32-norm
isiXhosa TTS — SLR32 prepared for VITS
Multi-speaker isiXhosa speech, resampled and text-normalised for VITS training.
Each audio file is paired with its transcript in metadata.csv.
Attribution (required by the licence)
Derived from OpenSLR SLR32, "High quality TTS data for four South African
languages (af, st, tn, xh)", created by North West University and Google
(2017), released under CC BY-SA 4.0.
Source: https://openslr.org/32/
This derivative is likewise CC… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-slr32-norm.turnbench-dev-no-backchannel
TurnBench Dev - Backchannels Removed
A derivative of mundo-ai/turn-benchmark-dev
with every majority-annotated backchannel removed from the audio: 1853 backchannels
across 38 conversations, 2077.0 seconds in total, cut out of the
speaker's own channel and replaced by background noise taken from elsewhere in that same channel.
Everything else is the original recording, sample for sample. Same conversations, same duration,
same timeline, same annotator tracks, same speech -- only… See the full description on the dataset page: https://huggingface.co/datasets/JSALT2026-Conv-AI-Simulator/turnbench-dev-no-backchannel.keep-it-simple-multimodal
keep-it-simple-multimodal
A mini, standalone multimodal dataset: image+caption, audio+caption, video+caption, lidar, IMU, and
optimal-control state/action pairs. Companion to keep-it-simple
(text), built to feed KairosPretrainingDataset in kairos.
Structure
One generic schema for every row — no per-modality columns, no fixed shape/dtype assumptions:
Column
Type
Description
modality
string
image_caption | audio_caption | video_caption | lidar | imu |… See the full description on the dataset page: https://huggingface.co/datasets/ffurfaro/keep-it-simple-multimodal.simplified-google-speech-commands-wav2vec2-960hsimsamu
Dataset Card for the Simsamu dataset
This repository contains recordings of simulated medical dispatch dialogs in the french language, annotated for diarization and transcription. It is published under the MIT license.
These dialogs were recorded as part of the training of emergency medicine interns, which consisted in simulating a medical dispatch call where the interns took turns playing the caller and the regulating doctor.
Each situation was decided randomly in advance, blind to… See the full description on the dataset page: https://huggingface.co/datasets/diarizers-community/simsamu.SimpleDatasetaudsem-simple
AudSem Dataset without Semantic Descriptors
Accompanying paper: https://arxiv.org/abs/2505.14142.
GitHub repo: https://github.com/gljs/audsemthinker
Dataset Description
Overview
The AudSem-Simple dataset (audsem-simple) is a streamlined version of the AudSem dataset, designed to enhance the reasoning capabilities of Audio-Language Models (ALMs) through structured thinking over sound, but without explicit semantic element breakdowns.
The dataset includes four… See the full description on the dataset page: https://huggingface.co/datasets/gijs/audsem-simple.feji-first-200-simple-moodeddataset_MRtrain_de_en_de_not_similar_final
Dataset Card for "train_de_en_de_not_similar_final"
More Information needed
parler-large-v1-og_speaker_similaritydataset-tts-francais_HWK
Dataset Card for "dataset-tts-francais_HWK"
More Information needed
train_de_en_de_similar_final_4st_1000
Dataset Card for "train_de_en_de_similar_final_4st_1000"
More Information needed
commonvoice_13_0_pt_48kHz_simplificado
Dataset Card for "commonvoice_13_0_pt_48kHz_simplificado"
More Information needed
