datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vibevoice-quran_persian-single-speaker100-hour-Egyptian-dataset-single-speaker
Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus
A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts.
Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data.
Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home.
This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.AV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.vibevoice-gptinformal_persian-single-speakervibevoice-channelbpodcast-single-speakersingle_speakereurospeech-bg-single-speaker
EuroSpeech BG — single-speaker subset
Bulgarian parliamentary speech from disco-eth/EuroSpeech,
filtered down to clips containing exactly one speaker.
Why
EuroSpeech ships no speaker labels — its only identity-like field, video_id,
is a parliamentary session containing dozens of speakers. To build
LibriSpeechMix-style simulated mixtures for speaker-diarization training you
first need clean single-speaker source audio. This subset is that source.
How… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/eurospeech-bg-single-speaker.experiment-speaker-embeddinguzbek-multi-speaker-35hIndicDiarBench-speaker-retrieval
Indic DiarBench speaker retrieval (MTEB)
Indic DiarBench reshaped for speaker retrieval across the 22 scheduled languages
of India: given a clip of one speaker, find other clips of that same speaker.
Source: sarvamai/indic-diarbench at revision 92877ba, cc-by-4.0, official
test split. Turns are cut by their annotated times, restricted to 2 to 15
seconds, and turns overlapping a different speaker are dropped. identity pairs
the recording session with the speaker, because speaker… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/IndicDiarBench-speaker-retrieval.iv_speaker_disjoint_sociodem_unaware_dsSpeakerVerification_Aishell1Train
Dataset Card for "SpeakerVerification_AISHELL1Train"
More Information needed
russian-single-speaker-speech-datasetMSPP_WAV_speaker_splitvibevoice-movarekhpodcast-single-speakercommon_voice_17-pl-speakers-v1.0Available sources(voices):
tomasz
konrad
jan
wojciech
krystian
emilia
kacper
karol
henryk
marek
kinga
szymon
robert
milena
jagoda
marcel
edmund
weronika
maciej
hubert
aniela
artur
stefan
ireneusz
grzegorz
roman
leon
maksymilian
zygmunt
jerzy
piotr
edward
bernard
aleksander
arkadiusz
zuzanna
melania
ignacy
lucjan
iga
patryk
bogdan
adam
dominika
adrian
antoni
bartosz
marian
cezary
ludwik
zenon
ryszard
feliks
filip
franciszek
mateusz
witold
igor
marcin
dominik
julian
kazimierz
sebastian
mariusz… See the full description on the dataset page: https://huggingface.co/datasets/TeeZee/common_voice_17-pl-speakers-v1.0.ruv_tv_unknown_speakersDataset copied from http://hdl.handle.net/20.500.12537/191 by Reykjavik University.
Information can be found at that link.
RUV TV unknown speakers
About the RUV TV unknown speakers corpus
The RUV TV unknown speakers corpus is 281 hours of TV data from six RÚV TV
shows. The data continas 221,759 utterrances from various unlabelled speakers.
The text is normalized. The data is aligned and segmented, ready for ASR
training. Audio conditions vary between recordings. This data set is… See the full description on the dataset page: https://huggingface.co/datasets/tiro-is/ruv_tv_unknown_speakers.speech_with_speaker_idsingle_speaker_en_test_librivox
Dataset Card for "single_speaker_en_test_librivox"
Created for testing, not suggested for production
Dataset Summary
The corpus consists of a single speaker extracted frrom LibriVox audiobook.
Languages
The audio is in English.
Source Data
Initial Data Collection and Normalization
The voices used in my Datasets are volenteers who have donated their time and voices to open source LibriVox projects. Please respect their privacy.… See the full description on the dataset page: https://huggingface.co/datasets/sjdata/single_speaker_en_test_librivox.speaker-datasetsSpeakerVerification_Tedlium2Train
Dataset Card for "SpeakerVerification_TEDLIUM2Train"
More Information needed
romanian_speech_dataset_with_15_percent_6_speakers_synthetic_datahungarian-single-speaker-tts
Dataset Card for CSS10 Hungarian: Single Speaker Speech Dataset
Dataset Summary
The corpus consists of a single speaker, with 4515 segments extracted
from a single LibriVox audiobook.
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in Hungarian.
Dataset Structure
[Needs More Information]
Data Instances
[Needs More Information]
Data Fields
[Needs More Information]
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/KTH/hungarian-single-speaker-tts.speaker_identification_100_speakerscommon_voice_17-pl-speakers-v2.1speaker_datasets_parler_v1resampled_16KHrz_vctk_speakers_splitar-eg-speech-tts-multi-speakersaudio-3-speaker-dataset-v2
Three-Speaker Audio Dataset with Timbral Speaker Embeddings
A teaching dataset maintained by AI-Academy. It pairs single-speaker English
utterances with precomputed timbral speaker embeddings, and is intended for
coursework and exercises rather than for benchmarking or production systems.
The dataset deliberately contains one injected inconsistency; locating it is one of
the intended exercises (see The injected anomaly).
Overview
Property
Value
Examples… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/audio-3-speaker-dataset-v2.synthetic-multilingual-speaker-diarization
Synthetic Multilingual Speaker Diarization Dataset
This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples.
Dataset Structure
├── audio/ # WAV audio files (16kHz) - 3417 files
├── all_samples_combined.csv # Complete dataset annotations (with silence)
└── all_visible_combined.csv # Visible dataset annotations (without silence)
Statistics
Total samples: 3417 audio… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/synthetic-multilingual-speaker-diarization.
