datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seamless-align-enA-viA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.w2vbert-600mseamless-align-enA-frA.speaker-embedding.hubert-xlseamless-align-enA-jaA.speaker-embedding.w2vbert-600mseamless-align-deA-enA.speaker-embedding.xlsr-2bseamless-align-enA-hiA.speaker-embedding.hubert-xlseamless-align-enA-frA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.xlsr-2bseamless-align-enA-zhA.speaker-embedding.w2vbert-600mseamless-align-enA-frA.speaker-embedding.w2vbert-600mseamless-align-enA-zhA.speaker-embedding.xlsr-2bseamless-align-enA-zhA.speaker-embedding.hubert-xlseamless-align-enA-koA.speaker-embedding.w2vbert-600mseamless-align-enA-hiA.speaker-embedding.w2vbert-600mseamless-align-enA-viA.speaker-embedding.w2vbert-600m100-hour-Egyptian-dataset-single-speaker
Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus
A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts.
Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data.
Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home.
This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.seamless-align-enA-jaA.speaker-embedding.hubert-xlseamless-align-enA-hiA.speaker-embedding.xlsr-2bseamless-align-enA-jaA.speaker-embedding.xlsr-2bvibevoice-quran_persian-single-speakerAV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.seamless-align-enA-koA.speaker-embedding.hubert-xlseamless-align-deA-enA.speaker-embedding.w2vbert-600mseamless-align-enA-koA.speaker-embedding.xlsr-2bvibevoice-gptinformal_persian-single-speakerseamless-align-enA-esA.speaker-embedding.hubert-xlvibevoice-channelbpodcast-single-speakersingle_speakerseamless-align-enA-viA.speaker-embedding.hubert-xleurospeech-bg-single-speaker
EuroSpeech BG — single-speaker subset
Bulgarian parliamentary speech from disco-eth/EuroSpeech,
filtered down to clips containing exactly one speaker.
Why
EuroSpeech ships no speaker labels — its only identity-like field, video_id,
is a parliamentary session containing dozens of speakers. To build
LibriSpeechMix-style simulated mixtures for speaker-diarization training you
first need clean single-speaker source audio. This subset is that source.
How… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/eurospeech-bg-single-speaker.
