datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.SEAR
SEAR: Spoofing Evidence-Grounded Audio Reasoning
SEAR is an audio question-answering benchmark for testing whether audio language models
can identify and quantify signal-level acoustic anomalies and use them as evidence for
audio deepfake detection. Instead of evaluating only a final authenticity verdict, SEAR
separates deepfake detection, forgery-cue identification, acoustic measurement, and
forensic rationale generation.
SEAR contains four complementary tasks covering acoustic… See the full description on the dataset page: https://huggingface.co/datasets/SEAR-benchmark/SEAR.MSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/Khalilah-Shields/MSU-Benchmark.darija-asr-benchmark-6speaker
Darija ASR 6-Speaker Benchmark
A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3
female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus),
used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi)
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
Consent and anonymization
Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.MMAG-Benchmark
MMAG: A Multi‑Control Mixed Audio Generation Benchmark
MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment.
Dataset Structure
The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2026082026/MMAG-Benchmark.hi-en-noisy-vad-benchmark
Hindi-English Noisy VAD Benchmark
Version 0.1.0 is a deterministic, evaluation-only benchmark with 78
mono PCM16 WAV files at 16 kHz: six clean speech controls and 72 mixtures spanning
six speech sources, three real noise categories, and four SNRs (20, 10, 5, 0 dB).
Intended use
Use this dataset to compare voice-activity detectors under matched Hindi/English
noise conditions and to tune thresholds. It is too small and insufficiently diverse
for model training… See the full description on the dataset page: https://huggingface.co/datasets/Aakash22134/hi-en-noisy-vad-benchmark.emotion-preserving-s2st-benchmark
Emotion-Preserving Speech-to-Speech Translation Benchmark
Overview
First benchmark for evaluating emotion preservation in speech-to-speech translation systems.
Pipeline: English speech → Emotion Detection → Translation (EN→HI) → Emotion-conditioned TTS.
Key Results (1440 samples, RAVDESS dataset)
Metric
Score
Emotion Detection Accuracy
36.3% (4-class on 8-class data)
Emotion Preservation Rate
43.5%
Preservation Gain over Flat TTS
+7.2%
F0… See the full description on the dataset page: https://huggingface.co/datasets/primal-sage/emotion-preserving-s2st-benchmark.
