datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BEAT2AV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.beat2-additional-annotations
BEAT2 Official Release + Additional Annotations
This is a fork of H-Liu1997/BEAT2
that adds annotations contributed by the
RAG-Gesture (CVPR 2025)
and MIBURI (CVPR 2026) projects.
The base BEAT2-English data (motion, audio, TextGrids, semantic labels,
pretrained motion-autoencoder weights) is inherited verbatim from upstream;
the additional annotations from RAG-Gesture and MIBURI are pushed on top.
Citations
If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.serena-synthetic-it-28h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.TUT2018-ov2AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.dia-earning21-all
Earnings 21
The Earnings 21 dataset ( also referred to as earnings21 ) is a 39-hour corpus of earnings calls containing entity dense speech from nine different financial sectors. This corpus is intended to benchmark automatic speech recognition (ASR) systems in the wild with special attention towards named entity recognition (NER).
This work has been recently accepted to Interspeech 2021!
File Format Overview
In the following section, we provide an overview of the file… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-earning21-all.sada2022
Dataset Card for SADA صدى
Dataset Summary
يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر.
ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر مجموعة… See the full description on the dataset page: https://huggingface.co/datasets/khaledalganem/sada2022.sinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.TUT2018-ov3Copyright (c) 2018 Tampere University of Technology and its licensors
All rights reserved.
Permission is hereby granted, without written agreement and without license or royalty
fees, to use and copy the TUT Sound Events 2018 - Ambisonic, Reverberant and Real-life Impulse Response Dataset (“Work”)
described in this document and composed of audio and metadata. This grant is only
for experimental and non-commercial purposes, provided that the copyright notice
in its entirety appear in all… See the full description on the dataset page: https://huggingface.co/datasets/labhamlet/TUT2018-ov3.TUT2018-ov1TUT2017
TUT2017
This is an audio classification dataset for Acoustic Scene Classification.
Classes = 15 , Split = four-fold
Structure
audios folder contains audio files.
csv_files folder contains CSV files for four-fold cross-validation.
To perform cross-validation on fold 1, train_1.csv will be used for the training split and test_1.csv for the testing split, with the same pattern followed for the other folds.
To perform training and testing witout cross-validation, use… See the full description on the dataset page: https://huggingface.co/datasets/MahiA/TUT2017.serena-synthetic-it-27h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.sada2022-arabic-tts
SADA 2022 - Saudi Arabic Dataset for TTS
مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech)
المصدر الأصلي
Kaggle: sdaiancai/sada2022
الاستخدام
# طريقة 1: Git Clone
!git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset
# طريقة 2: مكتبة datasets
from datasets import load_dataset
dataset = load_dataset("aalshalfi/sada2022-arabic-tts")
الملفات
valid.csv - ملف البيانات الرئيسي
wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts.bhavvani
Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning
This repository contains the BhavVani dataset introduced in the INTERSPEECH 2024 Paper :
Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning
Please fill this form for accessing the audio files associated with the BhavVani dataset: Form Link
Overview
In our work, we propose the following contributions:… See the full description on the dataset page: https://huggingface.co/datasets/ag2003/bhavvani.SADA_khaledalganemsada2022_Rawdate
Dataset Card for SADA صدى
Dataset Summary
يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر.
ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر… See the full description on the dataset page: https://huggingface.co/datasets/Sundus246/SADA_khaledalganemsada2022_Rawdate.DEAF
DEAF
DEAF is a collection of audio data, aligned text metadata, and data-generation scripts accompanying the paper DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models.
This repository is organized as a Hugging Face dataset repository and contains the locally hosted resources used in the paper: BSC audio, SIC audio, paired text metadata, and the scripts used to generate the speech-related subsets.
Repository structure… See the full description on the dataset page: https://huggingface.co/datasets/Wonder239/DEAF.TUT2018-ov2-SELDhuman-perception-audio-deepfake-2026
Human Audio Deepfake Perception 2026
A large-scale listening study evaluating how well humans detect modern audio
deepfakes. The dataset contains 35,532 deepfake-detection judgments from
1,768 anonymous participants across 138 TTS and voice-conversion systems,
collected via a publicly accessible online listening game in 2025–2026.
This is the successor to the 2021 ASVspoof-2019 perception study
(Müller, Pizzi & Williams, 2022)
and extends the same paradigm to modern systems… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/human-perception-audio-deepfake-2026.arabic-tts-saudi-multi-speaker-xtts
Arabic Saudi TTS Dataset (LJSpeech Format) 🇸🇦
This dataset is designed for training Text-to-Speech (TTS) models such as XTTS_v2 using the LJSpeech format.
📌 Overview
Language: Arabic (Saudi Dialect)
Format: LJSpeech
Use Case: TTS training (XTTS_v2, YourTTS, Tacotron, etc.)
Speakers: Multi-speaker (Male & Female)
Audio Format: WAV (mono recommended)
Sample Rate: 22050 Hz (recommended)
📂 Structure
all_data/
│
├── wavs/
│ ├── sample_0.wav
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman2922/arabic-tts-saudi-multi-speaker-xtts.YDX07_Multilingual_Corpus_2026aatc2025simVN-SpeechMix_Dataset
VN-SpeechMix: A Large-Scale Multi-Dialect Vietnamese Speech Mixture Dataset
VN-SpeechMix is a large-scale, multi-dialect Vietnamese speech mixture
dataset for two-speaker speech separation research. It is built from the
ViMD corpus (Van Dinh et al., EMNLP 2024)
using a loudness-aware mixing pipeline (LUFS normalization + two-stage
anti-clipping) and a dialect-aware pairing strategy across Vietnam's three
macro-dialect regions (North / Central / South).
26,000 two-speaker… See the full description on the dataset page: https://huggingface.co/datasets/pervasiveaidataresearchlab2025/VN-SpeechMix_Dataset.Yadonay-YDX07_Multilingual_Corpus_2026amharic-speech-dataset-2026
Amharic Speech Dataset 2026
Overview
This dataset contains Amharic speech recordings collected using the Leyu Platform for the Leyu Platform Competition 2026.
Language
Amharic (am)
Dialect
Standard Addis Ababa Amharic
Speaker Information
Number of Speakers: 1
Speaker IDs: SPK001
Audio Format
Format: M4A
Duration: 10–60 seconds per recording
Directory Structure
audio/
metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/ofc-its-phyla/amharic-speech-dataset-2026.emotion-voice-dataset
emotion voice dataset
Developed by Aryan Singh Chandel (Shiro) at Rustamji Institute of Technology (RJIT).
📝 Overview
This repository contains assets for emotion voice dataset. It is a professional research component of the Shiro AI ecosystem.
🚀 Status
The core files are live. Detailed usage instructions and technical benchmarks are currently being compiled for the elite release.
2026-dwesui-g01-neurologia
DWESUI 2026 - Grupa 1 - neurologia
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 1 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Domena: neurologia
Licencja zrodla: nagrania YouTube CC-BY + synteza TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował zbiór… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g01-neurologia.CommonVoices20_ro
Common Voices Corpus 20.0 (Romanian)
Common Voices is an open-source dataset of speech recordings created by
Mozilla to improve speech recognition technologies.
It consists of crowdsourced voice samples in multiple languages, contributed by volunteers worldwide.
Challenges: The raw dataset included numerous recordings with incorrect transcriptions
or those requiring adjustments, such as sampling rate modifications, conversion to .wav format, and other refinements
essential… See the full description on the dataset page: https://huggingface.co/datasets/TransferRapid/CommonVoices20_ro.vox2-veri-full
VoxCeleb 2
VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
5,994
5,994
118
# of samples
982,808
109,201
36,237
Data Fields
ID (string): The ID of the sample with format <spk_id--utt_id_start_stop>.
duration (float64): The duration of the segment in seconds.
wav (string): The filepath of the waveform.
start (int64): The… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/vox2-veri-full.dwesui-grupa-2-kulinarna
G2-Polish-Culinary-ASR-Evaluation-Corpus
Korpus do ewaluacji systemow ASR jezyka polskiego (domena kulinarna) stworzony
w ramach warsztatow Ewaluacja Systemow Rozpoznawania Mowy (UAM WMI, edycja 2026,
zespol 2). Publikowany podzbior to mowa naturalna z wideo kulinarnych YouTube
(licencja CC-BY) - sluzy do badania odpornosci ASR na szum kuchenny oraz dopasowania
domenowego do specjalistycznego slownictwa (zapozyczenia, miary, liczby).
Pelny eksperyment ewaluacyjny zespolu… See the full description on the dataset page: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna.
