datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Earnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.VietSuperSpeech
VietSuperSpeech
Vietnamese Speech Recognition Dataset
Dataset Information
Total samples: 32,267
Train samples: 29,041
Dev samples: 3,226
Total duration: 103.18 hours
Sample rate: 16000 Hz
Average segment length: ~12 seconds
Source Datasets
asr_dataset_nguoivietdailynews
asr_dataset_nguyenkhangofficial
asr_dataset_trinhlieu
Format
The dataset follows Icefall format:
train.json: Training samples
dev.json: Development samples
manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.StreamAudio-2M
StreamAudio-2M
Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a
stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips
are organised into six task subsets.
Subsets
Subset
Rows
Description
Stream_Audio_Understanding
90,738
Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA
Real_time_ASR
28,109
Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.ImageEval-ArabicNLP26
ImageEval-ArabicNLP26 👁️
ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026.
It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation.
The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.gdpval_preference_rubricsdahih-tts2-demucs-cleanedphoneme2Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.DCASE2026-Task5-DevSet
DCASE 2026 Task 5 Audio-Dependent Question Answering (ADQA) Development Set
This is the official Development Set for DCASE 2026 Challenge Task 5: Audio-Dependent Question Answering (ADQA).
The ADQA task focuses on addressing "Textual Hallucination" in Large Audio-Language Models (LALMs) — where models pass audio understanding benchmarks by relying on text prompts and internal linguistic priors rather than actual audio perception. ADQA introduces a rigorous evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Harland/DCASE2026-Task5-DevSet.Sympatheia-18k
Sympatheia-18k
Sympatheia-18k is an emotion-aware spoken dialogue dataset for empathetic speech synthesis research.
It contains 18,000 query–response pairs across 12 emotion categories, each accompanied by
synthesized audio and text transcripts.
Dataset Structure
Subset
Unique Queries
Responses
Description
Emotional
8,400 train / 3,600 eval
8,400 train / 3,600 eval
Emotional queries with emotionally-matched responses
Neutral
350 train / 150 eval
4,200 train /… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2222/Sympatheia-18k.music2chords_v2comfyui-wan22-assetsbharatvani-hindi-speech-corpus
BharatVani Hindi Speech Corpus (150-Hour Studio Dataset)
Proprietary Speech Asset • TheCreatorOS • BharatVani AI
1. Overview
The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered specifically for training sovereign neural Text-to-Speech (TTS) models, voice cloning engines, and speech foundation models in Devanagari Hindi.
Audio Clips: 103,784 Verified Studio Audio Clips (24,000 Hz, 16-bit Mono… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.VoxSafeBench
VoxSafeBench
This dataset is uploaded as raw files (JSONL + audio), not parquet.
Subset: Safety-tier1
Split: No_jailbreak (8708 samples)
Columns: system_prompt, clean_audio_file_name, diverse_audio_file_name, transcript, super_category, task_type, language, query
Split: Singleturn_jailbreak (2516 samples)
Columns: system_prompt, audio_file_name, transcript, source_text, super_category, jailbreak_type, task_type, language, query… See the full description on the dataset page: https://huggingface.co/datasets/nips26/VoxSafeBench.valor32k-avqa-v2
Valor32k-AVQA v2.0
Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position.
Links
Paper: ACM Digital Library
Project page: inesriahi.github.io/valor32k-avqa-2
Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.nightjar-flight-20260814
nightjar-flight-20260814
Elevation datum (fit_el_datum, 2026-08-16)
Status: FEW_ON_DRONE — this day's elevation datum is honestly UNSOLVABLE from
banked data (camera never/rarely locked on the drone).
Close-range elevation truth for this day remains datum-limited (~25 deg
floor). Tool: sirch613/subhunt-v2 v3/fit_el_datum.py; summary:
joshruby/acoustic-knowledge -> v3/assets/el_datum/SUMMARY.json.
cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.OpenGameArt-GPL-2.0
Dataset Card for OpenGameArt-GPL-2.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 2.0 (GPL-2.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-2.0.August25_ChickenZucchini
Filename structure
e.g. chicken_zucchini_speed-10_22g-12cm-shiba_stethoscope_2025-08-04_19.28.10
meaning: material-speed-needle_size-needle_length-needle_type-microphone_type-timestamp
Structure
Method of puncturing (Dobot)
Needle types
aishell1mix-ver2-n100-per-mix
AISHELL-1 Mix ver2 — 100 clips per mix
This is AISHELL-1 Mix ver2, not ver1. 8 kHz mono test subset: 100 mixtures per speaker count (N=1\ldots5) (50 mix_clean + 50 mix_both each) → 500 clips.
Derived from the local aishell1mix_ver2 test SCPs (data/scp/scp_aishell1mix_ver2). Includes mixture + oracle speaker stems and transcripts.
Split
Count
1mix / 2mix / 3mix / 4mix / 5mix
100 each
clean / both
250 each
Files
manifests/test.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/aishell1mix-ver2-n100-per-mix.bharatvani-hindi-showcase
BharatVani Hindi Speech Corpus • Public Interactive Showcase
150-Hour Enterprise Devanagari Hindi Speech Corpus & Precomputed Latents
Curated & Mastered by BharatVani AI • TheCreatorOS
1. Interactive Dataset Preview
This repository is the official public evaluation showcase for the 150-Hour BharatVani Hindi Speech Corpus (103,784 Studio Clips).
Use the Dataset Viewer above to play real audio clips, inspect the word-level timestamp alignments, and… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-showcase.thauMMAG-Benchmark
MMAG: A Multi‑Control Mixed Audio Generation Benchmark
MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment.
Dataset Structure
The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2026082026/MMAG-Benchmark.hi-en-noisy-vad-benchmark
Hindi-English Noisy VAD Benchmark
Version 0.1.0 is a deterministic, evaluation-only benchmark with 78
mono PCM16 WAV files at 16 kHz: six clean speech controls and 72 mixtures spanning
six speech sources, three real noise categories, and four SNRs (20, 10, 5, 0 dB).
Intended use
Use this dataset to compare voice-activity detectors under matched Hindi/English
noise conditions and to tune thresholds. It is too small and insufficiently diverse
for model training… See the full description on the dataset page: https://huggingface.co/datasets/Aakash22134/hi-en-noisy-vad-benchmark.Video_AMME_ci
Video-AMME CI
Video-AMME is a 50-case CI dataset derived from zhaochenyang20/Video_MME_ci.
Each example keeps the Video-MME video and moves the question, answer
choices, and answer-format instruction into a spoken WAV file.
Files
data/test.jsonl: metadata and source Video-MME references.
audios/*.wav: spoken question/options/instruction.
videos/*.mp4: present only when built with --copy-videos.
Generation
TTS model: fishaudio/s2-pro
Max samples requested: 50… See the full description on the dataset page: https://huggingface.co/datasets/zhaochenyang20/Video_AMME_ci.EGYSpeak
EGYSpeak
A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline.
Quick Start
1. Download the dataset:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="MohamedGomaa30/EGYSpeak",
repo_type="dataset",
local_dir="EGYSpeak",
)
2. Extract the dataset:
from… See the full description on the dataset page: https://huggingface.co/datasets/mahir2111/EGYSpeak.TRIAD
TRIAD: Benchmarking Omni-Modal Ambiguity in Multimodal Large Language Models
TRIAD is a diagnostic benchmark for evaluating whether omni-modal models can resolve ambiguity by jointly integrating text, image, and audio.
Overview
TRIAD targets a failure mode common in multimodal evaluation: a model may appear to solve a multimodal task while actually relying on only one dominant modality. In TRIAD, each example is constructed so that the full tri-modal input is needed… See the full description on the dataset page: https://huggingface.co/datasets/triad-26/TRIAD.stage1a_smoke_data
stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh)
Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP →
frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend:
WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT
encoder — no raw-audio decoding at train time.
113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.yt2_chunked_tokenized
yt2_chunked_48k_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_tokenized.cmi-annotate
