datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Earnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.laser-vibrations
Laser Vibrations
Dataset of laser speckle vibration recordings used to locate objects hidden inside a cardboard box.
A 10×10 grid of lasers shines on the side of a box; as loudspeakers excite the box, each laser's
speckle pattern shifts in proportion to the local surface vibration. The goal is to reconstruct the
shape and location of an object inside the box from the vibration signals alone.
Per-sample viewer metadata lives in data/metadata.jsonl.
Full signal data and media files… See the full description on the dataset page: https://huggingface.co/datasets/eturok-weizmann/laser-vibrations.Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.asr-reference-set-eval-temp
Temporary ASR evaluation audio
Temporary public audio files used for hosted ASR evaluation.
cs-envi-dual-encoder-60EMOaudio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.EGYSpeak
EGYSpeak
A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline.
Quick Start
1. Download the dataset:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="MohamedGomaa30/EGYSpeak",
repo_type="dataset",
local_dir="EGYSpeak",
)
2. Extract the dataset:
from… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/EGYSpeak.OST-Diagbench
OST-DiagBench
OST-DiagBench contains exactly four diagnostic operations with 510 cases per operation: 2,040 cases in total. Every case includes an MP4 input and a paired M4A audio file.
Academic research use only, no commercial use. The clips are short, transformed excerpts derived from public academic audio-visual corpora whose material originates from user-uploaded YouTube content; the donor sounds come from ESC-50 under CC BY-NC 3.0. We are not the rights holders of the… See the full description on the dataset page: https://huggingface.co/datasets/EnjunDu/OST-Diagbench.hawza-asr-evalA test dataset for evaluate ASR (Automatic Speech Recognition) models in the domain of Islamic lectures and specialized Hawza courses.
Audio files are mono 16khz wav.
Texts are verified.
alimeeting-eval-8k
AliMeeting Eval — 8 kHz CH0 clips
Unknown-(N) eval clips from AliMeeting Eval (M2MeT / OpenSLR 119), far-field channel 0, resampled to 8 kHz.
Manifest
Clips
manifests/eval_all.jsonl
1280 (full local eval)
manifests/eval_n100.jsonl
100 stratified subset
manifests/eval_n200.jsonl
200 stratified subset
manifests/eval_*mix.jsonl
by (N=1\ldots4)
Headset s{k}.wav stems (near, TextGrid-gated) are present for the n200 subset (eval_n200_headset.jsonl). Other clips… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/alimeeting-eval-8k.uzbek-asr-curated-701h
Uzbek ASR Curated Dataset (701 hours)
A curated multi-source Uzbek speech dataset for automatic speech recognition (ASR) training and evaluation.
Dataset Description
Language
Uzbek (Latin script with okina ʻ)
Total utterances
337,920
Total duration
~701 hours
Audio format
16 kHz mono WAV (PCM_16)
Manifest format
NeMo JSONL
Splits
train (94%) / val (3%) / test (3%)
Splits
Split
Utterances
Hours
Train
317,655… See the full description on the dataset page: https://huggingface.co/datasets/uzinfocom-edu-ai/uzbek-asr-curated-701h.multilingual-speech
Multilingual Indian Conversational Speech
A dataset of naturalistic, spontaneous two-speaker conversations across
13 Indian languages, with segment-level transcripts, speaker profiles,
timestamps, and recording metadata. Designed for ASR, TTS, speaker
diarization, and conversational speech research.
Languages (13)
Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali,
Odia, Punjabi, Tamil, Telugu, Urdu.
Content
Conversations… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/multilingual-speech.event-bench
Event Bench
29-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as an event planning assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as an event planning assistant managing venue bookings, catering, and guest logistics. The conversation features cascading changes — a venue switch triggers catering… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/event-bench.vocence_eval_corpus
Dataset Card for vocence_eval_corpus
Dataset Summary
Everything produced by evaluating 7 PromptTTS systems (gemini, voxcpm, qwen3, maya1,
openai, parler-tts, elevenlabs) on the 200-item balanced subset of
vocence_corpus: the
generated audio, every per-clip judge output, and the aggregated leaderboards/charts.
For the prompt corpus alone (no eval data), see vocence_corpus.
Scoring library: vocencebench
(source)
Benchmark code + full methodology write-up:… See the full description on the dataset page: https://huggingface.co/datasets/vocence/vocence_eval_corpus.telugu-indicf5-evaluationEnglish-Hebrew-Mixed-Sentences
English-Hebrew Mixed Sentences Dataset
A dataset of English sentences with Hebrew words and phrases interspersed, designed for speech-to-text training and evaluation for English speakers in Israel.
Overview
This dataset addresses a common challenge for English-speaking immigrants in Israel: standard speech-to-text (STT) systems struggle to accurately transcribe code-switched speech where Hebrew words are mixed into primarily English sentences.
Example: "I need to pick up… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/English-Hebrew-Mixed-Sentences.echo-embeddings-vctk-tar
VCTK Speaker Embeddings (tarred)
Items: 109
This dataset ships as a single tar at the repo root. Members preserve paths like
VCTK/<id>/audio.mp3 and VCTK/<id>/speaker_latent.safetensors.
See loader.py for example loading.
Attribution:
Contains audio and embeddings derived from the CSTR VCTK Corpus. Distributed under CC BY 4.0; attribution required.
echo-embeddings-expresso-tar
Expresso Speaker Embeddings (tarred)
Items: 17
This dataset ships as a single tar at the repo root. Members preserve paths like
Expresso/<id>/audio.mp3 and Expresso/<id>/speaker_latent.safetensors.
See loader.py for example loading.
Attribution:
Contains audio and embeddings derived from the Expresso dataset (INTERSPEECH 2023). Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted.
echo-embeddings-custom
Custom Speaker Embeddings
Contains speaker folders within HF-Custom, each with:
a precomputed speaker embedding (speaker_latent.safetensors)
its corresponding audio (audio.mp3)
a metadata file describing the voice and licensing (metadata.json)
Licensing:There is no single license for this dataset. Each voice has its own terms stored
in its metadata.json. You must check the metadata for any voice you use.
odyssey-v0.5.1-16h-eval-preview-public
🇿🇦 Odyssey V0.5.1 — 16H Evaluation Preview (Code-Switched ASR)
Odyssey AI Labs presents a 16-hour Enterprise Evaluation Preview of the V0.5 corpus. This release is designed to surface the hardest South African ASR realities directly: native urban code-switching, multi-speaker turn-taking, and overlapping conversational speech.
A structured public preview of the Odyssey SA Voice Corpus, designed for researchers, data buyers, and speech teams evaluating multilingual South African… See the full description on the dataset page: https://huggingface.co/datasets/ODYSSEYAILABS/odyssey-v0.5.1-16h-eval-preview-public.echo-embeddings-ears-tar
EARS Speaker Embeddings (tarred)
Items: 2568
This dataset ships as a single tar at the repo root. Members preserve paths like
EARS/<id>/audio.mp3 and EARS/<id>/speaker_latent.safetensors.
See loader.py for example loading.
Attribution:
Contains audio and embeddings derived from the EARS dataset. Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted.
thud-eval
THUD-Eval · audio-visual Clever Hans benchmark
Evaluation benchmark accompanying the paper
When Vision Speaks for Sound.
This dataset probes the audio-visual Clever Hans effect — the tendency
of video-capable MLLMs to appear to listen while really just reading
visual cues. We test the same source clips under three audio
interventions:
Task
Intervention
What it tests
sync
audio temporally shifted (early / delay)
Can the model detect a time offset?
mute
audio replaced… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/thud-eval.EGYSpeak
EGYSpeak
A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline.
Quick Start
1. Download the dataset:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="MohamedGomaa30/EGYSpeak",
repo_type="dataset",
local_dir="EGYSpeak",
)
2. Extract the dataset:
from… See the full description on the dataset page: https://huggingface.co/datasets/mahir2111/EGYSpeak.asr-evaluationshi-en-noisy-vad-benchmark
Hindi-English Noisy VAD Benchmark
Version 0.1.0 is a deterministic, evaluation-only benchmark with 78
mono PCM16 WAV files at 16 kHz: six clean speech controls and 72 mixtures spanning
six speech sources, three real noise categories, and four SNRs (20, 10, 5, 0 dB).
Intended use
Use this dataset to compare voice-activity detectors under matched Hindi/English
noise conditions and to tune thresholds. It is too small and insufficiently diverse
for model training… See the full description on the dataset page: https://huggingface.co/datasets/Aakash22134/hi-en-noisy-vad-benchmark.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.emotion-preserving-s2st-benchmark
Emotion-Preserving Speech-to-Speech Translation Benchmark
Overview
First benchmark for evaluating emotion preservation in speech-to-speech translation systems.
Pipeline: English speech → Emotion Detection → Translation (EN→HI) → Emotion-conditioned TTS.
Key Results (1440 samples, RAVDESS dataset)
Metric
Score
Emotion Detection Accuracy
36.3% (4-class on 8-class data)
Emotion Preservation Rate
43.5%
Preservation Gain over Flat TTS
+7.2%
F0… See the full description on the dataset page: https://huggingface.co/datasets/primal-sage/emotion-preserving-s2st-benchmark.voxtral-emotion-temporal
VoxTral Emotion Temporal Dataset
Dataset for training emotion transition detection in speech. ~500 clips with frame-level emotion annotations at 20ms resolution.
Overview
Property
Value
Clips
~500
Sample Rate
16000 Hz
Frame Resolution
20ms
Emotions
6 (neutral, happy, angry, sad, surprise, fear)
Clip Distribution
Single Emotion (40%): One emotion throughout
Transitions (40%): Emotion changes mid-sentence (2-3 segments)
No Emotion (20%):… See the full description on the dataset page: https://huggingface.co/datasets/MrlolDev/voxtral-emotion-temporal.
