datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.Sympatheia-18k
Sympatheia-18k
Sympatheia-18k is an emotion-aware spoken dialogue dataset for empathetic speech synthesis research.
It contains 18,000 query–response pairs across 12 emotion categories, each accompanied by
synthesized audio and text transcripts.
Dataset Structure
Subset
Unique Queries
Responses
Description
Emotional
8,400 train / 3,600 eval
8,400 train / 3,600 eval
Emotional queries with emotionally-matched responses
Neutral
350 train / 150 eval
4,200 train /… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2222/Sympatheia-18k.multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.Sympatheia-18k
Sympatheia-18k
Sympatheia-18k is an emotion-aware spoken dialogue dataset for empathetic speech synthesis research.
It contains 18,000 query–response pairs across 12 emotion categories, each accompanied by
synthesized audio and text transcripts.
Dataset Structure
Subset
Unique Queries
Responses
Description
Emotional
8,400 train / 3,600 eval
8,400 train / 3,600 eval
Emotional queries with emotionally-matched responses
Neutral
350 train / 150 eval
4,200… See the full description on the dataset page: https://huggingface.co/datasets/susameddin/Sympatheia-18k.SVHalluc
SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models
TL;DR: Speech content is not necessarily visual evidence. SVHalluc is the first benchmark that tests whether audio-visual large language models (AV-LLMs) can distinguish what is said from what is actually seen.
SVHalluc was introduced in “SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models,” CVPR 2026.
Project page
Paper (arXiv)
Code… See the full description on the dataset page: https://huggingface.co/datasets/zcs15/SVHalluc.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.jam-actions-v1-probe
jam-actions-v1-probe
Schema: jam-actions-v1-probe/1.0.0 · Records: 24, all split: test · Evaluation only ·
Companion to: jam-actions-v1
Why it exists
An adapter trained on an earlier version of the corpus scored 47/54 on held-out acoustic takes.
Its completions, which state the comparison before the label, showed that it wrote against a 50-cent gate whenever it saw a minus sign — and negative cents occurred in exactly one class of that
corpus. The main split could… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1-probe.Music
Made by herza For APP.t
agentic-asr
Agentic ASR
Public consolidated audio and ASR result dataset for the OSWorld and
WildClawBench benchmark families.
Layout
osworld/: synthetic raw/colloquial speech, human recordings, DNS-noise
pairs, task images, ASR results, and reports.
wildclawbench/: 60 formal colloquialized prompts, synthetic speech,
20 synthetic ASR condition tables, and ten-participant human recordings.
task0_template derivatives are excluded.
metadata/conditions.jsonl: model, variant… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/agentic-asr.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.Chords1217open-vi-dialog-synthetic-100h
OpenDialog Vietnamese Synthetic Dialogue 100h
Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments.
12,000 chunks
30 seconds per chunk
100.0 hours total
Each item contains S1/S2 speaker labels, turn timings, target text,
relationship, pronouns, environment, topic, mood, and source reference IDs.
Audio renderer: vLLM-Omni VoxCPM2
Audio format: mono WAV, 48 kHz, 30 seconds per chunk
This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.PICSAFEv1
Speech Quality Test Labels
PICSAFEv1 is a multi-source annotated test dataset for evaluating speech quality assessment and audio data filtering methods. It contains 10,728 audio samples drawn from 14 source datasets, with annotations from a vocabulary of 33 tags. These tags describe recording provenance, speech styles, speaking rate and pitch, speaker attributes, noise, reverberation, distortion, and transcript errors. These annotations support benchmarking quality metrics and… See the full description on the dataset page: https://huggingface.co/datasets/AudioCC-Lab/PICSAFEv1.rgad-crosslingual-tts-10h
RGAD Cross-Lingual TTS 10h
This is a 10-hour cross-lingual TTS dataset for prompt-conditioned Chinese TTS fine-tuning.
Format
The dataset contains:
train.jsonl
dev.jsonl
metadata.csv
audio/prompts/*.wav
audio/targets/*.wav
Each JSONL row has this format:
{"id":"sample_000001","prompt_wav":"audio/prompts/sample_000001.wav","target_wav":"audio/targets/sample_000001.wav","text":"中文目标文本。","prompt_language":"en-US","target_language":"zh-CN"… See the full description on the dataset page: https://huggingface.co/datasets/isabeth/rgad-crosslingual-tts-10h.meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Duckyle/meow-10k.fdb-v1-outputs-v1
Full-Duplex-Bench v1.0 model outputs (fdb-v1-outputs-v1)
7997 model responses = 11 benchmark runs × the 727 stimuli of
Full-Duplex-Bench v1.0 (pause handling,
backchannel, smooth turn-taking, user interruption). The runs cover 7 systems;
several differ only in voice prompt, prompting regime or weights, which is the point — those are
controlled pairs. For every stimulus and run you get the
model's own reply channel as lossless FLAC — time-synchronous with the stimulus, so… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/fdb-v1-outputs-v1.aishell1mix-ver2-n100-per-mix
AISHELL-1 Mix ver2 — 100 clips per mix
This is AISHELL-1 Mix ver2, not ver1. 8 kHz mono test subset: 100 mixtures per speaker count (N=1\ldots5) (50 mix_clean + 50 mix_both each) → 500 clips.
Derived from the local aishell1mix_ver2 test SCPs (data/scp/scp_aishell1mix_ver2). Includes mixture + oracle speaker stems and transcripts.
Split
Count
1mix / 2mix / 3mix / 4mix / 5mix
100 each
clean / both
250 each
Files
manifests/test.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/aishell1mix-ver2-n100-per-mix.odyssey-v0.5.1-16h-eval-preview-public
🇿🇦 Odyssey V0.5.1 — 16H Evaluation Preview (Code-Switched ASR)
Odyssey AI Labs presents a 16-hour Enterprise Evaluation Preview of the V0.5 corpus. This release is designed to surface the hardest South African ASR realities directly: native urban code-switching, multi-speaker turn-taking, and overlapping conversational speech.
A structured public preview of the Odyssey SA Voice Corpus, designed for researchers, data buyers, and speech teams evaluating multilingual South African… See the full description on the dataset page: https://huggingface.co/datasets/ODYSSEYAILABS/odyssey-v0.5.1-16h-eval-preview-public.ONOTE
ONOTE: Omnimodal Notation Objective Topology Examination
ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature.
📂 Dataset Structure
The dataset is organized into two primary sub-directories based on the notation and instrument type:
1. pitch_Jianpu_dataset (Staff & Numbered Notation)
This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/Weisiqing123/ONOTE.ifbench-conversations-v1
IF-Bench conversations (ifbench-conversations-v1)
3000 synthetic full-duplex spoken conversations: 15 examiner configurations
(9 speech models), each holding the same 200-task set of Full-Duplex-Bench v2 staged scenarios as the
examiner (the model under study — it carries a role, a topic and four goals to hit in order)
against a PersonaPlex-7B examinee that is never told the topic. Per dialogue you get both
channels as lossless mono FLAC, the exact prompts/voices/sampling… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/ifbench-conversations-v1.hans-10k
Hans-10K · DPO recipe for the audio-visual Clever Hans
DPO training data accompanying the paper
When Vision Speaks for Sound.
Like the original Clever Hans 🐎 —
the horse that looked like he could do arithmetic but was actually reading
his trainer's body language — video-capable MLLMs often look like they
can hear: they answer audio questions by reading visual cues and never
verifying the audio stream.
Hans-10K is the 10,383-sample best-recipe preference-pair dataset
that cures this… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-10k.thud-eval
THUD-Eval · audio-visual Clever Hans benchmark
Evaluation benchmark accompanying the paper
When Vision Speaks for Sound.
This dataset probes the audio-visual Clever Hans effect — the tendency
of video-capable MLLMs to appear to listen while really just reading
visual cues. We test the same source clips under three audio
interventions:
Task
Intervention
What it tests
sync
audio temporally shifted (early / delay)
Can the model detect a time offset?
mute
audio replaced… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/thud-eval.advanced-soundscapes-stage-1
Advanced Soundscapes Stage 1 — Raw Components
This dataset contains Stage 1 output from the LAION Universal Audio Annotation Pipeline (UAAP) data generation plan.
Contents
0 shard(s) containing 0 soundscape recipes with raw audio components
Each soundscape row includes:
recipe.json — full recipe with timeline, events, loudness, speaker IDs, overlap/density settings
spkN.flac / spkN.json — raw speech components + full source metadata
musicN.flac / musicN.json —… See the full description on the dataset page: https://huggingface.co/datasets/ChristophSchuhmann/advanced-soundscapes-stage-1.vggsync-3k
VGGSync-3K · out-of-domain audio-visual sync benchmark
Out-of-domain evaluation set used in the paper
When Vision Speaks for Sound.
Derived from VGGSoundSync,
this 3,000-clip slice tests whether a video-capable MLLM can detect
audio temporal offsets on everyday sound events outside the THUD
in-domain training distribution.
Each item is one VGGSound clip in one of three conditions:
Condition
Count
gt_synced
gt_direction
gt_offset_sec
Audio aligned (no shift)
1,000
true… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/vggsync-3k.IFAO-lalmCaptionStew10M-Qwen3Omnihans-sft-4k
Hans-SFT-4K · SFT recipe for the audio-visual Clever Hans
Supervised fine-tuning (SFT) data accompanying the paper
When Vision Speaks for Sound.
Like the original Clever Hans —
the horse that looked like he could do arithmetic but was actually reading
his trainer's body language — video-capable MLLMs often look like they
can hear: they answer audio questions by reading visual cues and never
verifying the audio stream.
Hans-SFT-4K is the 3,834-sample SFT mix that teaches models to… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-sft-4k.odyssey-v0.5.1-16h-eval-audio-gated
🔒 Odyssey V0.5.1 — 16H Gated Audio Evaluation Repo
This repository contains the gated audio companion to the public Odyssey evaluation preview.
All requests are reviewed manually.
Public Preview Repository
For transcripts, schema preview, and metadata inspection, see:
ODYSSEYAILABS/odyssey-v0.5.1-16h-eval-preview-public
About Odyssey AI Labs
Odyssey AI Labs builds premium African data infrastructure for next-generation AI systems.
meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Dramazy/meow-10k.
