datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JALMBench
About the Dataset
📦 JALMBench contains 245,355 audio samples and 11,316 text prompts to benchmark jailbreak attacks against audio-language models (ALMs). It consists of three main categories:
🔥 Harmful Query Category:Includes 246 harmful text queries ($T_{Harm}$), their corresponding audio ($A_{Harm}$), and a diverse audio set ($A_{Div}$) with 9 languages, 2 genders, 3 accents, and 3 TTS methods.
📒 Text-Transferred Jailbreak Category:Features adversarial texts generated by… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousUser000/JALMBench.MoM-CLAM-dataset
Melody or Machine: Benchmarking and Detecting Synthetic Music via Cross-Modal Contrastive Alignment
To support robust detection of AI-generated songs under diverse manipulations and model types, we present the Melody or Machine (MoM) Dataset—a large-scale benchmark reflecting the evolving landscape of song-level deepfakes. MoM spans three authenticity tiers: genuine recordings, synthetic audio with real lyrics/voice, and fully synthetic tracks, enabling evaluation across progressive… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2212/MoM-CLAM-dataset.vocalgrad
VocalGrad
VocalGrad is an audio benchmark for evaluating whether a model can detect the
direction of gradual perceptual change in speech. This public release contains
the test split only.
Each example contains one audio clip and one target attribute. The task is to
answer whether that attribute increases or decreases over time.
Task
Given an audio clip and an attribute name, predict one of two labels:
increase
decrease
The ground-truth label is derived from the metadata… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-user-592888/vocalgrad.Sympatheia-18k
Sympatheia-18k
Sympatheia-18k is an emotion-aware spoken dialogue dataset for empathetic speech synthesis research.
It contains 18,000 query–response pairs across 12 emotion categories, each accompanied by
synthesized audio and text transcripts.
Dataset Structure
Subset
Unique Queries
Responses
Description
Emotional
8,400 train / 3,600 eval
8,400 train / 3,600 eval
Emotional queries with emotionally-matched responses
Neutral
350 train / 150 eval
4,200 train /… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2222/Sympatheia-18k.anonymous-storybench
Omni-StoryBench
Omni-StoryBench is a context-aware omnimodal story generation benchmark.Each sample provides a current story page and requires generating the next page's image, narration text, and speech utterance.
Dataset Structure
The dataset contains:
data/testset.jsonl: Main benchmark file.
images/: Page images.
texts/: Page text files.
speech/: Generated speech audio files.
instruction/: Source-level instruction metadata.
Data Fields
Each JSONL sample… See the full description on the dataset page: https://huggingface.co/datasets/omnibench/anonymous-storybench.Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.CaReCoS
CaReCoS
A medical acoustic question-answering dataset for reasoning over mel spectrograms
of heart, lung, and cough sounds. Each record provides a clinical question, the
mel-spectrogram image of a recording, a ground-truth answer, and the
recording's clinical metadata.
The task is purely visual: a model receives the spectrogram image together with the
question and must reason over the spectrogram to produce the answer. The raw audio is
not used as model input - the original .wav… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-dataset-1/CaReCoS.mellow-podcast-datasoundsense-webtoon
Webtoon Sound-Moment Annotations
Human-curated annotations of where a vertical-scroll webtoon should make a sound, with a
reference audio clip attached to each moment. Built to evaluate whether vision-language models can
judge, from a static drawing alone, that a depicted moment is audible. Companion resource to the
paper "SoundSense: Visual Sound Grounding in Comics and Webtoons".
Webtoons have no prior sound-effect annotation, and unlike manga onomatopoeia these labels are… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-webtoon-sfx/soundsense-webtoon.News
AnonymousContinuousBench — News
A news-grounded QA benchmark built from Common Crawl News (CC-NEWS) articles
crawled in September 2025. QAs are generated by Gemini 2.5 from clusters
of related articles, then filtered for answerability and grounded with a
retrieval-based set of supporting articles drawn from the corpus.
What's inside
Config
Splits
Size
What it's for
qa (default)
val (1,189), test (1,415)
233 MB
Evaluate QA on news, post-event
corpus_large… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousContinuousBench/News.RVCBench
Benchmark Dataset
This is the benchmark dataset used for studying robustness in voice cloning and related audio generation pipelines.
Set RVCBENCH_HF_DATASET_ID if you want the codebase to fetch a remote dataset release from Hugging Face.
Each subset is exposed as its own Hugging Face dataset configuration. Most subsets contain:
metadata.parquet
audios/
The canonical metadata stores one row per benchmark pair with columns such as:
speaker_id
prompt_file_name
prompt_text… See the full description on the dataset page: https://huggingface.co/datasets/anonymous65432184/RVCBench.ETbird3m
Bird3M Dataset
Dataset Description
Bird3M is the first synchronized, multi-modal, multi-individual dataset designed for comprehensive behavioral analysis of freely interacting birds, specifically zebra finches, in naturalistic settings. It addresses the critical need for benchmark datasets that integrate precisely synchronized multi-modal recordings to support tasks such as 3D pose estimation, multi-animal tracking, sound source localization, and vocalization attribution.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission000/bird3m.MuseBench-part1
MuseBench: Sample Dataset for Distinguishing Human vs. AI-Generated Music
This repository contains a sample preview of the MuseBench dataset used for benchmarking human-vs-AI music classification. Each folder includes a single representative file so users can inspect the layout before downloading the full release. The complete dataset is hosted on Hugging Face under Anonymousv22222, split across MuseBench-part1 … MuseBench-part5 (see the Download section below).… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousv22222/MuseBench-part1.same-text-different-speechWorldGUI-BenchMMAG-Benchmark
MMAG: A Multi‑Control Mixed Audio Generation Benchmark
MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment.
Dataset Structure
The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2026082026/MMAG-Benchmark.ViVoice34
ViVoice-34: Vietnamese Speech Dataset
Dataset Description
ViVoice-34 is a Vietnamese speech dataset featuring recordings from speakers across provinces of Vietnam. This repository contains a preview subset with playable WAV samples and metadata.
The preview data is stored as Parquet with the audio column encoded as Hugging Face Audio, so Dataset Viewer and Data Studio can render an audio player instead of plain file paths.
Data Fields
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-vivoice34/ViVoice34.SEABED
SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning
SEABED (SouthEast Asian Benchmark for Evaluating Audio Reasoning) covers
six audio-reasoning tasks: speech emotion recognition, speech affective
interpretation, dialect and language identification, dialectal speech
comprehension, prosodic ambiguity resolution, and long-form audio
reasoning. This repository releases a stratified 10% sample (541 of 5,404
records) of its QA data for anonymous peer review, so reviewers… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-1/SEABED.MuseBench-part3MuseBench-Examplestar1_jailbreakRepresentative_Samples
Representative Samples
This repository contains representative Vietnamese speech samples organized by speaker folders.
Dataset structure
Representative_Samples/
train/
spk1/
*.wav
spk2/
*.wav
val/
spk1/
*.wav
spk2/
*.wav
Fields
audio: audio waveform file.
label: speaker folder name, for example spk1, spk2, etc.
Splits
train: training representative samples.
validation: validation representative samples.
Intended use… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-vivoice34/Representative_Samples.audio-flamingo-3-hf_advwave_50000_500merged_qwen_audio_dataset_subsetVarta-DF
Varta-DF: A Dataset for Partial Audio Deepfake Localization (Sample)
🚨 NeurIPS 2026 Double-Blind Review Notice 🚨
This dataset is currently hosted on an anonymous account to strictly comply with the double-blind review policies of the Datasets and Benchmarks track. Upon acceptance, the repository will be transferred to the official laboratory organization account.
Overview
This repository contains the < 4GB representative sample of the Varta-DF dataset, provided as… See the full description on the dataset page: https://huggingface.co/datasets/anonymous19submission/Varta-DF.advbenchaudio-flamingo-3-hf_advwave_100000_500audio-flamingo-3-hf_noise2_10_0.5_advwave_50000_500star1_audio
