datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cosyvoice-instructlalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.cosyvoice-clone
LALM Emotional Vulnerability Dataset
Overview
This dataset contains synthesized malicious speech instructions across multiple emotions and intensity levels to evaluate the safety responsiveness of Large Audio-Language Models (LALMs). The dataset aims to examine how speaker emotion and intensity influence the safety and robustness of AI responses.
Dataset Composition
Total samples: 8,320
Emotion categories:
Neutral: 520 samples
Angry: 1560 samples… See the full description on the dataset page: https://huggingface.co/datasets/LALM-emotional-vulnerability/cosyvoice-clone.cosyvoice-clonetemporal-lalm
Temporal LALM: Relative Temporal Audio MCQA
Multiple-choice questions probing relative temporal reasoning over audio:
identifying which sound event starts earliest, ends latest, or has the longest
duration within a clip. Built on the TACOS
audio collection.
Tasks
task
question
#MCQs
earliest_start
Which sound event starts earliest?
528
latest_end
Which sound event ends latest?
499
longest_duration
Which sound event has the longest duration?
630… See the full description on the dataset page: https://huggingface.co/datasets/gamma-lab-umd/temporal-lalm.cosyvoice-instructIFAO-lalmgpt-4o-mini-ttsSAKE
SAKE: A Benchmark for Editing Auditory Perceptual Knowledge in Large Audio-Language Models
This is the dataset for the SAKE benchmark, the first benchmark for auditory perceptual knowledge editing in LALMs.
The dataset is organized into four subsets: train, val, single-test, and sequential-test, which refer to the training set, validation set, single editing test set, and sequential editing test set, respectively. Each subset is stored in a separate config within this Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/sake-lalm/SAKE.bouba-kiki-lalm
Bouba-Kiki LALM Evaluation
Stimuli, prompts, and human reference data for evaluating the bouba-kiki effect in
large audio-language models. Replicates Ćwiek et al. (2022) and McCormick (2015) / Lacey (2020).
config
what
modality
exp1_stimuli / exp1_trials / exp1_human_reference
Ćwiek bouba/kiki 2AFC, 26 prompt languages, with human GT
audio→text
exp2_stimuli / exp2_trials
537 McCormick pseudowords rated rounded/pointed (1–7)
audio→text
exp3_stimuli / exp3_trials… See the full description on the dataset page: https://huggingface.co/datasets/dlion168/bouba-kiki-lalm.
