CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MBZUAI /AudioJailbreak Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly. 📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.audioquestion-answering1K<n<10K9 likes2.4k downloads1y agoHugging Face02m-a-p /GSaudio1K<n<10K1 likes1.7k downloads1y agoHugging Face03MrSupW /ContextASR-Bench ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.textautomatic-speech-recognition10K<n<100K38 likes1.6k downloads1y agoHugging Face04smgjch /meow-10k Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology. Dataset Summary Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.audio10K<n<100K3 likes1.5k downloads5mo agoHugging Face05pollen-robotics /microduck-emotions Microduck Emotions A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.audioroboticsn<1K6 likes937 downloads17d agoHugging Face06McGill-NLP /speech-translation-and-summarization English-Centric Multilingual Audio Dataset This dataset contains generated article and summary audio for English-centric multilingual directions. Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits. Included directions amharic_english / english_amharic arabic_english / english_arabic bengali_english / english_bengali chinese_simplified_english / english_chinese_simplified english_english french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.audioautomatic-speech-recognition10K<n<100K6 likes760 downloads1mo agoHugging Face07medkit /simsamu Simsamu dataset This repository contains recordings of simulated medical dispatch dialogs in the french language, annotated for diarization and transcription. It is published under the MIT license. These dialogs were recorded as part of the training of emergency medicine interns, which consisted in simulating a medical dispatch call where the interns took turns playing the caller and the regulating doctor. Each situation was decided randomly in advance, blind to who was playing the… See the full description on the dataset page: https://huggingface.co/datasets/medkit/simsamu.audioautomatic-speech-recognitionn<1K8 likes662 downloads11mo agoHugging Face08ASLP-lab /MSU-Benchmark MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.audioaudio-classification1K<n<10K1 likes622 downloads3mo agoHugging Face09KZL96 /ModalityFaultLines-SCEval SCEval — Modality Fault Lines Data for Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning (Findings of EMNLP 2026). SCEval is a human-verified benchmark for omni-modal robustness. Text, vision, and audio all remain present, but controlled corruptions make the evidence inside a channel unreliable. Each corrupted item is paired with its clean counterpart at the example level, so clean-to-corrupted comparisons are made on the same underlying question… See the full description on the dataset page: https://huggingface.co/datasets/KZL96/ModalityFaultLines-SCEval.audiomultiple-choice10K<n<100K0 likes479 downloads28d agoHugging Face10nvidia /MMOU MMOU Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos Project Page · HuggingFace · Videos (Community Hosted) · Paper · Evaluator MMOU evaluates joint audio-visual understanding and reasoning in long and complex real-world videos. Dataset Summary MMOU is a benchmark for evaluating whether multimodal models can jointly reason over video, speech, sound, music, and long-range temporal context in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/MMOU.textvideo-text-to-text10K<n<100K20 likes405 downloads2mo agoHugging Face11rookie9 /MMAG MMAG: A Multi‑Control Mixed Audio Generation Benchmark MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment. Dataset Structure The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/rookie9/MMAG.audiotext-to-audio1K<n<10K0 likes365 downloads2mo agoHugging Face12taohu /music2chords_v2audio1K<n<10K0 likes362 downloads7mo agoHugging Face13zed-m97 /nano4m-Audio nano4M-Audio — Team (week-1) Week-1 data preparation for nano4M-Audio, an extension of EPFL's nano4M (the educational nano version of 4M / 4M-21) that adds audio as a fifth modality alongside RGB, depth, surface normals and captions. This dataset covers all 12 VGGSound classes assigned to the three-person team: person classes 1 (Hassan) lions roaring, horse neighing, pig oinking, cow lowing 2 (Ziyad) dog barking, cat meowing, coyote howling, elephant trumpeting 3… See the full description on the dataset page: https://huggingface.co/datasets/zed-m97/nano4m-Audio.textaudio-classification1K<n<10K0 likes334 downloads5mo agoHugging Face14nyuuzyou /OpenGameArt-Mixed-Licenses Dataset Card for OpenGameArt-Mixed-Licenses Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are available under multiple licenses simultaneously. This dataset includes assets where creators have made their work available under two or more license options. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata, all… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-Mixed-Licenses.audioimage-classification1K<n<10K0 likes302 downloads1y agoHugging Face15MML-Group /AVE-Speechgated AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals Abstract AVE Speech is a large-scale Mandarin speech corpus that pairs synchronized audio, lip video and surface electromyography (EMG) recordings. The dataset contains 100 sentences read by 100 native speakers. Each participant repeated the full corpus ten times, yielding over 55 hours of data per modality. These complementary signals enable… See the full description on the dataset page: https://huggingface.co/datasets/MML-Group/AVE-Speech.audion<1K7 likes296 downloads1y agoHugging Face16nvidia /MF-Skills🚨 Please request access with your institutional email to get access to the dataset. MF-Skills Dataset Project page | Paper | Code Dataset Description MF-Skills is a large-scale dataset for advancing expert-level music understanding and reasoning in (large) audio-language models. It builds upon audio samples from LAION-DISCO and augments them with rich metadata extracted using a suite of open-source large audio-language models (LALMs) and specialized music analysis… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/MF-Skills.audioaudio-text-to-text1M<n<10M9 likes296 downloads6mo agoHugging Face17ground-truth /multichannel-meetings-10h GroundTruth Multi-Channel Meeting Audio Dataset (10h) Summary This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant. Each meeting includes: One full meeting recording (room microphone) Individual close-talk recordings for each participant (one file per speaker) Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.audioautomatic-speech-recognitionn<1K1 likes287 downloads5mo agoHugging Face18MixEval /MixEval-X 🚀 Project Page | 📜 arXiv | 👨‍💻 Github | 🏆 Leaderboard | 📝 blog | 🤗 HF Paper | 𝕏 Twitter MixEval-X encompasses eight input-output modality combinations and can be further extended. Its data points reflect real-world task distributions. The last grid presents the scores of frontier organizations’ flagship models on MixEval-X, normalized to a 0-100 scale, with MMG tasks using win rates instead of Elo. Section C of the paper presents example data samples and model responses.… See the full description on the dataset page: https://huggingface.co/datasets/MixEval/MixEval-X.audioimage-to-text1K<n<10K10 likes249 downloads2y agoHugging Face19MrJackTung /cs-envi-dual-encoder-60audion<1K0 likes237 downloads4mo agoHugging Face20lilonghao /MM-ContextASR-Bench MM-ContextASR Bench Metadata and evaluation splits for Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark. Dataset summary Config Examples Audio Context Primary metric mm_contextasr 1,250 (250 current utterances × 5 histories) 1,439 WAV files included Controlled user-assistant dialogue entity Recall kespeech 19,212 Source ID only Same-speaker speech and transcript CER, SER, entity Recall cv_yue 3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.audioautomatic-speech-recognition10K<n<100K1 likes233 downloads8d agoHugging Face21m-a-p /EMOaudion<1K1 likes214 downloads1y agoHugging Face22mcp-tool-shop /jam-actions-v1 jam-actions-v1 Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) · Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE · Source repo: mcp-tool-shop-org/ai-jam-sessions The successor to jam-actions-v0. Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from what the tools return — and it exists in its current shape because, seven training runs in a row, the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.texttext-generationn<1K0 likes188 downloads15d agoHugging Face23MLSpeech /hillenbrand_vowels Hillenbrand Vowel Dataset The Hillenbrand dataset contains recordings of American English vowels produced by speakers from four demographic groups: men, women, boys, and girls.Each audio sample is accompanied by frame-level formant tracks (F1, F2, F3, F4) extracted every 10 ms. This repository provides the dataset in a structure compatible with the Hugging Face datasets library for easy loading and processing. Dataset Summary Sampling rate: 16 kHz Frame shift: 10… See the full description on the dataset page: https://huggingface.co/datasets/MLSpeech/hillenbrand_vowels.audio1K<n<10K0 likes175 downloads10mo agoHugging Face24ICTNLP /MutiEmo-Test MultiEmo-Test MultiEmo-Test is an English evaluation set for instruction-following multi-emotion text-to-speech synthesis. It accompanies HybridEmo, a system for modeling sequential emotion trajectories and simultaneous emotion blending within an utterance. The dataset is intended for evaluation only. It contains synthesis text, natural-language emotion instructions, emotion annotations, and prompt audio for speaker-timbre conditioning. It does not contain target synthesized… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/MutiEmo-Test.audion<1K1 likes174 downloads24d agoHugging Face25zsy814 /instructtts-three-model-gemini-zh InstructTTSEval 三模型 Gemini 评测数据 本目录整理了 InstructTTSEval 中文集上三个 TTS 模型的生成音频和 Gemini 一致性评测结果:Qwen3-TTS-12Hz-1.7B-VoiceDesign、Seed-Audio-1.0、VoxCPM2。 字段 records.jsonl 每行对应一个模型和一种控制格式(APS、DSD 或 RP): id:InstructTTSEval 样本 ID mode:控制格式 model、model_name:模型标识 text:合成文本 instruction:历史评测记录中的输入控制指令,按本行 APS/DSD/RP 格式保留;不等同于各模型 API 的完整请求封装 generated_audio:该模型生成音频的相对路径 reference_audio:原始参考音频的相对路径 gemini_consistent:Gemini judge 的一致性判断 inconsistency_reason:判断为不一致时的原因… See the full description on the dataset page: https://huggingface.co/datasets/zsy814/instructtts-three-model-gemini-zh.audio1K<n<10K0 likes163 downloads8d agoHugging Face26mcp-tool-shop /jam-actions-v1-probe jam-actions-v1-probe Schema: jam-actions-v1-probe/1.0.0 · Records: 24, all split: test · Evaluation only · Companion to: jam-actions-v1 Why it exists An adapter trained on an earlier version of the corpus scored 47/54 on held-out acoustic takes. Its completions, which state the comparison before the label, showed that it wrote against a 50-cent gate whenever it saw a minus sign — and negative cents occurred in exactly one class of that corpus. The main split could… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1-probe.textothern<1K0 likes158 downloads14d agoHugging Face27MohamedGomaa30 /EGYSpeak EGYSpeak A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline. Quick Start 1. Download the dataset: from huggingface_hub import snapshot_download snapshot_download( repo_id="MohamedGomaa30/EGYSpeak", repo_type="dataset", local_dir="EGYSpeak", ) 2. Extract the dataset: from… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/EGYSpeak.textautomatic-speech-recognition100K<n<1M1 likes156 downloads5mo agoHugging Face28nekoyama12 /Music Made by herza For APP.t audion<1K0 likes156 downloads4mo agoHugging Face29mcp-tool-shop /jam-actions-acoustic-v0 Dataset Card for jam-actions-acoustic-v0 Version: 1.0.2 Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI. Summary 108 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence). This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.texttext-generationn<1K0 likes153 downloads16d agoHugging Face30rorosese /my-voxtral-datasetaudion<1K0 likes152 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.