datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AudioSkills
AudioSkills-XL Dataset
Project page | Paper | Code
Dataset Description
AudioSkills-XL is a large-scale audio question-answering (AQA) dataset designed to develop (large) audio-language models on expert-level reasoning and problem-solving tasks over short audio clips (≤30 seconds). It expands upon the original AudioSkills collection by adding approximately 4.5 million new QA pairs, resulting in a total of ~10 million diverse examples. The release includes the full dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/AudioSkills.Nemotron-Content-Safety-Audio-Dataset
Nemotron Content Safety Audio Dataset
Dataset Description
The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories.
LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.Numb3rs
Numb3rs - Numbers Speech Benchmark (Dataset)
A speech dataset for text normalization (TN) and inverse text normalization (ITN) tasks, containing paired written/spoken forms with corresponding synthetic audio.
Dataset Creation
This dataset was created through the following pipeline:
Source Data: Text normalization pairs were derived from the Google Text Normalization dataset, containing written forms (e.g., "$100") and their spoken equivalents (e.g., "one hundred… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Numb3rs.MMOU
MMOU
Massive Multi-Task Omni Understanding and Reasoning
Benchmark for Long and Complex Real-World Videos
Project Page
·
HuggingFace
·
Videos (Community Hosted)
·
Paper
·
Evaluator
MMOU evaluates joint audio-visual understanding and reasoning in long and complex real-world videos.
Dataset Summary
MMOU is a benchmark for evaluating whether multimodal models can jointly reason over video, speech, sound, music, and long-range temporal context in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/MMOU.MF-Skills🚨 Please request access with your institutional email to get access to the dataset.
MF-Skills Dataset
Project page | Paper | Code
Dataset Description
MF-Skills is a large-scale dataset for advancing expert-level music understanding and reasoning in (large) audio-language models. It builds upon audio samples from LAION-DISCO and augments them with rich metadata extracted using a suite of open-source large audio-language models (LALMs) and specialized music analysis… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/MF-Skills.video-full-duplex-benchmark
VideoFDB: Video-Full-Duplex-Benchmark
Project Page · HuggingFace · Paper (arXiv)
Dataset Description
A benchmark dataset of annotated, two-person video conference recordings designed to support the evaluation of multimodal AI agents in conversational settings. The dataset covers 11 distinct conversational dynamics — spanning verbal, nonverbal, and mixed-modality behavior — annotated through a three-pass human-in-the-loop pipeline.
The benchmark consists of trimmed… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/video-full-duplex-benchmark.av-skillsAV-Skills
Audio-visual instruction and temporally grounded reasoning data for Nemotron-Labs-Audio-Visual Flamingo
AV-Skills supports joint understanding of video, speech, sound, music, and long-range temporal context in real-world videos.
Dataset Summary
AV-Skills is the audio-visual instruction and reasoning dataset for
Nemotron-Labs-Audio-Visual Flamingo, an open audio-visual language model for long and
complex… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/av-skills.Audio2Face-3D-Dataset-v1.0.0-claire
Dataset Description:
NVIDIA Audio2Face-3D-dataset-v1.0.0-claire includes audio files, blendshape data, animated geometry caches, geometry files, and transform files.
This dataset is for demonstration purposes and not for production usage.
For source code, documentation, helper scripts, packaged builds, and links to all components in the Audio2Face-3D technology stack, visit the Audio2Face-3D GitHub repository
Dataset Owner:
NVIDIA Corporation
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Audio2Face-3D-Dataset-v1.0.0-claire.nvidia-brain-noise-evaluation-dataset
Nvidia Brain Noise Evaluation Dataset
Dataset Description
This dataset contains 64 samples organized across multiple splits and 32 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
noisy-bg-snr-10: 2 samples
test: 2 samples
noisy-bg-snr-20: 2 samples
test: 2 samples
noisy-bg-snr-30: 2 samples
test: 2 samples
noisy-bg-snr-40: 2 samples
test: 2 samples
noisy-bg-snr-50: 2 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/nvidia-brain-noise-evaluation-dataset.
