datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MF-Skills🚨 Please request access with your institutional email to get access to the dataset.
MF-Skills Dataset
Project page | Paper | Code
Dataset Description
MF-Skills is a large-scale dataset for advancing expert-level music understanding and reasoning in (large) audio-language models. It builds upon audio samples from LAION-DISCO and augments them with rich metadata extracted using a suite of open-source large audio-language models (LALMs) and specialized music analysis… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/MF-Skills.av-skillsAV-Skills
Audio-visual instruction and temporally grounded reasoning data for Nemotron-Labs-Audio-Visual Flamingo
AV-Skills supports joint understanding of video, speech, sound, music, and long-range temporal context in real-world videos.
Dataset Summary
AV-Skills is the audio-visual instruction and reasoning dataset for
Nemotron-Labs-Audio-Visual Flamingo, an open audio-visual language model for long and
complex… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/av-skills.av-skillsAV-Skills
Audio-visual instruction and temporally grounded reasoning data for Nemotron-Labs-Audio-Visual Flamingo
AV-Skills supports joint understanding of video, speech, sound, music, and long-range temporal context in real-world videos.
Dataset Summary
AV-Skills is the audio-visual instruction and reasoning dataset for
Nemotron-Labs-Audio-Visual Flamingo, an open audio-visual language model for long and
complex… See the full description on the dataset page: https://huggingface.co/datasets/ViralNow/av-skills.af_skills
