datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SoccerHigh
⚽ SoccerHigh
This dataset provides annotations and pre-extracted features for the SoccerHigh benchmark introduced in:
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
Artur Díaz-Juan, Coloma Ballester, Gloria HaroACM MMSports 2025
📦 Contents
Highlight summary annotations
Train / validation / test splits
Pre-extracted visual features (no raw videos)
All data are provided as .npy feature arrays, .srt temporal annotations, and .json metadata files.… See the full description on the dataset page: https://huggingface.co/datasets/imva-upf/SoccerHigh.upfall-detection-actualallmusiccaps
AllMusicCaps
Music caption dataset built from professional AllMusic album reviews, cross-referenced with Discogs
releases and YouTube tracks. Captions are generated in two complementary styles by a two-stage LLM
pipeline.
Released with the ISMIR 2026 paper AllMusicCaps: Album Reviews as Complementary Supervision for
Music CLAP. Code and models: github.com/MTG/allmusiccaps.
This dataset contains no audio, only identifiers and captions. Recover audio from the
youtube_id field, or… See the full description on the dataset page: https://huggingface.co/datasets/mtg-upf/allmusiccaps.HumMusQA
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
Authors: Benno Weck, Pablo Puentes, Andrea Poltronieri, Satyajeet Prabhu, Dmitry Bogdanov
HumMusQA is a multiple-choice question answering dataset designed to test music understanding in Large Audio-Language Models (LALMs).
Dataset Highlights
✍️ 320 hand-written multiple-choice questions curated and validated by experts with musical training
🎵 108 Creative Commons-licensed music tracks sourced from… See the full description on the dataset page: https://huggingface.co/datasets/mtg-upf/HumMusQA.upfall-detection-alpacaUpFFSmoothLabelsR2RLoraupf_dataset
