CoolFace
Datasetpublicgated

HKUSTAudio/AudioX-IFcaps

[ICLR 2026] AudioX-IFcaps: Instruction-Following Audio Caption Dataset AudioX-IFcaps (Instruction-Following) is a large-scale, high-quality multimodal dataset designed for training unified audio and music generation models. The dataset contains over 7 million samples with fine-grained, structured annotations that enable precise control over audio generation, including sound event categories, counts, temporal ordering, and timestamps. 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/AudioX-IFcaps.

sourceHugging Facecc-by-nc-nd-4.0updated 6mo agoView on Hugging Face
7likes11downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.