HKUSTAudio/AudioX-IFcaps
[ICLR 2026] AudioX-IFcaps: Instruction-Following Audio Caption Dataset AudioX-IFcaps (Instruction-Following) is a large-scale, high-quality multimodal dataset designed for training unified audio and music generation models. The dataset contains over 7 million samples with fine-grained, structured annotations that enable precise control over audio generation, including sound event categories, counts, temporal ordering, and timestamps. 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/AudioX-IFcaps.
711
No commit history came back for main. The revision may not exist, or the source declined the request.
