datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WavCaps
WavCaps
WavCaps is a ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research, where the audio clips are sourced from three websites (FreeSound, BBC Sound Effects, and SoundBible) and a sound event detection dataset (AudioSet Strongly-labelled Subset).
Paper: https://arxiv.org/abs/2303.17395
Github: https://github.com/XinhaoMei/WavCaps
Statistics
Data Source
# audio
avg. audio duration (s)avg. text length
FreeSound… See the full description on the dataset page: https://huggingface.co/datasets/cvssp/WavCaps.CVSSShareGPT75CVSSShareGPT50Kbased on https://huggingface.co/datasets/Chaphie/CVSS
there is also a 170k dataset version here: https://huggingface.co/datasets/nbcv12/CVSSShareGPT
CVSSShareGPT100CVSSShareGPTbased on https://huggingface.co/datasets/Chaphie/CVSS
CVSSShareGPT0CVSSAlpacabased on https://huggingface.co/datasets/Chaphie/CVSS
CVSSShareGPT25CVSSShareGPT50CVSSClass
