yangwang825/audioset
AudioSet AudioSet[1] consists of an expanding ontology of 527 audio event classes and a collection of 2M human-labelled 10-second sound clips drawn from YouTube. Some clips are missing on YouTube, so the number of files downloaded is different from time to time. This repository contains 20550 / 22160 of the balanced train set, 1913637 / 2041789 of the unbalanced train set (separated into 41 parts), and 18887 / 20371 of the evaluation set. The pre-process script can be found at… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/audioset.
AudioSet
AudioSet<sup>[1]</sup> consists of an expanding ontology of 527 audio event classes and a collection of 2M human-labelled 10-second sound clips drawn from YouTube. Some clips are missing on YouTube, so the number of files downloaded is different from time to time. This repository contains 20550 / 22160 of the balanced train set, 1913637 / 2041789 of the unbalanced train set (separated into 41 parts), and 18887 / 20371 of the evaluation set. The pre-process script can be found at qiuqiangkong's github<sup>[2]</sup>.
To improve training efficiency, we add a slightly more balanced subset AudioSet500K<sup>[3]</sup>.
References
- Gemmeke, Jort F., et al., Audio set: An ontology and human-labeled dataset for audio events, 2017
- Kong, Qiuqiang, et al., Panns: Large-scale pretrained audio neural networks for audio pattern recognition, 2020
- Nagrani, Arsha, et al., Attention bottlenecks for multimodal fusion, 2021
