datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Zeroshot-Audio-Classification-Instructions
Zeroshot-Audio-Classification-Instructions
Convert audio classification dataset into zero-shot format speech instructions, support both single label and multi-label,
VGGSound
FSD50k
Nonspeech7k
urbansound8K
VocalSound
Emotion
Gender
ESD Emotion
Age
Language
TAU Urban Acoustic Scenes 2022
CochlScene
BirdCLEF_2021
EmoBox
AudioSet
We also converted huge WAV files into MP3 16k sample rate to reduce storage size.To prevent leakage, please do not include test set in training session.… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Zeroshot-Audio-Classification-Instructions.audio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.audio_classification_dataset2waste-classification-audio-helsinki
Dataset Card for "waste-classification-audio"
english to italian translation was made with helsinki-NLP translation model.
More Information needed
waste-classification-audio-deepl2audio created from dataset "italian-dataset-deepl2"
waste-classification-audio-deepl-largeconcatentation of "thomasavare/waste-classification-audio-deepl" and "thomasavare/waste-classification-audio-deepl2", so its easier to use.
there are 3 duplicates, idk why but it's too long to remake the datasets (maybe later) and not necessary.
waste-classification-audio-helsinki2
Dataset Card for "waste-classification-audio-helsinki2"
More Information needed
audio_classification_kenya_dataset
Dataset Card for "audio_classification_kenya_dataset"
More Information needed
audio_classification_datasetwaste-classification-audio-deepl
Dataset Card for "waste-classification-audio-deepl"
More Information needed
waste-classification-audio-deepl-v3audio_classificationaudio-classification-checkpoint-downloadswaste-classification-audio-unseen-gptaudio_classification
