datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Zeroshot-Audio-Classification-Instructions
Zeroshot-Audio-Classification-Instructions
Convert audio classification dataset into zero-shot format speech instructions, support both single label and multi-label,
VGGSound
FSD50k
Nonspeech7k
urbansound8K
VocalSound
Emotion
Gender
ESD Emotion
Age
Language
TAU Urban Acoustic Scenes 2022
CochlScene
BirdCLEF_2021
EmoBox
AudioSet
We also converted huge WAV files into MP3 16k sample rate to reduce storage size.To prevent leakage, please do not include test set in training session.… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Zeroshot-Audio-Classification-Instructions.au.1.audio_classification
Dataset Card for "au.1.audio_classification"
More Information needed
audio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.Audio_for_age_classification_EvalAudio_for_age_classification_Traintamil-audio-emotion-classificationaudio_classification_dataset2waste-classification-audio-helsinki
Dataset Card for "waste-classification-audio"
english to italian translation was made with helsinki-NLP translation model.
More Information needed
tamil-audio-emotion-classificationwaste-classification-audio-deepl2audio created from dataset "italian-dataset-deepl2"
waste-classification-audio-helsinki2
Dataset Card for "waste-classification-audio-helsinki2"
More Information needed
waste-classification-audio-deepl-largeconcatentation of "thomasavare/waste-classification-audio-deepl" and "thomasavare/waste-classification-audio-deepl2", so its easier to use.
there are 3 duplicates, idk why but it's too long to remake the datasets (maybe later) and not necessary.
audio_classification_kenya_dataset
Dataset Card for "audio_classification_kenya_dataset"
More Information needed
Mel_Spectrogram_Images_for_Audio_Classification
Mel-Spectrogram Image Dataset (Generated via Custom Pipeline)
> This dataset was fully generated through my notebook
> “Building an Audio Classification Pipeline with DL” available on my profile.
> It represents a complete end-to-end transformation from raw audio to clean, balanced Mel-spectrogram images suitable for deep learning.
Dataset Summary
Property
Description
Number of Classes
13 distinct audio categories
Original Audio per Class
~40 raw recordings… See the full description on the dataset page: https://huggingface.co/datasets/AIOmarRehan/Mel_Spectrogram_Images_for_Audio_Classification.audio_classification_datasetaudio_classificationwaste-classification-audio-deepl
Dataset Card for "waste-classification-audio-deepl"
More Information needed
waste-classification-audio-deepl-v3audio-classification-checkpoint-downloadswaste-classification-audio-unseen-gptaudio-classification-datasetaudio_classificationaudio-scene-classification-dataaudio_classification_modelaudio-scene-classification-data
