datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Barkopedia_Dog_Sex_Classification_Dataset
📦 Dataset Description
This dataset is part of the Barkopedia Challenge: https://uta-acl2.github.io/barkopedia.html
Check training data on Hugging Face:
👉 ArlingtonCL2/Barkopedia_Dog_Sex_Classification_Dataset
This challenge provides a dataset of labeled dog bark audio clips:
29,345 total clips of vocalizations from 156 individual dogs across 5 breeds:
Shiba Inu
Husky
Chihuahua
German Shepherd
Pitbull
Training set: 26,895 clips
13,567 female13,328 male
Test set: 2,450… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_Dog_Sex_Classification_Dataset.Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset
Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset Description
This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets.
Training set: 17888 audio clips.
Test set: 4920 audio clips, further divided into:
Test Public (~40%): 1966 audio clips for live leaderboard updates.
Test Private (~60%): 2954 audio clips for final evaluation.
You will… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.ParlaSpeech-RS
The Serbian Parliamentary Spoken Dataset ParlaSpeech-RS 1.0
The master dataset can be found at http://hdl.handle.net/11356/1834.
Notice: ParlaSpeech corpora are currently in the process of enrichment with new features. Follow our progress here: http://clarinsi.github.io/parlaspeech
The ParlaSpeech-RS dataset is built from the transcripts of parliamentary proceedings available in the Serbian part of the ParlaMint corpus (http://hdl.handle.net/11356/1859), and the parliamentary… See the full description on the dataset page: https://huggingface.co/datasets/classla/ParlaSpeech-RS.ParlaSpeech-CZ
Dataset Card for "ParlaSpeech-CZ.v1.0"
The master dataset can be found at http://hdl.handle.net/11356/1785.
Notice: ParlaSpeech corpora are currently in the process of enrichment with new features. Follow our progress here: http://clarinsi.github.io/parlaspeech
The ParlaSpeech-CZ dataset is built from the transcripts of parliamentary proceedings available in the Czech part of the ParlaMint corpus (http://hdl.handle.net/11356/1859), and the parliamentary recordings available… See the full description on the dataset page: https://huggingface.co/datasets/classla/ParlaSpeech-CZ.ParlaSpeech-HR
The Croatian Parliamentary Spoken Dataset ParlaSpeech-HR 2.0
The master dataset can be found at http://hdl.handle.net/11356/1914.
Notice: ParlaSpeech corpora are currently in the process of enrichment with new features. Follow our progress here: http://clarinsi.github.io/parlaspeech
The ParlaSpeech-HR dataset is built from the transcripts of parliamentary proceedings available in the Croatian part of the ParlaMint corpus (http://hdl.handle.net/11356/1859), and the parliamentary… See the full description on the dataset page: https://huggingface.co/datasets/classla/ParlaSpeech-HR.AfriMCQA-category-classification
Afri-MCQA cross-modal cultural category classification (MTEB)
Classify the cultural category of an entry from its photograph and the question
about it spoken by a native speaker, across 16 African languages.
Labels index this list:
geography, building, and landmarks
public figure and pop culture
cooking and food
objects, materials, clothing
tranditions, art, and history
brands, products, and companies
plants and animals
people, and everyday life
vehicles and transportation… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-category-classification.Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset
Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset Description
This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets.
Training set: 17888 audio clips.
Test set: 4920 audio clips, further divided into:
Test Public (~40%): 1966 audio clips for live leaderboard updates.
Test Private (~60%): 2954 audio clips for final evaluation.
You… See the full description on the dataset page: https://huggingface.co/datasets/hlx1021/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.Vehicle_sounds_classification_datasetBarkopedia_DOG_BREED_CLASSIFICATION_DATASET
📦 Dataset Description
This dataset is part of the Barkopedia Challenge
🔗 https://uta-acl2.github.io/barkopedia.html
Check Training Data here:👉 ArlingtonCL2/Barkopedia_DOG_BREED_CLASSIFICATION_DATASET
This dataset contains 29,347 audio clips of dog barks labeled by dog breed.
The audio samples come from 156 individual dogs across 5 dog breeds:
shiba inu
husky
chihuahua
german shepherd
pitbull
📊 Per-Breed Summary
Breed
Train
Public TestPrivate Test
Test… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_DOG_BREED_CLASSIFICATION_DATASET.Zeroshot-Audio-Classification-Instructions
Zeroshot-Audio-Classification-Instructions
Convert audio classification dataset into zero-shot format speech instructions, support both single label and multi-label,
VGGSound
FSD50k
Nonspeech7k
urbansound8K
VocalSound
Emotion
Gender
ESD Emotion
Age
Language
TAU Urban Acoustic Scenes 2022
CochlScene
BirdCLEF_2021
EmoBox
AudioSet
We also converted huge WAV files into MP3 16k sample rate to reduce storage size.To prevent leakage, please do not include test set in training session.… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Zeroshot-Audio-Classification-Instructions.fma-genre-classification
FMA Genre Classification Dataset
The FMA Genre Classification Dataset is a subset of the Free Music Archive (FMA), containing audio samples and genre labels for music classification tasks. This version uses the "small" subset of FMA, which contains 8,000 tracks of 30 seconds each, evenly distributed across 8 genres.
Dataset Description
Dataset Summary
This dataset consists of 8,000 audio tracks from the Free Music Archive (FMA), each 30 seconds in length… See the full description on the dataset page: https://huggingface.co/datasets/rpmon/fma-genre-classification.wav-classical-musicClassification-Speech-Instructions
Classification Speech Instructions
Speech instructions for emotion, gender, age and language audio classification.
Source code
Source code at https://github.com/mesolitica/malaysian-dataset/tree/master/llm-instruction/speech-classification-instructions
Barkopedia_DOG_BREED_CLASSIFICATION_DATASET
📦 Dataset Description
This dataset is part of the Barkopedia Challenge
🔗 https://uta-acl2.github.io/barkopedia.html
Check Training Data here:👉 ArlingtonCL2/Barkopedia_DOG_BREED_CLASSIFICATION_DATASET
This dataset contains 29,347 audio clips of dog barks labeled by dog breed.
The audio samples come from 156 individual dogs across 5 dog breeds:
shiba inu
husky
chihuahua
german shepherd
pitbull
📊 Per-Breed Summary
Breed
Train
Public Test
Private Test… See the full description on the dataset page: https://huggingface.co/datasets/PuneettArora/Barkopedia_DOG_BREED_CLASSIFICATION_DATASET.Recorrected_Classification_Data_filtered_trainDOA_dataset_6_classes2
Dataset Card for "DOA_dataset_6_classes2"
More Information needed
ParlaSpeech-PL
The Polish Parliamentary Spoken Dataset ParlaSpeech-PL 1.0
The master dataset can be found at http://hdl.handle.net/11356/1686.
Notice: ParlaSpeech corpora are currently in the process of enrichment with new features. Follow our progress here: http://clarinsi.github.io/parlaspeech
The ParlaSpeech-PL dataset is built from the transcripts of parliamentary proceedings available in the Polish part of the ParlaMint corpus (http://hdl.handle.net/11356/1859), and the parliamentary… See the full description on the dataset page: https://huggingface.co/datasets/classla/ParlaSpeech-PL.audio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.CAMEO-emotion-classification
CAMEO multilingual speech emotion classification (MTEB)
Speech emotion recognition across five languages, drawn from the CAMEO collection.
Labels index this list:
anger
fear
happiness
neutral
sadness
surprise
Source: amu-cai/CAMEO at revision 38e9e96, cc-by-nc-sa-4.0. Split by
speaker so no speaker appears in both train and test. Only the six emotions common
to every included language are kept. Audio is 16 kHz Opus.
Built by scripts/data/cameo_emotion/create_data.py in the… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/CAMEO-emotion-classification.Recorrected_Classification_Data_filtered_train_22building_floor_classificationDataset_chunked_5 : chunks of 05 seconds obtained from expert samples
Dataset_chunked_10 : chunks of 10 seconds obtained from expert samples
Dataset_expanded : chunks of 10 seconds obtained from whole samples
Data.zip : original dataset
vocal-burst-classification-v2
Vocal Burst Classification V2 — laion/vocal-burst-classification-v2
The V2 training corpus for the Vocal Burst Classifier V2:
a single-label dataset over an 83-class vocal-burst taxonomy (82 non-speech human vocalizations
no_burst, index 82), shipped as precomputed VoiceCLAP-commercial embeddings plus the raw
vocal-bursts-clean audio.
The vocal-burst clips were generated with various synthetic text-to-audio models such as DramaBox,
then annotated and filtered as described… See the full description on the dataset page: https://huggingface.co/datasets/laion/vocal-burst-classification-v2.Genre_Classificationmaestro-classicalbinary-classifier-birdnet
Binary BirdNet Classifier
Contiene anotaciones y audios de 3s y 5s para clasificación binaria con rutas relativas.
Vehicle_sounds_classification_datasetopen-focus-classical-600
Open Focus and Classical 600
This repository contains two independently usable but analysis-aligned music collections. The default paired configuration loads both groups; the focus and classical configurations load either group independently. Each configuration preserves discovery, validation, and holdout splits.
from datasets import load_dataset
paired = load_dataset("OWNER/open-focus-classical-600", "paired")
focus = load_dataset("OWNER/open-focus-classical-600", "focus")… See the full description on the dataset page: https://huggingface.co/datasets/fisheryv/open-focus-classical-600.GTZAN_genre_classificationViSpeech-Gender-Dialect-Classificationimport datasets as hugDS
import pandas as pd
import os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from df.io import resample
from df.enhance import enhance, init_df
import torch
import warnings
df_model, df_state, _ = init_df()
SAMPLING_RATE = 16_000
def normalize_vietmed(example):
global vietmed_info
example["gender"] = vietmed_info[vietmed_info["Speaker ID"] == example["Speaker ID"]]["Gender"].values[0].lower()
example["dialect"] = vietmed_info[vietmed_info["Speaker ID"] ==… See the full description on the dataset page: https://huggingface.co/datasets/hr16/ViSpeech-Gender-Dialect-Classification.Audio_for_age_classification_Train
