datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-speech-commands-15lang
Multilingual Speech Commands Dataset (15 Languages, Augmented)
This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification.
Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.voicehub-arena-seed-tts-eval
VoiceHub Arena — native TTS evaluations
Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA
speaker SIM and UTMOS22 measurements. The full campaign is still running.
Each generation method is evaluated separately using its publisher's native API.
Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target
diagnostic pilots are stored separately and must not be treated as full scores.
Interactive demo ·
Source code
Layout… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/voicehub-arena-seed-tts-eval.symile-m3
Dataset Card for Symile-M3
Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice.
Paper: https://arxiv.org/abs/2411.01053
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.VaaniVAANI is an India-representative multi-modal multi-lingual dataset.
The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages.
From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts.
Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.big_bench_audio
Artificial Analysis Big Bench Audio
Dataset Summary
Big Bench Audio is an audio version of a subset of Big Bench Hard questions. The dataset can be used for evaluating the reasoning capabilities of models that support audio input.
The dataset includes 1000 audio recordings for all questions from the following Big Bench Hard categories. Descriptions are taken from Suzgun et al. (2022):
Formal Fallacies Syllogisms Negation (Formal Fallacies) - 250 questions
Given a context… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio.ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.Earnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.music-arena-dataset
Music Arena Dataset
This is the official dataset from Music Arena, an open platform for evaluating text-to-music (TTM) models.
How to Download (Recommended Method)
The most reliable way to get a complete local copy of all files, including the entire audio collection, is to clone the repository directly using Git. This method is ideal for offline access and workflows that require direct file manipulation.
Note: This repository uses Git LFS (Large File Storage) to… See the full description on the dataset page: https://huggingface.co/datasets/music-arena/music-arena-dataset.Barkopedia_Dog_Sex_Classification_Dataset
📦 Dataset Description
This dataset is part of the Barkopedia Challenge: https://uta-acl2.github.io/barkopedia.html
Check training data on Hugging Face:
👉 ArlingtonCL2/Barkopedia_Dog_Sex_Classification_Dataset
This challenge provides a dataset of labeled dog bark audio clips:
29,345 total clips of vocalizations from 156 individual dogs across 5 breeds:
Shiba Inu
Husky
Chihuahua
German Shepherd
Pitbull
Training set: 26,895 clips
13,567 female13,328 male
Test set: 2,450… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_Dog_Sex_Classification_Dataset.Vaani-Noise-Event-Dataset
Vaani Noise Event Timestamps
Dataset Summary
Vaani Noise Event Timestamps is a derived dataset from Project Vaani, a large-scale multilingual speech initiative by IISc Bangalore and ARTPARK that captures India's linguistic diversity across all districts.
This dataset provides noise event annotations with fine-grained timestamps for the subset audio recordings from the Vaani corpus. Each entry identifies background noise categories along with their precise start… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Noise-Event-Dataset.dialectal-arabic-lahgtna-v2
Dialectal Arabic Lahgtna v2
Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI.
Dataset Summary
~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech
**13 Arabic dialects **, labeled per utterance
16 kHz mono audio
Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.central-kurdish-pseudolabel
Central Kurdish → English Pseudo-Labeled Speech Translation Corpus
Dataset Summary
This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish).
The dataset was automatically generated using a pipeline composed of:
Speech segmentation
Automatic Speech Recognition (ASR)
Machine Translation (MT)
The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.multilingual-speech-commands-3lang-raw
Multilingual Speech Commands Dataset (3 Languages, Raw)
This dataset is a curated subset of previously published speech command datasets in Kazakh, Tatar, and Russian. It is intended for use in multilingual speech command recognition and keyword spotting tasks. No data augmentation has been applied.
All files are included in their original form as released in the cited works below. This repository simply reorganizes them for convenience and accessibility.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-3lang-raw.free-music-archive-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset
Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset Description
This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets.
Training set: 17888 audio clips.
Test set: 4920 audio clips, further divided into:
Test Public (~40%): 1966 audio clips for live leaderboard updates.
Test Private (~60%): 2954 audio clips for final evaluation.
You will… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.voicehub-arena-seed-tts-eval
VoiceHub Arena — full English Seed-TTS-Eval
35,904 synthesized WAV files: 33 model families × the same 1,088 target texts.
The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB.
All 198 shards and every WAV SHA256 were verified after backup.
Interactive leaderboard and all audio samples
· Source repository (access required).
Contents
audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.MASC-Arabic
MASC Arabic Dataset Card
Dataset Summary
MASC is a dataset that contains 1,000 hours of speech sampled at 16 kHz and crawled from over 700 YouTube channels.
The dataset is multi-regional, multi-genre, and multi-dialect intended to advance the research and development of Arabic speech technology with a special emphasis on Arabic speech recognition.
How to use
The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MASC-Arabic.ImageEval-ArabicNLP26
ImageEval-ArabicNLP26 👁️
ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026.
It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation.
The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.dialectal-arabic-voices
Dialectal Arabic Voices
An expanding collection of Arabic audio from YouTube, SoundCloud, and other sources. Currently labelled Palestinian Arabic (ps).
60,775 recordings · approximately 10,895.2 hours · 571.24 GB
Column
Description
audio
Original audio, embedded in the Parquet file
transcript_text
Empty; ASR transcripts are stored in a separate private dataset
language
Dialect code: ps (Palestinian)
source
Original channel or account name
Audio retains its… See the full description on the dataset page: https://huggingface.co/datasets/moaead/dialectal-arabic-voices.arc-voicesamples-generatedartifactbench
ArtifactBench — AI-Generated Music Detection Benchmark
ArtifactBench v2 — lineage-aware frozen protocol
ArtifactBench v2 adds a metadata-first, lineage-aware evaluation protocol while
preserving the v1 and v1.1 releases below. Its frozen primary cohort contains
828 entries (605 AI-generated and 223 real) across 15 source strata, split by
content lineage into calibration, validation, and sealed-test partitions before
final model comparison.
The v2 package is under… See the full description on the dataset page: https://huggingface.co/datasets/intrect/artifactbench.arc-speeches-refinedarabic-audio-collection-algerian-loubna-stories
Loubna Stories Arabic Speech Dataset
Dataset Summary
The Loubna Stories Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 237 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-loubna-stories.free-music-archive-medium
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-medium.Arabic-Quran-ASR-datasetlibrispeech_asr_arrowVaani-transcription-partThis dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages.
This table represents the audio and transcription duration data for various languages.
Language
Angami
Angika
Ao
Assamese
Awadhi
Bajjika
Bearybashe
Bengali
Bhili
Bhojpuri
Bundeli
Chakhesang
Chakma
Chhattisgarhi
English
Garhwali
Garo
Gondi
Gujarati
Halbi
Haryanvi
Hindi
IduMishmi
Kannada
Kashmiri
Karbi
Khariboli
Khortha
Kokborok
Konkani… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.linto-dataset-audio-ar-tn
LinTO DataSet Audio for Arabic Tunisian A collection of Tunisian dialect audio and its annotations for STT task
This is the first packaged version of the datasets used to train the Linto Tunisian dialect with code-switching STT
(linagora/linto-asr-ar-tn).
Dataset Summary
Dataset composition
Sources
Data Table
Data sources
Content Types
Languages and Dialects
Example use (python)
License
Citations
Dataset Summary
The LinTO DataSet Audio for Arabic Tunisian is a diverse… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn.ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.ADI17
