datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LAION-Audio-300Msoundscapeslaions_got_talent
LAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel.
Dataset Composition
The dataset includes:
Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.conceptual-captions-12m-webdatasetlaion-audio-previewlaion-3m
WebDataset shards
Each sample contains only:
{image_hash}.png
{image_hash}.json
laions_got_talent_rawEmolia
Dataset Card for Emolia
Dataset Description
This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?".
The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
synthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository.
https://huggingface.co/datasets/sleeping-ai/Vocal-burst
We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories.
It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts.
majestrino-dataCS-Arxiv-PDFs-08-25laions_got_talent_enhanced_no_metadatalaions_got_talent_embs_only
laions_got_talent Whisper Embeddings (Embeddings + Metadata Only)
This dataset contains Whisper embeddings (NPY) and metadata (JSON). The original audio files (MP3) are NOT included.
Embeddings computed with: mkrausio/EmoWhisper-AnS-Small-v0.1
Includes original audio: No
Includes metadata: Yes (JSON)
Includes embeddings: Yes (NPY)
Creation date: 2025-05-11
Emilia-with-Emotion-Annotations4laions_got_talent_german_bicodecEmilia-with-Emotion-Annotations5unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1clevr-webdatasetvoiceclap-data
VoiceCLAP Data
The audio + dense-caption mixture used to train
laion/voiceclap-small and
laion/voiceclap-large.
Each tar shard is a WebDataset of
paired <key>.flac (48 kHz mono audio) + <key>.json (caption + metadata)
samples. Captions and structured attribute annotations are produced
automatically by a pipeline of audio-aware LLMs — Qwen-Audio, Gemini Flash 2.5,
and a thinking-mode reasoning model that scores emotion under the EmoNet
taxonomy plus per-clip vocal-burst, timbre… See the full description on the dataset page: https://huggingface.co/datasets/laion/voiceclap-data.common-voice-subset-for-clapEmilia-with-Emotion-Annotations3laions_got_talent_previewlaion-coco-13m-tarEmilia-with-Emotion-Annotations2talent_plus_rl_groups_of_50_with_audiobox_scoresmajestrino-1.00-16xk5-sae-features
Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5)
Top-2000 activating audio samples for each feature in the
Majestrino 1.00 SAE.
Overview
Metric
Value
SAE Architecture
16x expansion, k=5, d_model=768
Total Features
12,288
Alive Features
10,684
Audio per Feature
Up to 2,000 highest-activating
Audio Format
Opus (24 kbps OGG container)
Total TAR Files
1069
Source Dataset
laion/majestrino-data
File Structure
Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.imagenetpp-laion-t2iDataset Card for ImageNet++'s LAION Text-to-Image Split
emonet-face-big
EmoNet-Face: A Fine-Grained, Expert-Annotated Benchmark for Facial Emotion Recognition
Dataset Summary
EmoNet-Face is a comprehensive benchmark suite designed to address critical gaps in facial emotion recognition (FER). Current benchmarks often have a narrow emotional spectrum, lack demographic diversity, and use uncontrolled imagery. EmoNet-Face provides a robust foundation for developing and evaluating AI systems with a deeper, more nuanced understanding of human… See the full description on the dataset page: https://huggingface.co/datasets/laion/emonet-face-big.timbre-audio-caption-pairs
