datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LAION-Audio-300Msoundscapeslaions_got_talent
LAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel.
Dataset Composition
The dataset includes:
Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.BVD-A-10M-URLs
LAION-BVD — 10M Audio Clip URLs
This repository contains the metadata and captions for ~10 million audio clips randomly sampled
from BVD-V-55M for large-scale audio pre-training.
The audio itself is not included in this repository — every clip is described by the URL of
its source video plus the start_time/end_time offsets needed to reproduce it. The
corresponding clip files are available in the gated
laion/BVD-A-10M repository.
Dataset structure
One row per audio… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-10M-URLs.laion-audio-previewlaions_got_talent_rawBVD-A-1.7M-URLs
LAION-BVD — 1.7M Audio Clip URLs
This repository contains the metadata and captions for ~1.7 million audio clips taken from
BVD-V-55M and sampled for uniqueness of the
source video, so that the subset maximises source diversity rather than clip count.
The audio itself is not included in this repository — every clip is described by the URL of
its source video plus the start_time/end_time offsets needed to reproduce it. The
corresponding clip files are available in the gated… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-1.7M-URLs.freesound-laion-640k
About this Repository
This repository is a re-upload of the FreeSound.org dataset as curated by LAION for the larger LAION-Audio-630k dataset, with the following changes:
Limited columns to only the audio and basic metadata.
Incorporated necessary information for licensing and attribution.
Removed ambiguously licensed samples, amounting to around 1,000 total samples.
What about download links?
Links were ommitted for the sake of size, as they can be constructed from… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/freesound-laion-640k.laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants.
Updated Composition
Voices and Languages
English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.Emolia
Dataset Card for Emolia
Dataset Description
This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?".
The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
freesound-laion-640k-commercial-16khz-full
About this Repository
This repository is the training split of the complete FreeSound LAION 640k dataset, limited only to licenses that permit commercial works, resampled to 16khz using torchaudio.transforms.Resample.
This is ideal for use cases where a variety of audio is desired but fidelity and labels are unnecessary, such as background audio for augmenting other datasets.
Dataset Versions
You are looking at the full dataset which contains 403,146 unique sounds… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/freesound-laion-640k-commercial-16khz-full.majestrino-datasynthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository.
https://huggingface.co/datasets/sleeping-ai/Vocal-burst
We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories.
It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts.
laions_got_talent_enhanced_no_metadataemolia-thinking-balanced-buckets
Emolia-Thinking — Balanced Per-Dimension Bucket Subset
A balanced, per-dimension bucket subset of
VoiceNet/emolia-thinking,
derived from that dataset's zero-shot VoiceNet-dimension labels.
For every VoiceNet voice/prosody/timbre/style dimension, this subset draws a
roughly equal number of clips from each ordinal bucket (0–6), so that
downstream training / probing sees a balanced distribution along each axis
instead of the strongly skewed natural distribution.
How… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-thinking-balanced-buckets.Emilia-with-Emotion-Annotations4laions_got_talent_german_bicodecEmilia-with-Emotion-Annotations5unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1common-voice-subset-for-clapvoiceclap-data
VoiceCLAP Data
The audio + dense-caption mixture used to train
laion/voiceclap-small and
laion/voiceclap-large.
Each tar shard is a WebDataset of
paired <key>.flac (48 kHz mono audio) + <key>.json (caption + metadata)
samples. Captions and structured attribute annotations are produced
automatically by a pipeline of audio-aware LLMs — Qwen-Audio, Gemini Flash 2.5,
and a thinking-mode reasoning model that scores emotion under the EmoNet
taxonomy plus per-clip vocal-burst, timbre… See the full description on the dataset page: https://huggingface.co/datasets/laion/voiceclap-data.laion_audio_contrastive
⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models.
Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive
📜 Citation
If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/laion_audio_contrastive.Emilia-with-Emotion-Annotations3voice-acting-data-annotated
Voice Acting Data - Annotated
Post-processed version of laion/voice-acting-data.
Processing Pipeline
RE-USE Speech Enhancement (nvidia/RE-USE) - Applied to non-singing samples for noise reduction
LavaSR Super Resolution (YatharthS/LavaSR) - Audio bandwidth extension to 48kHz
Whisper Turbo ASR - Full transcript with word-level timestamps
Scene Split - Audio split at CUT TO: transition into two parts (Part 1 + Part 2)
VoiceCLAP Large Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-data-annotated.laions_got_talent_previewlaion-coco-13m-tarEmilia-with-Emotion-Annotations2Emilia-Annotated-WIPStill a WIP, full dataset is still being annotated
talent_plus_rl_groups_of_50_with_audiobox_scores
