datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
small-overlapping-speech-bench
Small Overlapping Speech Bench
A tiny, fully-reproducible benchmark for multilingual overlapping speech. Each of the
100 clips contains three people speaking at the same time, each in a different European
language, with ground-truth per-speaker timestamps, languages, and transcripts.
It is a deliberately hard "cocktail-party" stress test: how much of each simultaneous speaker can
an ASR (speech-to-text) model recover, and can a model tell how many people are talking?
100 clips… See the full description on the dataset page: https://huggingface.co/datasets/laion/small-overlapping-speech-bench.soundscape-bench
SoundScape-Bench
200 held-out multilingual soundscapes with exact, automatically-gradable answer keys for evaluating
"universal audio annotation" — the task of describing everything audible in a clip (speech, who/when/
what/which-language/how-it-is-said, sound effects, music, and vocal bursts) as one structured JSON list.
It is the benchmark for the LAION Universal Audio Annotation Pipeline (UAAP).
Why it exists
Every clip is built by gluing together pieces we… See the full description on the dataset page: https://huggingface.co/datasets/laion/soundscape-bench.eurospeech-enhanced-dacvae
EuroSpeech parliamentary speech converted to DAC VAE latents
Source
disco-eth/EuroSpeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/laion/eurospeech-enhanced-dacvae.
