datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Neko_Audio-30K_Longaudiosnippets_long_2_5MLongAudioSpan
LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension
Introduction
LongAudioSpan is a benchmark for long-form audio comprehension, spanning
diverse durations and cognitive depths.
Questions come from two complementary paths:
Native QA: questions drawn from the audio's natural content.
Anchor QA: questions built around acoustic anchors planted into the
audio.
Each path is scored in its own mode:
Accuracy: multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/holvan/LongAudioSpan.audiosnippets_long_1Mmidi-audio-abc_longmidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5 min - 2 hours, less than 5 min data are in 300s
and there are several subsets with smaller duration
60s
30s
10s)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_long.LongAudioaudiosnippets_long_50klong_audio_it_full_finalblab_long_audio
BLAB: Brutally Long Audio Bench
Dataset Summary
Brutally Long Audio Bench (BLAB) is a challenging long-form audio benchmark that evaluates audio LMs on localization, duration estimation, emotion, and counting tasks using audio segments averaging 51 minutes in length. BLAB consists of 833+ hours of diverse, full-length Youtube audio clips, each paired with human-annotated, text-based natural language questions and answers. Our audio data were collected from permissively… See the full description on the dataset page: https://huggingface.co/datasets/oreva/blab_long_audio.tedlium3_long_form_test@inproceedings{hernandez2018ted,
title={TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation},
author={Hernandez, Fran{\c{c}}ois and Nguyen, Vincent and Ghannay, Sahar and Tomashenko, Natalia and Esteve, Yannick},
booktitle={Speech and Computer: 20th International Conference, SPECOM 2018, Leipzig, Germany, September 18--22, 2018, Proceedings 20},
pages={198--208},
year={2018},
organization={Springer}
}
@article{wang2024audiobench… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/tedlium3_long_form_test.long_audio_it_fulllong_audio_itlongaudioSoprano-Long-Audio-10s-Plus
Soprano Long Audio Dataset (> 10s)
This dataset is a consolidated collection of high-quality speech samples filtered specifically for long duration (> 10 seconds). It is designed for training TTS models on long-context speech.
Statistics
Total Samples: 95305
Total Duration: 395.94 hours
Minimum Duration: 10.0 seconds
Composition
The dataset is merged from the following sources, strictly filtering for audio > 10s:
Source
Count
Description
source… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Soprano-Long-Audio-10s-Plus.long_vocal_audioTranscribed long audio vocals . We need to split them into short speech
