CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01liumindmind /Neko_Audio-30K_Longaudio10K<n<100K7 likes3k downloads4mo agoHugging Face02mitermix /audiosnippets_long_2_5Maudio1M<n<10M3 likes1.4k downloads2y agoHugging Face03holvan /LongAudioSpan LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension Introduction LongAudioSpan is a benchmark for long-form audio comprehension, spanning diverse durations and cognitive depths. Questions come from two complementary paths: Native QA: questions drawn from the audio's natural content. Anchor QA: questions built around acoustic anchors planted into the audio. Each path is scored in its own mode: Accuracy: multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/holvan/LongAudioSpan.textaudio-text-to-text1K<n<10K9 likes1.1k downloads9d agoHugging Face04mitermix /audiosnippets_long_1Maudio100K<n<1M0 likes823 downloads2y agoHugging Face05Yi3852 /midi-audio-abc_longmidi, synthesized audio, ABC code triples (this dataset contains those with audio duration in 5 min - 2 hours, less than 5 min data are in 300s and there are several subsets with smaller duration 60s 30s 10s) (token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer) midi files are from bread-midi-dataset synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer abc notation: mid2abc by EasyABC (midi2abc.py)… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_long.audio10K<n<100K1 likes379 downloads1y agoHugging Face06wcyh /LongAudioaudio0 likes205 downloads1y agoHugging Face07mitermix /audiosnippets_long_50kaudio10K<n<100K0 likes132 downloads2y agoHugging Face08AlexWortega /long_audio_it_full_finalaudio1K<n<10K0 likes120 downloads9mo agoHugging Face09oreva /blab_long_audio BLAB: Brutally Long Audio Bench Dataset Summary Brutally Long Audio Bench (BLAB) is a challenging long-form audio benchmark that evaluates audio LMs on localization, duration estimation, emotion, and counting tasks using audio segments averaging 51 minutes in length. BLAB consists of 833+ hours of diverse, full-length Youtube audio clips, each paired with human-annotated, text-based natural language questions and answers. Our audio data were collected from permissively… See the full description on the dataset page: https://huggingface.co/datasets/oreva/blab_long_audio.audioaudio-text-to-text1K<n<10K2 likes100 downloads5mo agoHugging Face10AudioLLMs /tedlium3_long_form_test@inproceedings{hernandez2018ted, title={TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation}, author={Hernandez, Fran{\c{c}}ois and Nguyen, Vincent and Ghannay, Sahar and Tomashenko, Natalia and Esteve, Yannick}, booktitle={Speech and Computer: 20th International Conference, SPECOM 2018, Leipzig, Germany, September 18--22, 2018, Proceedings 20}, pages={198--208}, year={2018}, organization={Springer} } @article{wang2024audiobench… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/tedlium3_long_form_test.audion<1K0 likes96 downloads2y agoHugging Face11AlexWortega /long_audio_it_fullaudio1K<n<10K1 likes26 downloads9mo agoHugging Face12AlexWortega /long_audio_itaudio1K<n<10K0 likes11 downloads9mo agoHugging Face13AlexWortega /longaudioaudion<1K0 likes10 downloads9mo agoHugging Face14humair025 /Soprano-Long-Audio-10s-Plus Soprano Long Audio Dataset (> 10s) This dataset is a consolidated collection of high-quality speech samples filtered specifically for long duration (> 10 seconds). It is designed for training TTS models on long-context speech. Statistics Total Samples: 95305 Total Duration: 395.94 hours Minimum Duration: 10.0 seconds Composition The dataset is merged from the following sources, strictly filtering for audio > 10s: Source Count Description source… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Soprano-Long-Audio-10s-Plus.text10K<n<100K0 likes8 downloads8mo agoHugging Face15metinovadilet /long_vocal_audiogatedTranscribed long audio vocals . We need to split them into short speech audion<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.