CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Telecom-Paris /iamd_v0 Internet Archive Music Dataset (IAMD v0) ~4.2M thirty-second music segments (34,469 hours) sourced from Creative-Commons audio on the Internet Archive, each paired with machine-generated natural-language captions and the original item metadata. Segments 4.2M Audio 34k hours Segment length 30 s nominal (mean 29.22 s) Format MP3, 320 kbps CBR, native channels + sample rate Shards 2,320 Parquet files Download size 4.53 TB Loading A… See the full description on the dataset page: https://huggingface.co/datasets/Telecom-Paris/iamd_v0.audioaudio-classification1M<n<10M6 likes2.1k downloads2mo agoHugging Face02IAsistemofinteres /vzaudio1K<n<10K1 likes210 downloads2y agoHugging Face03IAmSkyDra /LeViSQA-v1audio10K<n<100K0 likes164 downloads1y agoHugging Face04Bgeorge /Liv-IA-audiodatasetaudio10K<n<100K0 likes140 downloads11mo agoHugging Face05sengtha /iany-khmer-voice iAny Khmer Voice An open, community-contributed Khmer speech dataset for training speech-to-text (ASR). Recorded through the iAny "Contribute your voice" page (https://iany.app/voice), where people read short Khmer sentences aloud — with the community, for the community. Clips: 5629 Speakers: 79 (anonymous ids) Duration: 5.3 hours Audio: 16 kHz mono WAV Language: Khmer (km) License: CC-BY-SA-4.0 Structure Standard Hugging Face audiofolder:… See the full description on the dataset page: https://huggingface.co/datasets/sengtha/iany-khmer-voice.audioautomatic-speech-recognition1K<n<10K2 likes131 downloads2mo agoHugging Face06iashchak /ru_common_voice_sova_rudevices_golos_fleursaudio100K<n<1M0 likes120 downloads2y agoHugging Face07iamhuuphuc /vietnamese-music-datasetaudio1K<n<10K0 likes118 downloads1mo agoHugging Face08iamTangsang /OpenSLR54-Nepali-ASRaudio100K<n<1M0 likes115 downloads2y agoHugging Face09PaulChimzy01 /IARPA_BABEL_OP3_306audio10K<n<100K0 likes88 downloads11mo agoHugging Face10iamfortytwo /MusicAVQA-A2V-Retrieval MusicAVQA-A2V-Retrieval This is a derived retrieval benchmark from the test split of mteb/MUSIC-AVQA_cls-preprocessed at revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses audio queries and video corpus items. Construction The source clips are labelled with 22 musical-instrument classes. For every class, a deterministic seed (42) selects five clips as queries and ten distinct clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-A2V-Retrieval.audio1K<n<10K0 likes69 downloads9d agoHugging Face11iatop65 /Voice_Training_DataA database of the Audio Samples that was use to trained the Voice Models audio1K<n<10K0 likes47 downloads2y agoHugging Face12iamfortytwo /MusicAVQA-V2A-Retrieval MusicAVQA-V2A-Retrieval This is a derived retrieval benchmark from the test split of mteb/MUSIC-AVQA_cls-preprocessed at revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses video queries and audio corpus items. Construction The source clips are labelled with 22 musical-instrument classes. For every class, a deterministic seed (42) selects five clips as queries and ten distinct clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-V2A-Retrieval.audio1K<n<10K0 likes47 downloads9d agoHugging Face13Tairooonz /dataset_ia_strixaudio1K<n<10K0 likes40 downloads6mo agoHugging Face14IAMCB /elise-clone Custom Elise-like TTS Dataset Converted on 2025-06-18T05:56:51Z. Samples: 1043 Format : 10-s clips with text transcription (like MrDragonFox/Elise) Structure column type description audio audio 24kHz mono wav clip text string transcription audio1K<n<10K0 likes37 downloads1y agoHugging Face15IAmSkyDra /LeViSQAaudio10K<n<100K0 likes31 downloads1y agoHugging Face16IAmSkyDra /ViSQA-newaudion<1K0 likes25 downloads2y agoHugging Face17mmmsss1234 /iamnotthatkindoftalentaudion<1K0 likes25 downloads2mo agoHugging Face18fosters /iakub-kolas-kazki-zhytstsia-output_original Казкі жыцця — арыгінальнае аўдыё Аўтар / Author: Якуб КоласМова / Language: Беларуская (Belarusian) Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці. Частка калекцыі Ministerskija — корпус беларускіх аўдыёкніг. Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя): iakub-kolas-kazki-zhytstsia-output Доўгасць аўдыё 1h49m Радкоў у датасеце 504 Структура Кожны радок змяшчае: audio — арыгінальны аўдыёзапіс text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/iakub-kolas-kazki-zhytstsia-output_original.audioautomatic-speech-recognitionn<1K0 likes25 downloads3mo agoHugging Face19IAMCB /eightaudion<1K0 likes23 downloads1y agoHugging Face20iaaoli2 /florapaulitaariaudion<1K0 likes22 downloads2y agoHugging Face21Tairooonz /DS_STRIX_IAaudio1K<n<10K0 likes22 downloads7mo agoHugging Face22IamPre /ts_datasetaudio1K<n<10K0 likes21 downloads1y agoHugging Face23pengyizhou /IALP-2026-data IALP-2026: Whisper Open-Set Data-Selection — Query / Dev / Test Sets Supporting data for the study "Whisper-Based Open-Set Data Selection for NSC Adaptation." This repository holds the fixed target-query, validation, and evaluation sets used across all experiments. Each part is a self-contained .tar.gz. All audio is 16 kHz mono. Each split ships with: audio/ — audio files (FLAC, except GigaSpeech which is WAV PCM_16) wav.scp — <utt_id> audio/<file> (Kaldi-style, relative paths)… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/IALP-2026-data.audioautomatic-speech-recognition10K<n<100K0 likes19 downloads3mo agoHugging Face24BoltR /Cover.IAaudion<1K0 likes18 downloads6mo agoHugging Face25IAmSkyDra /ViSQA_plusaudion<1K0 likes18 downloads2y agoHugging Face26IamV /fluent_slu_v1.0 Dataset Card for "fluent_slu_v1.0" More Information needed audio10K<n<100K0 likes17 downloads3y agoHugging Face27IAMCB /24khz_800clipsaudion<1K0 likes17 downloads1y agoHugging Face28iamjamuna /AffectHuman-43Kgated AffectHuman-43K AffectHuman-43K is an emotion-aligned multimodal benchmark for controlled human affect generation and evaluation. The benchmark contains 42,469 usable samples with complete image, reference-image, audio, and text coverage. Identity is specified through a visual reference image, while text, audio, and emotion labels provide affective control signals. This design separates identity preservation from affective control, enabling evaluation of whether a model can preserve… See the full description on the dataset page: https://huggingface.co/datasets/iamjamuna/AffectHuman-43K.audioimage-to-image10K<n<100K0 likes17 downloads4mo agoHugging Face29IAMCB /emotional_tts_datasetaudion<1K0 likes16 downloads1y agoHugging Face30archivartaunik /ianka-sipakou-zialeny-listok-na-planetse-ziamlia-maryia-zakharevich Зялёны лісток на планеце Зямля Metadata Author: Янка Сіпакоў Title: Зялёны лісток на планеце Зямля Narrator: Марыя Захарэвіч Source Group: Аўдыёкнігі Source: Notes The original audio files are preserved as-is: no conversion; no re-encoding; no filename changes inside each split folder, except removing one common top-level archive folder when present. To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/ianka-sipakou-zialeny-listok-na-planetse-ziamlia-maryia-zakharevich.audion<1K0 likes16 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.