CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ming030890 /youtube_caption_yue YouTube ASR Caption Dataset (Cantonese) This dataset was built from YouTube videos with manually provided captions in Cantonese. We used SenseVoice to re-transcribe the audio and filtered segments to build a high-quality collection of audio-caption pairs. What’s included Segments where the ASR output is identical to the original caption — likely clean. Segments where differences are only homophones (同音字) or English words — likely ASR mistakes. This combination supports… See the full description on the dataset page: https://huggingface.co/datasets/ming030890/youtube_caption_yue.audio10K<n<100K2 likes3.5k downloads1y agoHugging Face02capleaf /viVoicegated Important Note ⚠️ This dataset is only to be used for research purposes. Access requests must be made via your school, institution, or work email. Requests from common email services will be rejected. We apologize for any inconvenience. viVoice: Enabling Vietnamese Multi-Speaker Speech Synthesis For a comprehensive description, please visit https://github.com/thinhlpg/viVoice This dataset is licensed under CC-BY-NC-SA-4.0 and is intended for research purposes only.… See the full description on the dataset page: https://huggingface.co/datasets/capleaf/viVoice.audiotext-to-speech100K<n<1M93 likes2.5k downloads2y agoHugging Face03laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M15 likes1.9k downloads11mo agoHugging Face04TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes1.8k downloads6mo agoHugging Face05ExylosAi /egocentric-vr-capture-20h-multimodal-sample Egocentric VR Capture — 20-Hour Multimodal Inspection Sample 195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package. This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.tabularrobotics1M<n<10M0 likes803 downloads3d agoHugging Face06yuriyvnv /capes_synthetic_audio_filteredaudio10K<n<100K0 likes626 downloads1y agoHugging Face07MYJOKERML /audioset_captionaudio10K<n<100K0 likes496 downloads1y agoHugging Face08mrfakename /cc0-music-captionedCollected from various CC0 music sites. These are all instrumental - none have lyrics. They are annotated with a description of the song. Some of the music comes from FreePD, a site that shared public domain music. The FreePD website has since been taken down. All songs were created by humans, not AI-generated. audio1K<n<10K2 likes384 downloads9mo agoHugging Face09TMICCProj /Closed_Captioning_Lecture_Datasetaudio10K<n<100K0 likes369 downloads7mo agoHugging Face10morateng /CapTTS-SFT-ears-cleanedaudio10K<n<100K2 likes351 downloads1y agoHugging Face11mispeech /MECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks 📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF) Dataset Description MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks: Audio Captioning: Generating textual descriptions for given audio Audio Question Answering: Answering questions about given audio Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.audioaudio-classification10K<n<100K4 likes328 downloads5mo agoHugging Face12gijs /tacos-captioningaudio10K<n<100K1 likes290 downloads1y agoHugging Face13JackyHoCL /common_voice_22_yue_w_background_captionMerged JackyHoCL/urban-noise-uganda-61k-caption, OpenSound/AudioCaps TODO: convert to MP3, reduce size audio100K<n<1M0 likes284 downloads5mo agoHugging Face14surindersinghssj /gurbani-sehajpath-yt-captions-canonical Gurbani Sehajpath — Canonical-aligned ASR corpus Stage-1 + Stage-2 canonical pipeline output for sehaj-path (calm recitation of the Guru Granth Sahib). Built from publicly available audio recordings with aligned transcripts, chunked by caption timing and aligned against the canonical Guru Granth Sahib Ji text (SGGS). Columns Schema is auto-inferred from the parquet shards. Primary columns: audio — 16 kHz mono waveform final_text — canonical Gurmukhi transcription (post… See the full description on the dataset page: https://huggingface.co/datasets/surindersinghssj/gurbani-sehajpath-yt-captions-canonical.audioautomatic-speech-recognition10K<n<100K0 likes240 downloads5mo agoHugging Face15introvoyz042 /egocentric-vr-capture-1h-multimodal-sample Egocentric VR Capture — 1-Hour Multimodal Inspection Sample 13 real-world task episodes / 108,029 frames / approximately 60 minutes captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package. This publicly accessible dataset is an inspection slice produced by the EXYLOS real-world data pipeline. It demonstrates capture quality, synchronization, schema, and QA metadata… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/egocentric-vr-capture-1h-multimodal-sample.tabularrobotics100K<n<1M0 likes239 downloads28d agoHugging Face16vucinatim /spectrogram-captionsDataset of captioned spectrograms (text describing the sound). audiotext-to-image1K<n<10K4 likes217 downloads4y agoHugging Face17laion /timbre-audio-caption-pairsaudio100K<n<1M2 likes206 downloads9mo agoHugging Face18laion /freesound-commercially-permissive-subset-with-captionsaudio100K<n<1M2 likes194 downloads11mo agoHugging Face19Evan-Lin /captioneraudion<1K0 likes151 downloads5mo agoHugging Face20JackyHoCL /cv22_yue_captionaudio100K<n<1M0 likes142 downloads6mo agoHugging Face21mitermix /audioset-with-grounded-captionsaudio1M<n<10M4 likes139 downloads1y agoHugging Face22mohammadhossein /1125_trad_sample_16k_caption_RoformerSeparateaudio1K<n<10K0 likes137 downloads2y agoHugging Face23hidude562 /operation-cappicinoaudio10K<n<100K0 likes130 downloads5mo agoHugging Face24JiaWeiDong /XCT_Data_of_Retired_Large_Capacity_Prismatic_BatteriesThis dataset provides 3D XCT reconstruction results for eight retired and one fresh LiFePO4 (LFP) prismatic batteries. These samples were manufactured by REPT BATTERO Energy Co., Ltd. in 2021 and sourced from Dongfeng E70 electric vehicles. Each battery has a nominal voltage of 3.2 V and a rated capacity of 135 Ah (dimensions: 148 mm × 80 mm × 105 mm). Compared to the fresh control sample (136.7 Ah), the retired batteries exhibit varying degrees of degradation, with measured capacities ranging… See the full description on the dataset page: https://huggingface.co/datasets/JiaWeiDong/XCT_Data_of_Retired_Large_Capacity_Prismatic_Batteries.audion<1K0 likes125 downloads10mo agoHugging Face25mb23 /music_caps_4sec_wave_typeaudio10K<n<100K4 likes121 downloads3y agoHugging Face26seastar105 /emo_speech_caption_testaudio1K<n<10K1 likes98 downloads2y agoHugging Face27capa2000 /binary-classifier-birdnet Binary BirdNet Classifier Contiene anotaciones y audios de 3s y 5s para clasificación binaria con rutas relativas. audion<1K0 likes94 downloads1y agoHugging Face28seungheondoh /music-audio-pseudo-captions Dataset Card for Music-Audio-Pseudo Captions Pseudo Music and Audio Captions from LP-MusicCaps, Music Negation/Temporal Ordering WavCaps Dataset Summary Compared to other domains, music and audio domains cannot obtain well-written web caption data, and caption annotation is expensive. Therefore, we use the Music (LP-MusicCaps), (Music Negation/Temporal Ordering) and Audio (Wavcaps) datasets created with ChatGPT to re-organize them in the form of instructions, input… See the full description on the dataset page: https://huggingface.co/datasets/seungheondoh/music-audio-pseudo-captions.text1M<n<10M4 likes92 downloads3y agoHugging Face29TTS-AGI /majestrino-unified-detailed-captions-temporal Majestrino Unified Detailed Captions with Temporal Aspects Filtered subset of laion/majestrino-data containing only samples with unified_detailed_caption_with_temporal_aspects. Stats 4,128,665 samples 826 tar files (~1.1 GB each) ~878 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption with temporal aspects caption_type — always unified_detailed_caption_with_temporal_aspects transcription — speech… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions-temporal.audioaudio-classification1M<n<10M0 likes81 downloads6mo agoHugging Face30laion /audioset-with-captionsaudio1M<n<10M2 likes72 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.