CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes11k downloads10h agoHugging Face02FBK-MT /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.audioautomatic-speech-recognition1K<n<10K70 likes1.2k downloads2mo agoHugging Face03Rijgersberg /YouTube-Commons-nl-audio YouTube Commons NL Audio This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions, all under a CC BY 4.0 license. It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB. Source The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.audioautomatic-speech-recognition10K<n<100K1 likes424 downloads1y agoHugging Face04Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes411 downloads3mo agoHugging Face05humyn-labs /APAC-Egocentric-Residential-Voiceover APAC Egocentric Residential (with Voiceover) Ten narrated first-person recordings of household chores, each shipping the original capture with spoken voiceover, a burned-in caption render, WebVTT captions, and an ASS annotation track. This is the only release in the HumynLabs egocentric collection that carries audio narration — the wearer describes each action as they perform it, and the captions align that speech to the video. Preview: 45 s from the cooking sample, captioned… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/APAC-Egocentric-Residential-Voiceover.tabularroboticsn<1K0 likes238 downloads2mo agoHugging Face06Rendra86318 /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/MCIF.audioautomatic-speech-recognition1K<n<10K0 likes93 downloads9mo agoHugging Face07vaishnavikedar4 /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavikedar4/MCIF.audioautomatic-speech-recognition1K<n<10K0 likes41 downloads9mo agoHugging Face08snorbyte /indic-audio-dialog-samplegated Dataset Card for Indic Dialog Sample Dataset Dataset Details Dataset Description The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.audioaudio-to-audio1K<n<10K1 likes23 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.