CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes11k downloads9h agoHugging Face02facebook /2M-Flores-ASL 2M-Flores As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest sentences in the original flores200 dataset. To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded. The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time. The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.tabulartranslation1K<n<10K2 likes1.6k downloads2y agoHugging Face03BAAI /Chinese-LiPS Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides ⭐ Introduction The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios. 🚀 Dataset Details Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.audioautomatic-speech-recognition10K<n<100K12 likes1.5k downloads10mo agoHugging Face04FBK-MT /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.audioautomatic-speech-recognition1K<n<10K70 likes1.2k downloads2mo agoHugging Face05tugrulbayrak /Real-TurnTurk Real-TurnTurk English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the… See the full description on the dataset page: https://huggingface.co/datasets/tugrulbayrak/Real-TurnTurk.tabularaudio-classification100K<n<1M2 likes504 downloads17h agoHugging Face06Rijgersberg /YouTube-Commons-nl-audio YouTube Commons NL Audio This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions, all under a CC BY 4.0 license. It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB. Source The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.audioautomatic-speech-recognition10K<n<100K1 likes424 downloads1y agoHugging Face07Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes411 downloads3mo agoHugging Face08Shofo /shofo-talking-head-engated Shofo Talking Head Dataset (English) A curated dataset of ~10,000 talking-head videos (~186 hours, mean ~67s/clip), filtered for clean single-speaker framing and paired with time-aligned transcripts. Built and released by Shofo. This dataset is short-form social video, designed to support modern avatar, lip-sync, dubbing, and TTS work. Every clip is curated by a multi-stage pipeline (face/framing analysis, on-screen-text detection, object-occlusion detection, voice/face matching… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-talking-head-en.videovideo-classification1K<n<10K1 likes386 downloads3mo agoHugging Face09H-oliday /TeleEgo-Source TeleEgo-Source Source Videos and Time-Aligned Transcripts for TeleEgo Official source release forTeleEgo: Benchmarking Egocentric AI Assistants in the Wild Overview TeleEgo is a multimodal benchmark for evaluating egocentric AI assistants in realistic, long-duration settings. It contains recordings from five participants over three days and covers four broad themes: Work & Study, Lifestyle &… See the full description on the dataset page: https://huggingface.co/datasets/H-oliday/TeleEgo-Source.videovideo-text-to-textn<1K0 likes348 downloads26d agoHugging Face10ohsn /darija_yt_2026 darija_yt_2026 Partition upload generated automatically. Namespace: ohsn Repo: ohsn/darija_yt_2026 Video count: 3511 Duration hours: 1565.31 This dataset contains raw audio files and per-video metadata generated from the Darija YouTube extraction pipeline. audioautomatic-speech-recognition1K<n<10K0 likes269 downloads19d agoHugging Face11IndabaXSudan /Sudan-MM Sudan-MM: A Multimodal Dataset of Sudanese Arabic Sudan-MM is the first publicly available multimodal dataset for Sudanese Arabic (السودانية), a low-resource dialect with no prior paired image-caption, video-caption, or voice-caption data. It was produced through a competitive shared task held in 2025, where five teams collected and annotated media depicting everyday Sudanese life. Each item in the dataset pairs a visual or video recording with: a written caption in Modern Standard… See the full description on the dataset page: https://huggingface.co/datasets/IndabaXSudan/Sudan-MM.audioimage-to-text1K<n<10K2 likes268 downloads4mo agoHugging Face12humyn-labs /APAC-Egocentric-Residential-Voiceover APAC Egocentric Residential (with Voiceover) Ten narrated first-person recordings of household chores, each shipping the original capture with spoken voiceover, a burned-in caption render, WebVTT captions, and an ASS annotation track. This is the only release in the HumynLabs egocentric collection that carries audio narration — the wearer describes each action as they perform it, and the captions align that speech to the video. Preview: 45 s from the cooking sample, captioned… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/APAC-Egocentric-Residential-Voiceover.tabularroboticsn<1K0 likes238 downloads2mo agoHugging Face13kadirnar /ai-researcher-roadmap-media AI Researcher Roadmap Media Optional video and subtitle assets for the AI Researcher Roadmap application. Repository layout manifest.json: file sizes and SHA-256 checksums used by the application. videos/<stem>.mp4: lecture video. subs/<stem>.<language>.vtt: subtitle tracks. subs/<stem>.asr.<language>.vtt: ASR-generated subtitle tracks. The application downloads only the selected lecture and its subtitle tracks. Files are cached locally and can be played offline… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/ai-researcher-roadmap-media.videoautomatic-speech-recognitionn<1K0 likes147 downloads2mo agoHugging Face14Fhrozen /CABankSakura CABank Japanese Sakura Corpus Susanne Miyata Department of Medical Sciences Aichi Shukotoku University smiyata@asu.aasa.ac.jp website: https://ca.talkbank.org/access/Sakura.html Important This data set is a copy from the original one located at https://ca.talkbank.org/access/Sakura.html. Details Participants: 31 Type of Study: xxx Location: Japan Media type: audio DOI: doi:10.21415/T5M90R Citation information Some citation here. In… See the full description on the dataset page: https://huggingface.co/datasets/Fhrozen/CABankSakura.videoaudio-classificationn<1K0 likes94 downloads4y agoHugging Face15Rendra86318 /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/MCIF.audioautomatic-speech-recognition1K<n<10K0 likes93 downloads9mo agoHugging Face16Arozo /aiza-malagasy-administrative-asr AIZA — Malagasy/French Code-Switched Administrative Speech (Benchmark Set) 10 short, self-recorded, consented audio clips of Malagasy/French code-switched questions about Malagasy administrative procedures (lost ID card, birth certificate, passport, land title, etc.) — created as the benchmark audio for AIZA, a submission to the Sahara CodeSwitch Africa Main Challenge. Why this dataset exists Neither Sahara's published language list nor the official… See the full description on the dataset page: https://huggingface.co/datasets/Arozo/aiza-malagasy-administrative-asr.audioautomatic-speech-recognitionn<1K0 likes84 downloads6d agoHugging Face17kurianbenoy /Indic-subtitler-audio_evals Indic_audio_evals As part of this project. We are evaluating our performance of various ASR models as well in a benchmarking dataset, we have created in various languages. This benchmarking dataset is more alligned to real-world use-cases rather than having any academic datasets. About Dataset Dataset Link in HuggingFace: kurianbenoy/Indic-subtitler-audio_evals This dataset contains audio file in .wav format and video file in .mp4. The respective groundtruth will be… See the full description on the dataset page: https://huggingface.co/datasets/kurianbenoy/Indic-subtitler-audio_evals.audioautomatic-speech-recognitionn<1K2 likes49 downloads2y agoHugging Face18vaishnavikedar4 /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavikedar4/MCIF.audioautomatic-speech-recognition1K<n<10K0 likes41 downloads9mo agoHugging Face19juliasdata /medical-audio-sample-brazilian-portuguese Julia's Data: Brazilian Portuguese Medical Audio Sample Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker. This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio. Full dataset and commercial licensing: juliasdata.com Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.audioautomatic-speech-recognitionn<1K1 likes40 downloads6mo agoHugging Face20ZamAI-Pashto /zamai-pashto-video ZamAI Pashto Video ZamAI Pashto Video is a video understanding dataset scaffold for multilingual Afghan media research, with support for scene segmentation, subtitle alignment, and temporal event annotation. Dataset Summary The repository is structured for raw and segmented video assets, subtitle generation, action labels, and media metadata needed for temporal analysis workflows. Languages Pashto Dari English Modalities Video… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-video.textvideo-classificationn<1K0 likes39 downloads2mo agoHugging Face21DigiGreen /AgricultureVideosTranscriptThis dataset consists of agriculture videos in hindi and oriya. The dataset consists of mulitple xls files and each xls file has column of video urls (youtube video links) and corresponding transcripts. The transcripts are: Generated by ASR models (for the purpose of benchmarking) Manual transcripts Time stamps Manual translations This dataset can be used for training and benchmarking domain specific models for ASR and translation. The time stamps serve as the soruce of audio and the… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/AgricultureVideosTranscript.videotranslation1K<n<10K0 likes32 downloads2y agoHugging Face22alj68 /2M-Flores-ASL 2M-Flores As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest sentences in the original flores200 dataset. To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded. The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time. The… See the full description on the dataset page: https://huggingface.co/datasets/alj68/2M-Flores-ASL.tabulartranslation1K<n<10K0 likes31 downloads9mo agoHugging Face23scubavoice /ibibio-efik-speech-corpus-sample Scuba Voice Dataset: Ibibio & Efik Sample (1 hour) Scuba is a voice data infrastructure company collecting, labeling, and licensing speech corpora for low-resource languages. This repository contains a 1-hour sample of labeled Ibibio and Efik speech data collected from native speakers in Akwa Ibom, Nigeria. This sample is provided for evaluation purposes only. For access to the full dataset (~200 hours across Ibibio and Efik), contact us at info@scubavoice.com.… See the full description on the dataset page: https://huggingface.co/datasets/scubavoice/ibibio-efik-speech-corpus-sample.audioautomatic-speech-recognitionn<1K1 likes25 downloads3mo agoHugging Face24snorbyte /indic-audio-dialog-samplegated Dataset Card for Indic Dialog Sample Dataset Dataset Details Dataset Description The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.audioaudio-to-audio1K<n<10K1 likes23 downloads1y agoHugging Face25egirma /afrivoice-swahili-agriculture-subset Dataset Card for the image text and voice dataset Dataset Description Subset of Afrivoice dataset: approx. 100 hours of train (sampled, stratified) Full dev + test from original repo (DigitalUmuganda/Afrivoice_Swahili) Audio: .webm Includes images + transcriptions from original repo (DigitalUmuganda/Afrivoice_Swahili) Includes metadata csv License CC-BY-4.0 (derived from DigitalUmuganda/Afrivoice_Swahili) Notes Train split sampled using… See the full description on the dataset page: https://huggingface.co/datasets/egirma/afrivoice-swahili-agriculture-subset.audioautomatic-speech-recognition10K<n<100K0 likes22 downloads6mo agoHugging Face26ud-nlp /portuguese-speech-recognition-dataset Portuguese Telephone Dialogues Dataset - 10 Hours Dataset comprises 10 hours of high-quality telephone audio recordings in Portuguese, featuring 20+ native speakers and achieving a 98% Word Accuracy Rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/portuguese-speech-recognition-dataset.audioautomatic-speech-recognitionn<1K0 likes20 downloads10mo agoHugging Face27Denhotech /swahili-video-text Swahili Video-Text Dataset Automatically processed Swahili video clips and transcriptions. videoautomatic-speech-recognition0 likes20 downloads7mo agoHugging Face28CGIAR /AgricultureVideosTranscriptThis dataset consists of agriculture videos in hindi and oriya. The dataset consists of mulitple xls files and each xls file has column of video urls (youtube video links) and corresponding transcripts. The transcripts are: Generated by ASR models (for the purpose of benchmarking) Manual transcripts Time stamps Manual translations This dataset can be used for training and benchmarking domain specific models for ASR and translation. The time stamps serve as the soruce of audio and the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosTranscript.videotranslation1K<n<10K0 likes19 downloads2y agoHugging Face29octanove /moslagated Overview The MOSLA dataset ("MOSLA") is a longitudinal, multimodal, multilingual, and controlled dataset created by inviting participants to learn one of three target languages (Arabic, Spanish, and Chinese) from scratch over a span of two years, exclusively through online instruction, and recording every lesson using Zoom. The dataset is semi-automatically annotated with speaker/language IDs and transcripts by both human annotators and fine-tuned state-of-the-art speech models.… See the full description on the dataset page: https://huggingface.co/datasets/octanove/mosla.tabularautomatic-speech-recognition100K<n<1M5 likes14 downloads2y agoHugging Face30JacobLinCool /anime-2024gated Anime Video Dataset 2024 Overview This is an anime video dataset curated specifically for multimedia research purposes. It is sourced from anime series released in 2024 and includes both a complete combined file and four seasonal subsets (Winter, Spring, Summer, and Fall). Data Details Video Stream Codec: H.264 (High Profile) Format: YUV 4:2:0 (Progressive) Resolution: 640×360 (SAR 1:1, DAR 16:9) Frame Rate: 23.98 FPS Audio Stream Codec: AAC (LC) Sample… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/anime-2024.textvideo-classification1K<n<10K1 likes13 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.