CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mjuicem /StreamingBench StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.imagequestion-answering1K<n<10K13 likes12k downloads1y agoHugging Face02WenhaoWang /VidProM Summary This is the dataset proposed in our paper VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models (NeurIPS 2024). VidProM is the first dataset featuring 1.67 million unique text-to-video prompts and 6.69 million videos generated from 4 different state-of-the-art diffusion models. It inspires many exciting new research areas, such as Text-to-Video Prompt Engineering, Efficient Video Generation, Fake Video Detection, and Video Copy… See the full description on the dataset page: https://huggingface.co/datasets/WenhaoWang/VidProM.tabulartext-to-video1M<n<10M83 likes12k downloads1y agoHugging Face03Beakerman0101 /trex-visualizer T-Rex Dataset Visualizer A browseable subset of the T-Rex dataset — Tactile-Rich Bimanual Dexterous Manipulation — collected on a bimanual Dexmate Vega-1 robot equipped with two Sharpa Wave dexterous hands. This visualizer subset contains 3,838 short trajectory clips drawn from the full 100-hour T-Rex collection, organized by (verb, object, hand) so you can quickly inspect coverage across motion primitives and object categories. For the full dataset (multi-view RGB, robot… See the full description on the dataset page: https://huggingface.co/datasets/Beakerman0101/trex-visualizer.textrobotics1K<n<10K0 likes12k downloads3mo agoHugging Face04APRIL-AIGC /UltraVideo UltraVideo: High-Quality UHD 4K Video Dataset 🤓 Project    | 📑 Paper    | 🤗 Hugging Face (UltraVideo Dataset))   | 🤗 Hugging Face (UltraVideo-Long Dataset))   | 🤗 Hugging Face (UltraWan-1K/4K Weights)   UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions 🎋 Click below image to watch the 4K demo video. 🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo.tabularimage-to-video10K<n<100K66 likes11k downloads1y agoHugging Face05Linzhan /Mixamo-Animations-Characters Mixamo Animations and Characters A complete snapshot of the Mixamo library: 2,317 motion clips and 114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata. All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any compatible character without remapping. Use animation_motion/ and character_refined/. The full export contains 2,446 animation files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Mixamo-Animations-Characters.3dtext-to-3d1K<n<10K4 likes8k downloads14d agoHugging Face06tanish434 /Mixamo-Animations-Characters Mixamo Animations and Characters A complete snapshot of the Mixamo library: 2,317 motion clips and 114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata. All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any compatible character without remapping. Use animation_motion/ and character_refined/. The full export contains 2,446 animation files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Mixamo-Animations-Characters.3dtext-to-3d1K<n<10K0 likes7.6k downloads9d agoHugging Face07lightly-ai /epic-kitchens-100-clips EPIC-KITCHENS-100 Extracted Clips About Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset, more precisely the extension part not contained in EPIC-KITCHENS-55. For details, see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio. The clips folder contains one video for every narration from action annotations stored in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.tabular10K<n<100K2 likes6.8k downloads6mo agoHugging Face08skylenage /FilmBench FilmBench — Video Generation Benchmark Dataset 📢 Update (2026-08-02) Added English prompt files: filmbench_prompts_en.csv is now available with English prompts. filmbench_prompts_en.csv (1,169 rows): Prompt-level table with English prompts. Columns: uid, task, movie_type (English), en_prompt, reference_url. 📢 Update (2026-07-31) Fixed a batch of misaligned prompts in filmbench_videos.csv: the zh_prompt column has been recalibrated against the… See the full description on the dataset page: https://huggingface.co/datasets/skylenage/FilmBench.text1K<n<10K3 likes6.6k downloads2mo agoHugging Face09TencentARC /VPData VideoPainter This repository contains the implementation of the paper "VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control" Keywords: Video Inpainting, Video Editing, Video Generation Yuxuan Bian12, Zhaoyang Zhang1‡, Xuan Ju2, Mingdeng Cao3, Liangbin Xie4, Ying Shan1, Qiang Xu2✉ 1ARC Lab, Tencent PCG 2The Chinese University of Hong Kong 3The University of Tokyo 4University of Macau ‡Project Lead ✉Corresponding Author             Your… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/VPData.tabularimage-to-video100K<n<1M21 likes6.2k downloads1y agoHugging Face10Skyrmion /DAAD-XAbout Dataset The DAADX Dataset is derived from DAAD dataset (https://cvit.iiit.ac.in/research/projects/cvit-projects/daad#dataset), which contains all the captured videos for the Driver Intention Prediction task. We are introducing the first video based explanations dataset for driver intention prediction task. This will be help in further the research interms of making an explainable Autonomous Driving or ADAS System. DAAD-X contains explanations for each maneuver instance, these… See the full description on the dataset page: https://huggingface.co/datasets/Skyrmion/DAAD-X.tabularvideo-classification1K<n<10K1 likes4.9k downloads6mo agoHugging Face11stdKonjac /Sparkle Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou 📦 Dataset Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper. The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle.imagetext-to-video100K<n<1M1 likes4.4k downloads5mo agoHugging Face12jherng /wlasl_reduced Reduced WLASL Dataset This dataset is a task-specific reduced version of the WLASL dataset, constructed for American Sign Language (ASL) recognition experiments. Contents videos/Video clips organized by gloss label. metadata.csvPer-sample metadata including: file path gloss label fps (after normalization, if applied) video resolution normalized bounding box coordinates gloss_map.jsonMapping from gloss labels to integer class IDs. Dataset Construction… See the full description on the dataset page: https://huggingface.co/datasets/jherng/wlasl_reduced.tabularn<1K1 likes4.1k downloads8mo agoHugging Face13plnguyen2908 /AV-SpeakerBench AV-SpeakerBench Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning. Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/ Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench Paper: https://arxiv.org/abs/2512.02231 Files test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.audioquestion-answering1K<n<10K2 likes4.1k downloads9mo agoHugging Face14facebook /IntPhys2 IntPhys 2 Dataset   |   Hugging Face   |   Paper   |   Blog IntPhys 2 is a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2… See the full description on the dataset page: https://huggingface.co/datasets/facebook/IntPhys2.text1K<n<10K14 likes4k downloads1y agoHugging Face15yzy666 /SVBench Dataset Card for SVBench This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources. For details, see our Project, Paper and GitHub repository. Dataset Details Dataset Description SVBench is the first benchmark specifically designed to evaluate long-context streaming video understanding through temporal multi-turn question-answering (QA) chains. It addresses the limitations of existing video… See the full description on the dataset page: https://huggingface.co/datasets/yzy666/SVBench.textquestion-answering1K<n<10K7 likes3.9k downloads10mo agoHugging Face16APRIL-AIGC /UltraVideo-Long UltraVideo: High-Quality UHD 4K Video Dataset 🤓 Project    | 📑 Paper    | 🤗 Hugging Face (UltraVideo Dataset))   | 🤗 Hugging Face (UltraVideo-Long Dataset))   | 🤗 Hugging Face (UltraWan-1K/4K Weights)   UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions 🎋 Click below image to watch the 4K demo video. 🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo-Long.tabularimage-to-video10K<n<100K7 likes3.8k downloads1y agoHugging Face17Linzhan /UniML3D UniML3D UniML3D is the text-paired, topology-annotated motion dataset behind UniMate (SIGGRAPH Asia 2026): motion clips from three sources with very different skeletons — Mixamo humanoids, Truebones ZOO animals and rigged Objaverse-XL objects — brought into one canonical layout, captioned, and annotated with cleaned joint names, a body-plan category and a facing-direction joint pair per skeleton. Every annotation in it was generated by this project's own data… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/UniML3D.imagetext-to-3d10K<n<100K7 likes3.6k downloads16h agoHugging Face18THU-IAR /MIntRec Dataset details In real-world conversational interactions, we usually combine information from multiple modalities (e.g., text, video, audio) to help analyze human intentions. Though intent analysis has been widely explored in the Natural Language Processing community, there is a scarcity of data for multimodal intent analysis. Thus, we provide a novel multimodal intent benchmark dataset, MIntRec, to boom the research. To the best of our knowledge, it is the first multimodal intent… See the full description on the dataset page: https://huggingface.co/datasets/THU-IAR/MIntRec.text1K<n<10K1 likes3.6k downloads2y agoHugging Face19tanish434 /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.tabular1K<n<10K0 likes3.1k downloads9d agoHugging Face20larrylarrylarry1 /VideoArtifactDetectionVideo Artifact Detection dataset using source videos from LongVideoBench. Video level labels are given in labels.csv (training) and labels_test.csv (testing). Localized artifact regions for burst artifacts are given in the artifact_ranges column. Note that labels.csv contains additional source videos from LongVideoBench that are not included in this repository. You may visit the LongVideoBench page for the additional videos. For additional questions, please email palmerla@usc.edu. tabular1K<n<10K0 likes2.4k downloads4mo agoHugging Face21KlingTeam /FullBenchimage1K<n<10K7 likes2k downloads1y agoHugging Face22shi-labs /physical-ai-bench-conditional-generation Physical AI Bench - Conditional Generation Paper | Code This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below. Dataset Category Sample Nums Agibot World Robotics 200 OpenDV… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-conditional-generation.textvideo-to-videon<1K0 likes1.9k downloads10mo agoHugging Face23Linzhan /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Truebones-ZOO-Annotations.tabular1K<n<10K1 likes1.7k downloads14d agoHugging Face24SII-KYW /CogStream CogStream Dataset Dataset for CogStream: Context-guided Streaming Video Question Answering. Overview CogStream is a streaming video QA dataset designed to evaluate context-guided video reasoning. Models must identify and utilize relevant historical context to answer questions about ongoing video streams. Statistics: Split Videos QA Pairs Train 852 55,623 Test 236 15,364 Total 1,088 70,987 Sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%)… See the full description on the dataset page: https://huggingface.co/datasets/SII-KYW/CogStream.tabularquestion-answering10K<n<100K1 likes1.6k downloads7mo agoHugging Face25strike20023 /VideoDRtextn<1K2 likes1.5k downloads7d agoHugging Face26BAAI /Chinese-LiPS Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides ⭐ Introduction The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios. 🚀 Dataset Details Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.audioautomatic-speech-recognition10K<n<100K12 likes1.5k downloads10mo agoHugging Face27Lixsp11 /Sekai Sekai: A Video Dataset towards World Exploration This repo contains the dataset proposed in Sekai: A Video Dataset towards World Exploration Zhen Li, Chuanhao Li, Xiaofeng Mao, Shaoheng Lin, Ming Li, Shitian Zhao, Zhaopan Xu, Xinyue Li, Yukang Feng, Jianwen Sun, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Zhixiang Wang, Yuwei Wu, Tong He, Jiangmiao Pang, Yu Qiao, Yunde Jia, Kaipeng Zhang Shanghai AI Laboratory, Beijing Institute of… See the full description on the dataset page: https://huggingface.co/datasets/Lixsp11/Sekai.texttext-to-video100K<n<1M45 likes1.4k downloads3mo agoHugging Face28imageomics /KABR Dataset Card for KABR: In-Situ Dataset for Kenyan Animal Behavior Recognition from Drone Videos Dataset Summary We present a novel high-quality dataset for animal behavior recognition from drone videos. The dataset is focused on Kenyan wildlife and contains behaviors of giraffes, plains zebras, and Grevy's zebras. The dataset consists of more than 10 hours of annotated videos, and it includes eight different classes, encompassing seven types of animal behavior and an… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/KABR.textvideo-classification1M<n<10M9 likes1.3k downloads9mo agoHugging Face29junhee1998 /gr00t-n15-robocasa-gr1-eval GR00T N1.5 on RoboCasa GR-1 Tabletop — Evaluation Trajectories Per-simulator-step recordings of 1,200 evaluation episodes (24 tasks × 50 episodes) of NVIDIA's GR00T N1.5 vision-language-action model on the RoboCasa GR-1 Tabletop Tasks benchmark. Each episode stores every low-level transition — ego-view frames, robot state, executed actions, rewards, full MuJoCo state, and 3D poses of every object and scene body — so that rollouts can be re-analyzed or re-rendered without… See the full description on the dataset page: https://huggingface.co/datasets/junhee1998/gr00t-n15-robocasa-gr1-eval.textrobotics1K<n<10K0 likes1.2k downloads21d agoHugging Face30glory-hyeok /robocurate-synth100 synth100 — 100 generated clips for validating Pre-Contact Level Filtering 100 episodes drawn (seed 20260824) from the 952-episode multi-object generation set, packaged so Stage-5 filtering can be run on them without re-deriving anything. Every input the filter needs travels with the package, in the space it is consumed in. Read section 1 before using this. The single most important fact about this data is not in the file layout, and getting it wrong invalidates any score… See the full description on the dataset page: https://huggingface.co/datasets/glory-hyeok/robocurate-synth100.imageroboticsn<1K0 likes1.2k downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.