CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes11k downloads6h agoHugging Face02mohantesting /video-quality-scored Image-to-Video Quality-Scored Clips A collection of prompted image-to-video samples with quality-evaluation metadata. Each sample pairs a first frame (the I2V conditioning image) with one or both of: a generated video produced by a video model from the first frame + prompt an original clip (the reference/source video the prompt was authored around) A subset of the samples also carry per-clip quality scores: an overall quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.imagetext-to-video1K<n<10K0 likes6.5k downloads3mo agoHugging Face03VLABench /vlabench_composite_ft_lerobot_videotabular1M<n<10M0 likes4.2k downloads9mo agoHugging Face04VLABench /vlabench_primitive_ft_lerobot_video VLABench Primitive Tasks Dataset - LeRobot v3.0 Dataset Description This dataset is organized in the LeRobot v3.0 format and is used for integrating VLABench into the LeRobot framework officially. Compared with the v2.0 version and the RLDS version of the dataset, this release stores visual observations in a video-compressed format rather than as individual image files. This design provides significant advantages in both storage efficiency and data loading performance.… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_ft_lerobot_video.tabularrobotics100K<n<1M3 likes3.8k downloads5mo agoHugging Face05juyil /AVQA-videos AVQA — Audio-Visual Question Answering (videos + annotations) A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life audio-visual question answering over short in-the-wild clips. The original release ships only the QA annotations and expects users to collect the source videos from VGGSound themselves. This repository bundles the source video clips together with the official train/val annotations, so the dataset is usable without any YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.tabularvisual-question-answering10K<n<100K1 likes2.8k downloads4mo agoHugging Face06andrewt28 /keystroke-typing-videos Keystroke Typing Videos of Reuters Recordings of typing randomly sampled sentences (<= 150 characters) from nltk Reuters dataset. Keystroke data is provided too. tabularvideo-text-to-textn<1K0 likes2.8k downloads1y agoHugging Face07minkyuchoi /Temporal-Logic-Video-Dataset Temporal Logic Video (TLV) Dataset Temporal Logic Video (TLV) Dataset Synthetic and real video dataset with temporal logic annotation Explore the GitHub » NSVS-TL Project Webpage · NSVS-TL Source Code Overview The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components: Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.tabularquestion-answeringn<1K1 likes2.4k downloads2y agoHugging Face08larrylarrylarry1 /VideoArtifactDetectionVideo Artifact Detection dataset using source videos from LongVideoBench. Video level labels are given in labels.csv (training) and labels_test.csv (testing). Localized artifact regions for burst artifacts are given in the artifact_ranges column. Note that labels.csv contains additional source videos from LongVideoBench that are not included in this repository. You may visit the LongVideoBench page for the additional videos. For additional questions, please email palmerla@usc.edu. tabular1K<n<10K0 likes2.4k downloads4mo agoHugging Face09lerobot /video-benchmark-resultstabular10K<n<100K2 likes1.6k downloads2mo agoHugging Face10beingbetter11643 /PaSBench-Video PaSBench-Video PaSBench-Video is an evaluation-only video benchmark for proactive safety warning in streaming video. The task is to watch a video progressively and decide whether a safety warning should be issued before an adverse event happens. A good system should warn after the risk becomes visually clear, warn early enough to be useful, explain the risk, recommend an action, and stay silent on no-risk videos. This repository contains the test/evaluation set only; it is not a… See the full description on the dataset page: https://huggingface.co/datasets/beingbetter11643/PaSBench-Video.tabularn<1K5 likes1.3k downloads23d agoHugging Face11Rapidata /text-2-video-human-preferences Rapidata Video Generation Preference Dataset This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set: Sora Hunyouan Pika 2.0 Runway ML Alpha Luma Ray 2 Explore our latest model rankings on our website. If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.imagetext-to-video1K<n<10K21 likes1.2k downloads2y agoHugging Face12Rapidata /text-2-video-human-preferences-wan2.1 Rapidata Video Generation Alibaba Wan2.1 Human Preference If you get value from this dataset and would like to see more in the future, please consider liking it. This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Overview In this dataset, ~45'000 human annotations were collected to evaluate Alibaba Wan 2.1 video generation model on our benchmark. The up to date benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-wan2.1.imagevideo-classificationn<1K20 likes1.2k downloads2y agoHugging Face13yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes1.1k downloads6mo agoHugging Face14TIGER-Lab /VideoFeedback📃Paper | 🌐Website | 💻Github | 🛢️Datasets | 🤗Model | 🤗Demo Overview VideoFeedback contains a total of 37.6K text-to-video pairs from 11 popular video generative models, with some real-world videos as data augmentation. The videos are annotated by raters for five evaluation dimensions: Visual Quality, Temporal Consistency, Dynamic Degree, Text-to-Video Alignment and Factual Consistency, in 1-4 scoring scale. VideoFeedback is used to for trainging of VideoScore Below we… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VideoFeedback.tabularvideo-classification10K<n<100K35 likes1k downloads2y agoHugging Face15blaccastro /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/blaccastro/tiktok-videos-4b.tabulartext-classification1B<n<10B2 likes891 downloads14d agoHugging Face16allenai /Molmo2-VideoPoint Molmo2-VideoPoint Molmo2-VideoPoint is a dataset of video pointing data collected from human annotators. It can be used to fine-tune vision-language models for video grounding by pointing. Molmo2-VideoPoint is part of the Molmo2 dataset collection and was used to train the Molmo2 family of models. Quick links: 📃 Paper 🎥 Blog with Videos Usage from datasets import load_dataset # Load entire dataset ds = load_dataset("allenai/Molmo2-VideoPoint", split="train") #… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-VideoPoint.tabular1M<n<10M10 likes754 downloads6mo agoHugging Face17Rapidata /text-2-video-human-preferences-seedance-1-pro Rapidata Video Generation Seedance 1 Pro Human Preference In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.imagevideo-classification1K<n<10K9 likes681 downloads1y agoHugging Face18Yanbin99 /Depth-Normal-Videos-42K Depth and Normal Videos Dataset 42,498 videos with depth and surface normals. Usage from huggingface_hub import hf_hub_download video = hf_hub_download( repo_id="Yanbin99/Depth-Normal-Videos-42K", filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4", repo_type="dataset" ) tabulardepth-estimation10K<n<100K1 likes680 downloads9mo agoHugging Face19DogNeverSleep /MME-VideoOCR_DatasetThis dataset is from the paper MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios. See https://github.com/FrankYang-17/MME-VideoOCR for more information. License Our MME-VideoOCR is released as CC-BY-NC4.0. The video samples are collected from a publicly available resource. Note MME-VideoOCR is only used for academic research. Commercial use in any form is prohibited. The copyright of all videos belongs to the video owners. If there… See the full description on the dataset page: https://huggingface.co/datasets/DogNeverSleep/MME-VideoOCR_Dataset.tabularvideo-text-to-text1K<n<10K3 likes673 downloads1y agoHugging Face20facebook /PLM-VideoBench Dataset Summary PLM-VideoBench is a collection of human-annotated resources for evaluating Vision Language models, focused on detailed video understanding. [📃 Tech Report] [📂 Github] Supported Tasks PLM-VideoBench includes evaluation data for the following tasks: FGQA In this task, a model must answer a multiple-choice question (MCQ) that probes fine-grained activity understanding. Given a question and multiple options that differ in a… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PLM-VideoBench.tabularmultiple-choice10K<n<100K13 likes646 downloads1y agoHugging Face21dghadiya /TAG-Bench-Video License The TAG-Bench dataset (generated videos + human ratings) is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. If you use this dataset, please cite our paper. TAG-Bench: Generative Action Tell-Tales Project page: https://xthomasbu.github.io/video-gen-evals/ TAG-Bench is a benchmark for human motion realism in video generative models. It consists of 828 generated video clips of human actions, together with human ratings collected… See the full description on the dataset page: https://huggingface.co/datasets/dghadiya/TAG-Bench-Video.tabularothern<1K1 likes646 downloads5mo agoHugging Face22kwakuobeng /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/kwakuobeng/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes637 downloads13d agoHugging Face23Gokalp35 /VideoGameSalestabular10K<n<100K0 likes585 downloads2y agoHugging Face24dams2005 /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/dams2005/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes585 downloads20d agoHugging Face25facebook /PLM-Video-Human Dataset Card for PLM-Video Human PLM-Video-Human is a collection of human-annotated resources for training Vision Language Models, focused on detailed video understanding. Training tasks include: fine-grained open-ended question answering (FGQA), Region-based Video Captioning (RCap), Region-based Dense Video Captioning (RDCap) and Region-based Temporal Localization (RTLoc). [📃 Tech Report] [📂 Github] Dataset Structure Fine-Grained Question Answering (FGQA)… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PLM-Video-Human.tabularmultiple-choice1M<n<10M29 likes562 downloads1y agoHugging Face26Rapidata /text-2-video-human-preferences-veo3 Rapidata Video Generation Veo 3 Human Preference In this dataset, ~46k human responses from ~20k human annotators were collected to evaluate Veo3 video generation model on our benchmark. This dataset was collected in roughly 35 minutes using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.imagevideo-classification1K<n<10K20 likes561 downloads1y agoHugging Face27KOM-00 /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/KOM-00/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes537 downloads12d agoHugging Face28KlingTeam /VideoGen-RewardBench 🏆 [VideoGen-RewardBench Leaderboard] Introduction VideoGen-RewardBench is a comprehensive benchmark designed to evaluate the performance of video reward models on modern text-to-video (T2V) systems. Derived from the third-party VideoGen-Eval (Zeng et.al, 2024), we constructing 26.5k (prompt, Video A, Video B) triplets and employing expert annotators to provide pairwise preference labels. These annotations are based on key evaluation dimensions—Visual Quality (VQ), Motion… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/VideoGen-RewardBench.tabular10K<n<100K8 likes516 downloads2y agoHugging Face29exiawsh /videobench Dataset Card for LongVideoBench Large multimodal models (LMMs) are handling increasingly longer and more complex inputs. However, few public benchmarks are available to assess these advancements. To address this, we introduce LongVideoBench, a question-answering benchmark with video-language interleaved inputs up to an hour long. It comprises 3,763 web-collected videos with subtitles across diverse themes, designed to evaluate LMMs on long-term multimodal understanding. The… See the full description on the dataset page: https://huggingface.co/datasets/exiawsh/videobench.tabularmultiple-choice1K<n<10K0 likes504 downloads1y agoHugging Face30trojblue /test-HunyuanVideo-pixelart-videos trojblue/test-HunyuanVideo-pixelart-images 👋 Heads up—this repository is just a PARTIAL dataset. For the full pixelart-images dataset, make sure to grab both parts: Images Part Video Part (this repo) What's in the Dataset? This dataset is all about anime-styled pixel art images that have been carefully selected to make your models shine. Here’s what makes these images special: Rich in detail: Pixelated, yes—but still full of life and not overly simplified.… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-videos.tabulartext-to-imagen<1K8 likes493 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.