CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MCG-NJU /VideoChat3-LV116k VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video data with supervision over longer temporal contexts, where evidence can be sparse, delayed, and distributed across multiple video segments. The dataset is constructed through a long-video synthesis pipeline. Candidate long videos are filtered for visual quality, semantic content, and temporal coherence. Videos are then split into manageable… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k.textvideo-text-to-text1K<n<10K15 likes22k downloads2mo agoHugging Face02Vividbot /vivid-video-instructtext10K<n<100K2 likes7k downloads2y agoHugging Face03mohantesting /video-quality-scored Image-to-Video Quality-Scored Clips A collection of prompted image-to-video samples with quality-evaluation metadata. Each sample pairs a first frame (the I2V conditioning image) with one or both of: a generated video produced by a video model from the first frame + prompt an original clip (the reference/source video the prompt was authored around) A subset of the samples also carry per-clip quality scores: an overall quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.imagetext-to-video1K<n<10K0 likes6.5k downloads3mo agoHugging Face04artificialguybr /veo3-video-prompts Veo 3 Video Generation Dataset English | Português do Brasil English Summary A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant. Videos: 5,811 Input images: 1,354 Configurations: 6 Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.imagetext-to-video1K<n<10K0 likes5.3k downloads1mo agoHugging Face05minuzero /VideoKR-Train VideoKR-Train 📄 ArXiv &nbsp;|&nbsp; 💻 Code &nbsp;|&nbsp; 🤗 Collection About This repository contains the VideoKR training data presented in VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding (ICML 2026 Spotlight). VideoKR is the first large-scale training corpus specifically designed for knowledge- and reasoning-intensive video understanding. It contains 315K video reasoning examples over 145K newly collected, CC-licensed… See the full description on the dataset page: https://huggingface.co/datasets/minuzero/VideoKR-Train.textvisual-question-answering100K<n<1M2 likes3k downloads2mo agoHugging Face06juyil /AVQA-videos AVQA — Audio-Visual Question Answering (videos + annotations) A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life audio-visual question answering over short in-the-wild clips. The original release ships only the QA annotations and expects users to collect the source videos from VGGSound themselves. This repository bundles the source video clips together with the official train/val annotations, so the dataset is usable without any YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.tabularvisual-question-answering10K<n<100K1 likes2.9k downloads4mo agoHugging Face07minkyuchoi /Temporal-Logic-Video-Dataset Temporal Logic Video (TLV) Dataset Temporal Logic Video (TLV) Dataset Synthetic and real video dataset with temporal logic annotation Explore the GitHub » NSVS-TL Project Webpage · NSVS-TL Source Code Overview The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components: Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.tabularquestion-answeringn<1K1 likes2.4k downloads2y agoHugging Face08MCG-NJU /VideoChat3-Academic2M VideoChat3-Academic2M VideoChat3-Academic2M is the academic video instruction data used by VideoChat3. It re-annotates public academic video datasets for video captioning, video question answering, and fine-grained motion understanding. The dataset follows an evidence-grounded annotation enhancement pipeline. Short answers, option-only labels, and concise captions are rewritten into richer instruction-following responses that mention visible objects, actions, scenes, temporal… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-Academic2M.textvideo-text-to-text10K<n<100K24 likes2.2k downloads2mo agoHugging Face09Video-Reason /VBVR-Bench-Data VBVR: A Very Big Video Reasoning Suite Overview Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over motion, interaction, and causality. Rapid progress in video models has focused primarily on visual quality. Systematically studying video reasoning and its scaling behavior suffers from a lack of… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Bench-Data.imagen<1K9 likes2.2k downloads6mo agoHugging Face10bigai-nlco /VideoHallucer VideoHallucer Paper: https://huggingface.co/papers/2406.16338 This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/bigai-nlco/VideoHallucer.textquestion-answering1K<n<10K4 likes2.2k downloads11mo agoHugging Face11anonymousasdf /video2mentaltext10K<n<100K0 likes2k downloads5mo agoHugging Face12wofmanaf /ego4d-videoEgoCOT is a large-scale embodied planning dataset, which selected egocentric videos from the Ego4D dataset and corresponding high-quality step-by-step language instructions, which are machine generated, then semantics-based filtered, and finally human-verified. For mored details, please visit EgoCOT_Dataset. If you find this dataset useful, please consider citing the paper, @article{mu2024embodiedgpt, title={Embodiedgpt: Vision-language pre-training via embodied chain of thought}… See the full description on the dataset page: https://huggingface.co/datasets/wofmanaf/ego4d-video.textquestion-answering100K<n<1M16 likes1.5k downloads2y agoHugging Face13Alexislhb /Video-IFBench Video-IFBench This release contains the evaluation split used for the Video-IFBench main experiments. Paper: Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Project page: https://alexios-hub.github.io/Video-IFBench/ Code: https://github.com/Alexios-hub/Video-IFBench textvisual-question-answeringn<1K1 likes1.3k downloads26d agoHugging Face14Enxin /Video-MMLU Video-MMLU Benchmark Resources Website arXiv: Paper GitHub: Code Huggingface: Video-MMLU Benchmark Features Benchmark Collection and Processing Video-MMLU specifically targets videos that focus on theorem demonstrations and probleming-solving, covering mathematics, physics, and chemistry. The videos deliver dense information through numbers and formulas, pose significant challenges for video LMMs in dynamic OCR… See the full description on the dataset page: https://huggingface.co/datasets/Enxin/Video-MMLU.textvideo-text-to-text1K<n<10K13 likes941 downloads1y agoHugging Face15kolerk /Video_Reality_Test VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization. Benchmark Structure This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation: real_hard: 100 samples.… See the full description on the dataset page: https://huggingface.co/datasets/kolerk/Video_Reality_Test.texttext-to-videon<1K10 likes763 downloads5mo agoHugging Face16wchai /Video-Detailed-Caption Video Detailed Caption Benchmark Resources Website arXiv: Paper GitHub: Code Huggingface: AuroraCap Model Huggingface: VDC Benchmark Huggingface: Trainset Features Benchmark Collection and Processing We building VDC upon Panda-70M, Ego4D, Mixkit, Pixabay, and Pexels. Structured detailed captions construction pipeline. We develop a structured detailed captions construction pipeline to generate extra detailed descriptions from various… See the full description on the dataset page: https://huggingface.co/datasets/wchai/Video-Detailed-Caption.textvideo-text-to-text1K<n<10K17 likes744 downloads2y agoHugging Face17MBZUAI /VideoGPT-plus_Training_Datasettext100K<n<1M8 likes684 downloads2y agoHugging Face18anonymous-video-benchmark /toc_bench TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models TOC-Bench is a diagnostic benchmark for evaluating whether Video Large Language Models maintain object identity, state, persistence, and temporal relations throughout a video. It focuses on object-centric phenomena including occlusion, disappearance, reappearance, repeated events, event order, temporal location, duration, conditional state, and relative movement. Anonymous-review notice. This… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-video-benchmark/toc_bench.textvideo-text-to-text1K<n<10K0 likes657 downloads2mo agoHugging Face19sfsdfsafsddsfsdafsa /Long-video-test-datatextn<1K2 likes646 downloads3y agoHugging Face20LIMinghan /FiVE-Fine-Grained-Video-Editing-Benchmark FiVE-Bench FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Diffusion and Rectified Flow Models Minghan Li1*, Chenxi Xie2*, Yichen Wu13, Lei Zhang2, Mengyu Wang1† 1Harvard University 2The Hong Kong Polytechnic University 3City University of Hong Kong *Equal contribution †Corresponding Author 💜 Leaderboard (coming soon)   |   💻 GitHub   |   🤗 Hugging Face   📝 Project Page   |   📰 Paper   |   🎥 Video Demo   FiVE is a benchmark comprising 100 videos for… See the full description on the dataset page: https://huggingface.co/datasets/LIMinghan/FiVE-Fine-Grained-Video-Editing-Benchmark.imagetext-to-videon<1K5 likes579 downloads1y agoHugging Face21videoSALMONN2 /video-SALMONN_2_testset video-SALMONN 2 Benchmark Generate the caption corresponding to the video and the audio with video_salmonn2_test.json Organize your results in the format like the following example: [ { "id": ["0.mp4"], "pred": "Generated Caption" } ] Replace res_file in eval.py with your result file. Run python3 eval.pytextn<1K3 likes558 downloads1y agoHugging Face22PediaMedAI /Social-IQ-Video Copy of Social-IQ 2.0 Challenge We are hiring collaborators to organize a similar challenge like Social-IQ 2.0. If you are interested in it, please contact us via xucao@pediamed.ai. text1K<n<10K3 likes546 downloads1y agoHugging Face23osazuwa /2d_dungeon_flier_video_balanced 2D Dungeon Flier Video: Balanced Causal Splits This dataset is a split-safe, balanced augmentation of osazuwa/2d_dungeon_flier_video. It reuses all 10,000 source episodes exactly once and adds 3,100 episodes from the same simulator. There is no clip overlap across splits. Each episode is a 14-second MP4 with 140 frames at 10 FPS and a stored resolution of 900 x 540 pixels. Matching NPZ files contain the nine-variable causal trace, action tokens, and intervention encoding. Every… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/2d_dungeon_flier_video_balanced.tabular10K<n<100K0 likes446 downloads1mo agoHugging Face24ziweix /Video_Reality_Test Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans? This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization. Benchmark Structure This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:… See the full description on the dataset page: https://huggingface.co/datasets/ziweix/Video_Reality_Test.texttext-to-videon<1K0 likes403 downloads7mo agoHugging Face25Dii2 /Video_Reality_Test Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans? This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization. Benchmark Structure This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:… See the full description on the dataset page: https://huggingface.co/datasets/Dii2/Video_Reality_Test.texttext-to-videon<1K0 likes399 downloads9mo agoHugging Face26VideoSearchR1 /charades-stage1_data VideoSearch-R1 Charades-STA This repository contains the prepared Charades-STA artifacts used by VideoSearch-R1. Paper: VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement Project Page: https://mlvlab.github.io/VideoSearch-R1/ Repository: https://github.com/mlvlab/VideoSearch-R1 Stage 1 Cold Start SFT Charades_STA_Stage1_ColdStart.jsonl is the Stage 1 cold-start SFT dataset used to train the VideoSearch-R1 verifier/reasoner on… See the full description on the dataset page: https://huggingface.co/datasets/VideoSearchR1/charades-stage1_data.texttext-retrieval1K<n<10K0 likes392 downloads3mo agoHugging Face27minuzero /VideoKR-Eval VideoKR-Eval 📄 ArXiv &nbsp;|&nbsp; 💻 Code &nbsp;|&nbsp; 🤗 Collection About This repository contains the VideoKR-Eval benchmark presented in VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding (ICML 2026 Spotlight). VideoKR-Eval is an expert-annotated evaluation benchmark for knowledge- and reasoning-intensive video understanding. Unlike existing benchmarks where a substantial fraction of questions can be answered from a single… See the full description on the dataset page: https://huggingface.co/datasets/minuzero/VideoKR-Eval.textvisual-question-answering1K<n<10K0 likes352 downloads4mo agoHugging Face28zengziyun /VideoArgusBench VideoArgusBench Sample-specific rubric benchmark for conditioned video generation. VideoArgusBench is the evaluation benchmark for VideoArgus, a framework that scores a generated video against a rubric written for that specific prompt rather than a fixed global metric. This dataset ships the inputs (conditioning assets + prompts) and, for each input, a rubric. It does not contain generated videos — you bring your own model's outputs and score them with the VideoArgus evaluation… See the full description on the dataset page: https://huggingface.co/datasets/zengziyun/VideoArgusBench.imagetext-to-video1K<n<10K1 likes348 downloads2mo agoHugging Face29DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes320 downloads2y agoHugging Face30MBZUAI /VideoInstruct-100KVideoInstruct100K is a high-quality video conversation dataset generated using human-assisted and semi-automatic annotation techniques. The question answers in the dataset are related to, Video Summariazation Description-based question-answers (exploring spatial, temporal, relationships, and reasoning concepts) Creative/generative question-answers For mored details, please visit Oryx/VideoChatGPT/video-instruction-data-generation. If you find this dataset useful, please consider citing the… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/VideoInstruct-100K.text100K<n<1M49 likes315 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.