CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ShareGPT4Video /ShareGPT4Video ShareGPT4Video 4.8M Dataset Card Dataset details Dataset type: ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos. It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora. sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video.imagevisual-question-answering10K<n<100K204 likes11k downloads2y agoHugging Face02buaaplay /SVCBench SVCBench: Streaming Video Counting Benchmark This dataset contains the clipped video segments for SVCBench, a Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance. It repositions counting as a minimal, controlled probe for diagnosing how video understanding models maintain world state along the video timeline. Project Page: https://buaa-colalab.github.io/SVCBench/ Code: https://github.com/buaa-colalab/SVCBench Dataset Description This… See the full description on the dataset page: https://huggingface.co/datasets/buaaplay/SVCBench.tabularvideo-classification1K<n<10K10 likes3k downloads3mo agoHugging Face03franky-veteran /SITE-BenchThis dataset contains image and video QA test sets for SITE-Bench evaluation. imagequestion-answering1K<n<10K3 likes2.4k downloads7mo agoHugging Face04bigai-nlco /VideoHallucer VideoHallucer Paper: https://huggingface.co/papers/2406.16338 This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/bigai-nlco/VideoHallucer.textquestion-answering1K<n<10K4 likes2.2k downloads11mo agoHugging Face05a8cheng /SR-3D-Bench Spatial Region 3D (SR-3D) Aware Benchmark Paper: https://arxiv.org/abs/2509.13317Project page: https://www.anjiecheng.me/sr3dCode: https://github.com/AnjieCheng/SR-3D [!IMPORTANT] [Feb. 18, 2026] UPDATE: To improve compatibility with general-purpose VLMs, the benchmark is reformulated into multiple-choice and numerical questions following the VSI-Bench evaluation protocol. Videos are annotated with set-of-marks to explicitly indicate regions. The benchmark will be compatible… See the full description on the dataset page: https://huggingface.co/datasets/a8cheng/SR-3D-Bench.textquestion-answering1K<n<10K1 likes1.3k downloads7mo agoHugging Face06Ustiniansy /SportsTimegated SportsTime SportsTime is a long-form sports video question answering benchmark for temporal compositional reasoning, accepted to ECCV 2026. It contains 14,326 open-ended QA pairs with 50,000+ step-wise temporal evidence annotations across 1,575 videos and five team sports: basketball, American football, ice hockey, soccer, and volleyball. Dataset This Hugging Face dataset repository provides the annotation files, official train/test split, and video files.… See the full description on the dataset page: https://huggingface.co/datasets/Ustiniansy/SportsTime.textvisual-question-answering10K<n<100K1 likes1.2k downloads25d agoHugging Face07jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.1k downloads4y agoHugging Face08michalsr /molmo2-moments Molmo-2 Moments (M2M) Long-video QA dataset where every question is anchored to a specific [start, end] clip interval in seconds. Released alongside the ToolMerge paper, "Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval". ⚠️ Source videos & ownership The videos/*.mp4 files in this repository were collected from YouTube. We do not own these videos and claim no copyright over them. All rights to the video content remain with the original… See the full description on the dataset page: https://huggingface.co/datasets/michalsr/molmo2-moments.tabularvideo-text-to-text10K<n<100K0 likes298 downloads2mo agoHugging Face09KamiKrafton /agentvidbench AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline) by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI) Layout . ├── README.md ├── questions.jsonl # 100 rows — one per question ├── videos.jsonl # 71 rows — one per unique video ├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench.textvideo-text-to-textn<1K0 likes297 downloads2mo agoHugging Face10rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_3 MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes259 downloads1y agoHugging Face11iesc /Ava-100 Empowering Agentic Video Analytics Systems with Video Language Models [🖥️ Project Code] [📖 arXiv Paper] [📊 Dataset] Introduction AVA-100 is an ultra-long video benchmark specially designed to evaluate video analysis capabilities Avas-100 consists of 8 videos, each exceeding 10 hours in length, and includes a total of 120 manually annotated questions. The benchmark covers four typical video analytics scenarios: human daily activities, city walking, wildlife… See the full description on the dataset page: https://huggingface.co/datasets/iesc/Ava-100.textmultiple-choicen<1K2 likes218 downloads11mo agoHugging Face12ov015 /VideoHallucer VideoHallucer Paper: https://huggingface.co/papers/2406.16338 This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/ov015/VideoHallucer.textquestion-answering1K<n<10K0 likes201 downloads5mo agoHugging Face13inesriahi /valor32k-avqa-v2 Valor32k-AVQA v2.0 Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position. Links Paper: ACM Digital Library Project page: inesriahi.github.io/valor32k-avqa-2 Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.tabularquestion-answering100K<n<1M0 likes198 downloads3mo agoHugging Face14agentvidbench /agentvidbench AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline) by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI) Layout . ├── README.md ├── questions.jsonl # 100 rows — one per question ├── videos.jsonl # 71 rows — one per unique video ├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/agentvidbench/agentvidbench.textvideo-text-to-textn<1K0 likes192 downloads2mo agoHugging Face15Inst-IT /Inst-It-Dataset Inst-IT Dataset: An Instruction Tuning Dataset with Multi-level Fine-Grained Annotations introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning 🌐 Homepage | Code | 🤗 Paper | 📖 arXiv Inst-IT Dataset Overview We create a large-scale instruction tuning dataset, the Inst-it Dataset. To the best of our knowledge, this is the first dataset that provides fine-grained annotations centric on specific… See the full description on the dataset page: https://huggingface.co/datasets/Inst-IT/Inst-It-Dataset.textquestion-answering10K<n<100K10 likes167 downloads2y agoHugging Face16shuaishuaicdp /OmniCoding OmniCoding A multimodal terminal-tool-use SFT/RL dataset. Each record is a question + verifiable answer + media (video/audio/image) — the target agent is expected to operate on the media via a Linux terminal (ffmpeg, ffprobe, whisper, python, etc.) rather than a GUI. Aggregated and filtered from four upstream sources, with a single unified schema, global dedup, and category-balanced sampling. Records Source n Omnimodal-Agent-SFT-2K (RUC-NLPIR) — agentic… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/OmniCoding.textquestion-answering10K<n<100K0 likes158 downloads2mo agoHugging Face17LIQIIIII /ViMU ViMU: Benchmarking Video Metaphorical Understanding Qi Li, Xinchao Wang* *Corresponding author xML Lab, National University of Singapore Our GitHub repository contains the evaluation scripts for ViMU, a benchmark for video metaphorical understanding. The code evaluates multimodal models on four tasks: Open-ended interpretation (OE) Evidence grounding (EG) Rhetoric mechanism identification (RM) Social value signal identification (SV) Directory Structure Expected… See the full description on the dataset page: https://huggingface.co/datasets/LIQIIIII/ViMU.imagevisual-question-answering1K<n<10K6 likes137 downloads4mo agoHugging Face18QLGalaxy /VUDG VUDG: A Dataset for Video Understanding Domain Generalization VUDG is a benchmark dataset for evaluating domain generalization (DG) in video understanding. It contains 7,899 video clips and 36,388 high-quality QA pairs, covering 11 diverse visual domains, such as cartoon, egocentric, surveillance, rainy, snowy, etc. Each video is annotated with both multiple-choice and open-ended question-answer pairs, designed via a multi-expert progressive annotation pipeline using large… See the full description on the dataset page: https://huggingface.co/datasets/QLGalaxy/VUDG.textquestion-answering10K<n<100K3 likes132 downloads8mo agoHugging Face19Yuan4629 /WereBench Anonymization For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information. WereBench WereBench is a benchmark dataset for evaluating language models in the Werewolf (similar to Mafia) social deduction setting. It focuses on human‑aligned strategic reasoning rather than only coarse metrics (e.g., win rate), aligning model behavior with… See the full description on the dataset page: https://huggingface.co/datasets/Yuan4629/WereBench.tabularquestion-answeringn<1K1 likes127 downloads9mo agoHugging Face20KamiKrafton /agentvidbench-sample AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above. Sample selection The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench-sample.textvideo-text-to-textn<1K0 likes120 downloads2mo agoHugging Face21n0nam4 /WereBench Anonymization For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information. WereBench WereBench is a benchmark dataset for evaluating language models in the Werewolf (similar to Mafia) social deduction setting. It focuses on human‑aligned strategic reasoning rather than only coarse metrics (e.g., win rate), aligning model behavior… See the full description on the dataset page: https://huggingface.co/datasets/n0nam4/WereBench.tabularquestion-answeringn<1K0 likes110 downloads9mo agoHugging Face22meituan-longcat /MineExplorer MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft Tianjie Ju · Yueqing Sun · Zheng Wu · Wei Zhang · Yaqi Huo · Xi Su · Qi Gu · Xunliang Cai · Gongshen Liu · Zhuosheng Zhang Abstract Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic open worlds remains unclear. Existing embodied and… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/MineExplorer.textreinforcement-learningn<1K3 likes107 downloads3mo agoHugging Face23LordUky /EMCompressEMCompress A Benchmark for Endomorphic Multimodal Compression on Long Cooking Videos 📰 News 2026.05   🎉 Dataset + reproduction code released on HuggingFace & GitHub. 2026.04   📝 Paper accepted to ACL 2026 Findings. 🧠 About EMCompress is the first benchmark dedicated to evaluating the Endomorphic Multimodal Compression (EMC) task: an endomorphic transformation F_EMC : (V, Q) → (v, q) that compresses a (video, question) pair into a shorter… See the full description on the dataset page: https://huggingface.co/datasets/LordUky/EMCompress.textvideo-text-to-text1K<n<10K2 likes93 downloads4mo agoHugging Face24iLearn-Lab /FineBadmintonBenchmark FineBadmintonBenchmark Fine-grained badminton video question answering benchmark. Dataset Structure hf_video_clips_qa/: one clip per QA item, named by video_uid (for example video_000001.mp4). finebadmintonbenchmark/: annotation JSON files. each item contains video_uid each QA item maps to exactly one video clip through video_uid Citation @inproceedings{he2025finebadminton, title={Finebadminton: A multi-level dataset for fine-grained badminton video… See the full description on the dataset page: https://huggingface.co/datasets/iLearn-Lab/FineBadmintonBenchmark.textquestion-answering1K<n<10K0 likes66 downloads7mo agoHugging Face25KHUjongseo /E2E_real_object E2E Real Object Direction A video-based benchmark for evaluating VideoLLMs' directional reasoning and object recognition on real-world objects. Conditions Condition Question Answer Purpose direction_only "In which direction is the object moving?" "Up" Baseline direction recognition direction_obj_in_q "In which direction is the car moving?" "Up" Does naming the object help? direction_obj_in_a "In which direction is the object moving?" "The car is moving up"… See the full description on the dataset page: https://huggingface.co/datasets/KHUjongseo/E2E_real_object.textvideo-classificationn<1K0 likes46 downloads7mo agoHugging Face26EgoLink /EgoLink2026 EgoLink2026 EgoLink 2026 evaluation benchmark for egocentric social reasoning. Dataset Structure Eval.jsonl: 4030 multiple-choice evaluation questions data/: 1055 egocentric video files referenced by the path field in each sample textquestion-answering1K<n<10K0 likes42 downloads3mo agoHugging Face27ContinuousPerceptionResearch /CP-Bench Continuous Perception Benchmark (CP-Bench) Overview The Continuous Perception Benchmark (CP-Bench) is a diagnostic dataset designed to evaluate whether modern vision-language and multimodal models can integrate continuous visual information over time—an ability that is central to human visual perception but largely absent in contemporary architectures. Inspired by the continuous, stream-based nature of human vision, CP-Bench isolates the core requirement of maintaining… See the full description on the dataset page: https://huggingface.co/datasets/ContinuousPerceptionResearch/CP-Bench.textquestion-answering1K<n<10K1 likes40 downloads10mo agoHugging Face28agentvidbench /agentvidbench-sample AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above. Sample selection The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/agentvidbench/agentvidbench-sample.textvideo-text-to-textn<1K0 likes40 downloads2mo agoHugging Face29UBC-ViL /BlackSwanSuite-MCQgated Black Swan Suite Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events Aditya Chinchure*, Sahithya Ravi*, Raymond Ng, Vered Shwartz, Boyang Li, Leonid Sigal (* equal) 🎉 Accepted at CVPR 2025 arXiv | Website Dataset Information Black Swan has three variants of questions. Please find the data in the appropriate repositories: BlackSwanSuite-Gen -- link BlackSwanSuite-MCQ (this) BlackSwanSuite-YN -- link This dataset contains questions for MCQ… See the full description on the dataset page: https://huggingface.co/datasets/UBC-ViL/BlackSwanSuite-MCQ.tabularvisual-question-answering1K<n<10K2 likes39 downloads2y agoHugging Face30KHUjongseo /E2E_20K_test E2E 20K Synthetic Direction A large-scale synthetic video benchmark for evaluating VideoLLMs' directional reasoning.20,000 videos of colored geometric shapes moving in four cardinal directions, with automatically generated MCQ annotations. Directions Direction Description up Object moves toward the top of the frame down Object moves toward the bottom of the frame left Object moves toward the left of the frame right Object moves toward the right of the… See the full description on the dataset page: https://huggingface.co/datasets/KHUjongseo/E2E_20K_test.textvideo-classificationn<1K0 likes38 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.