CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MLVU /MVLUgatedMLVU: Multi-task Long Video Understanding Benchmark This repo contains the annotation data and evaluation code for the paper "MLVU: A Comprehensive Benchmark for Multi-Task Long Video Understanding". 🔔 News: 🆕 7/28/2024: The data for the MLVU-Test set has been released (🤗 Link)! The test set includes 11 different tasks, featuring our newly added Sports Question Answering (SQA, single-detail LVU) and Tutorial… See the full description on the dataset page: https://huggingface.co/datasets/MLVU/MVLU.videoquestion-answering39 likes15k downloads2y agoHugging Face02mjuicem /StreamingBench StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.imagequestion-answering1K<n<10K13 likes12k downloads1y agoHugging Face03ShareGPT4Video /ShareGPT4Video ShareGPT4Video 4.8M Dataset Card Dataset details Dataset type: ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos. It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora. sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video.imagevisual-question-answering10K<n<100K204 likes11k downloads2y agoHugging Face04ONE-Lab /GUI-World GUI-World: A Dataset for GUI-Orientated Multimodal Large Language Models Dataset: GUI-World Overview GUI-World introduces a comprehensive benchmark for evaluating MLLMs in dynamic and complex GUI environments. It features extensive annotations covering six GUI scenarios and eight types of GUI-oriented questions. The dataset assesses state-of-the-art ImageLLMs and VideoLLMs, highlighting their limitations in handling dynamic and multi-step tasks. It provides… See the full description on the dataset page: https://huggingface.co/datasets/ONE-Lab/GUI-World.videoquestion-answering10K<n<100K45 likes8.1k downloads1y agoHugging Face05yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M52 likes6.2k downloads6mo agoHugging Face06RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes5k downloads7mo agoHugging Face07plnguyen2908 /AV-SpeakerBench AV-SpeakerBench Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning. Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/ Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench Paper: https://arxiv.org/abs/2512.02231 Files test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.audioquestion-answering1K<n<10K2 likes4.1k downloads9mo agoHugging Face08yzy666 /SVBench Dataset Card for SVBench This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources. For details, see our Project, Paper and GitHub repository. Dataset Details Dataset Description SVBench is the first benchmark specifically designed to evaluate long-context streaming video understanding through temporal multi-turn question-answering (QA) chains. It addresses the limitations of existing video… See the full description on the dataset page: https://huggingface.co/datasets/yzy666/SVBench.textquestion-answering1K<n<10K7 likes3.9k downloads10mo agoHugging Face09HumanBehaviorAtlas /human_behavior_atlas Human Behavior Atlas A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features. This dataset was used to train OmniSapiens, a foundation model for social behavior processing. Papers: Human Behavior Atlas: Benchmarking Unified Psychological and… See the full description on the dataset page: https://huggingface.co/datasets/HumanBehaviorAtlas/human_behavior_atlas.textvideo-classification100K<n<1M3 likes3.5k downloads4mo agoHugging Face10ShareGPTVideo /train_video_and_instruction ShareGPTVideo Training Data All dataset and models can be found at ShareGPTVideo. Contents: Train 300k video frames: contains video frames used for SFT and DPO model, which is a subset of total 900k. ActivityNet 50k + vidal 150k + webvid 100k. Train 600k video frames: contains the rest 600k frames, the total 900k frames are used for pre-training stage. If you just do finetuning using our video QA, you can just download the 300k above. 900k composition is 400k WebVid +… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPTVideo/train_video_and_instruction.videoquestion-answering34 likes3.1k downloads2y agoHugging Face11buaaplay /SVCBench SVCBench: Streaming Video Counting Benchmark This dataset contains the clipped video segments for SVCBench, a Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance. It repositions counting as a minimal, controlled probe for diagnosing how video understanding models maintain world state along the video timeline. Project Page: https://buaa-colalab.github.io/SVCBench/ Code: https://github.com/buaa-colalab/SVCBench Dataset Description This… See the full description on the dataset page: https://huggingface.co/datasets/buaaplay/SVCBench.tabularvideo-classification1K<n<10K10 likes3k downloads3mo agoHugging Face12open-social-world /EgoNormia EgoNormia: Benchmarking Physical-Social Norm Understanding MohammadHossein Rezaei*,  Yicheng Fu*,  Phil Cuvin*,  Caleb Ziems,  Yanzhe Zhang,  Hao Zhu,  Diyi Yang,  🌎Website | 🤗 Dataset | 📄 arXiv | 📄 HF Paper EgoNormia EgoNormia is a challenging QA benchmark that tests VLMs' ability to reason over norms in context. The datset consists of 1,853 physically grounded egocentric interaction clips from Ego4D… See the full description on the dataset page: https://huggingface.co/datasets/open-social-world/EgoNormia.imagevisual-question-answering1K<n<10K7 likes2.7k downloads1y agoHugging Face13NJU-LINK /OmniVideoBenchgated OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs ✨ Overview Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction. 🎬 OmniVideoBench fills this gap.It’s a large-scale, rigorously curated… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.texttext-generation1K<n<10K5 likes2.7k downloads6mo agoHugging Face14franky-veteran /SITE-BenchThis dataset contains image and video QA test sets for SITE-Bench evaluation. imagequestion-answering1K<n<10K3 likes2.4k downloads7mo agoHugging Face15bigai-nlco /VideoHallucer VideoHallucer Paper: https://huggingface.co/papers/2406.16338 This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/bigai-nlco/VideoHallucer.textquestion-answering1K<n<10K4 likes2.2k downloads11mo agoHugging Face16SCAI-JHU /MUMA-TOM-BENCHMARK MuMA-ToM: Multi-modal Multi-Agent Theory of Mind AAAI 2025 (Oral) [🏠Homepage] [💻Code] [📝Paper] MuMA-ToM is the first multi-modal Theory of Mind benchmark designed to evaluate mental reasoning in embodied multi-agent interactions. The benchmark was designed with several key features in mind: It is factually correct, concise, and readable. It requires integrating information from multiple modalities to answer the questions. It tests understanding of multi-agent interactions… See the full description on the dataset page: https://huggingface.co/datasets/SCAI-JHU/MUMA-TOM-BENCHMARK.textquestion-answeringn<1K3 likes2.1k downloads1y agoHugging Face17Zhongzhi1228 /Terminal-Bench-Hard Terminal-Bench Hard Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks. The tasks cover software engineering, debugging, data processing, system administration, security, scientific computing, and related command-line workflows. Contents tasks/: runnable tasks in Harbor format. metadata/tasks.parquet: searchable task metadata and instructions. Each task directory contains task.toml, instruction.md, an environment/ directory, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Terminal-Bench-Hard.imagequestion-answeringn<1K0 likes1.8k downloads2mo agoHugging Face18KD-TAO /LVOmniBenchgated LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs LVOmniBench is a new audio-visual understanding evaluation benchmark in long-form audio-video inputs. 🌟 🔥 News 2026.03.19 🌟 We are very proud to launch LVOmniBench, the pioneering comprehensive evaluation benchmark of OmniLLMs in Long Audio-Video Understanding Evaluation! ✨ LVOmniBench Introduction Recent advancements in omnimodal large language models… See the full description on the dataset page: https://huggingface.co/datasets/KD-TAO/LVOmniBench.videoquestion-answeringn<1K9 likes1.8k downloads5d agoHugging Face19SII-KYW /CogStream CogStream Dataset Dataset for CogStream: Context-guided Streaming Video Question Answering. Overview CogStream is a streaming video QA dataset designed to evaluate context-guided video reasoning. Models must identify and utilize relevant historical context to answer questions about ongoing video streams. Statistics: Split Videos QA Pairs Train 852 55,623 Test 236 15,364 Total 1,088 70,987 Sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%)… See the full description on the dataset page: https://huggingface.co/datasets/SII-KYW/CogStream.tabularquestion-answering10K<n<100K1 likes1.6k downloads7mo agoHugging Face20shuzhig /nextgqa NExT-GQA (mirror) A redistribution of the NExT-GQA benchmark, packaged as a single self-contained repo (annotations + the 1,570 videos needed to run it) for convenience. This is not the official release. All credit goes to the original authors. Official code and data: https://github.com/doc-doc/NExT-GQA Dataset description NExT-GQA extends NExT-QA with temporal grounding labels: for each multiple-choice question it annotates the video segment(s) that actually… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/nextgqa.videoquestion-answering1K<n<10K0 likes1.5k downloads1mo agoHugging Face21yali30 /findingdory FindingDory: A Benchmark to Evaluate Memory in Embodied Agents Karmesh Yadav*, Yusuf Ali*, Gunshi Gupta, Yarin Gal, Zsolt Kira Current vision-language models (VLMs) struggle with long-term memory in embodied tasks. To address this, we introduce FindingDory, a benchmark in Habitat that evaluates memory-based reasoning across 60 long-horizon tasks. In this repo, we release the FindingDory Video Dataset. Each video contains images collected from a… See the full description on the dataset page: https://huggingface.co/datasets/yali30/findingdory.textquestion-answering10K<n<100K3 likes1.4k downloads1y agoHugging Face22rajjanardhan00 /Seamless_Dummy_Dataset_Fixed MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes1.3k downloads1y agoHugging Face23a8cheng /SR-3D-Bench Spatial Region 3D (SR-3D) Aware Benchmark Paper: https://arxiv.org/abs/2509.13317Project page: https://www.anjiecheng.me/sr3dCode: https://github.com/AnjieCheng/SR-3D [!IMPORTANT] [Feb. 18, 2026] UPDATE: To improve compatibility with general-purpose VLMs, the benchmark is reformulated into multiple-choice and numerical questions following the VSI-Bench evaluation protocol. Videos are annotated with set-of-marks to explicitly indicate regions. The benchmark will be compatible… See the full description on the dataset page: https://huggingface.co/datasets/a8cheng/SR-3D-Bench.textquestion-answering1K<n<10K1 likes1.3k downloads7mo agoHugging Face24ShijianW01 /VMemArena VMemArena Open-ended QA benchmark for long-video memory. 1000 questions over 561 videos (680 h of footage). Every question is free-form: there are no options to choose from. Files file contents vmemarena.json 561 video entries, each with its questions videos.jsonl per-video metadata (duration, fps, resolution, codec, size) videos/ 561 .mp4 files, named by video_id Schema // vmemarena.json — a list of video entries { "video_id":… See the full description on the dataset page: https://huggingface.co/datasets/ShijianW01/VMemArena.videovideo-text-to-textn<1K1 likes1.2k downloads9d agoHugging Face25FBK-MT /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.audioautomatic-speech-recognition1K<n<10K70 likes1.2k downloads2mo agoHugging Face26Ustiniansy /SportsTimegated SportsTime SportsTime is a long-form sports video question answering benchmark for temporal compositional reasoning, accepted to ECCV 2026. It contains 14,326 open-ended QA pairs with 50,000+ step-wise temporal evidence annotations across 1,575 videos and five team sports: basketball, American football, ice hockey, soccer, and volleyball. Dataset This Hugging Face dataset repository provides the annotation files, official train/test split, and video files.… See the full description on the dataset page: https://huggingface.co/datasets/Ustiniansy/SportsTime.textvisual-question-answering10K<n<100K1 likes1.2k downloads25d agoHugging Face27jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.1k downloads4y agoHugging Face28yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes1.1k downloads6mo agoHugging Face29RUC-NLPIR /OmniGAIA OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is designed to evaluate long-horizon, multi-hop, open-form problem solving in realistic settings rather than short perception-only QA. Benchmark Construction The OmniGAIA construction… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/OmniGAIA.audioquestion-answeringn<1K6 likes1.1k downloads7mo agoHugging Face30wangyz1999 /GameplayQA GameplayQA: A Decision-Dense POV-Synced Multi-Video Understanding Benchmark of 3D Virtual Agents Yunzhe Wang   Runhui Xu   Kexin Zheng   Tianyi Zhang Jayavibhav N. Kogundi   Soham Hans   Volkan Ustun University of Southern California ACL 2026 Corresponding Author: yunzhewa@usc.edu Overview GameplayQA is the first benchmark for POV-Synced Multi-Video Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/GameplayQA.tabularvideo-text-to-text1K<n<10K7 likes1k downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.