CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01minkyuchoi /Temporal-Logic-Video-Dataset Temporal Logic Video (TLV) Dataset Temporal Logic Video (TLV) Dataset Synthetic and real video dataset with temporal logic annotation Explore the GitHub » NSVS-TL Project Webpage · NSVS-TL Source Code Overview The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components: Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.tabularquestion-answeringn<1K1 likes2.4k downloads2y agoHugging Face02bigai-nlco /VideoHallucer VideoHallucer Paper: https://huggingface.co/papers/2406.16338 This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/bigai-nlco/VideoHallucer.textquestion-answering1K<n<10K4 likes2.1k downloads11mo agoHugging Face03yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes1.1k downloads7mo agoHugging Face04wofmanaf /ego4d-videoEgoCOT is a large-scale embodied planning dataset, which selected egocentric videos from the Ego4D dataset and corresponding high-quality step-by-step language instructions, which are machine generated, then semantics-based filtered, and finally human-verified. For mored details, please visit EgoCOT_Dataset. If you find this dataset useful, please consider citing the paper, @article{mu2024embodiedgpt, title={Embodiedgpt: Vision-language pre-training via embodied chain of thought}… See the full description on the dataset page: https://huggingface.co/datasets/wofmanaf/ego4d-video.textquestion-answering100K<n<1M16 likes667 downloads2y agoHugging Face05OpenMOSS-Team /VideoThinkBench [CVPR 2026] Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm 🎊 News [2026.02] 🔥🔥Our work has been accepted by CVPR 2026! 🎉🎉🎉 [2025.11] Our paper "Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm" has been released on arXiv! 📄 [Paper] On HuggingFace, it has achieved "#1 Paper of the Day"! [2025.11] 🔥We release "minitest" of our VideoThinkBench, including 500… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/VideoThinkBench.imagetext-to-video1K<n<10K19 likes605 downloads2mo agoHugging Face06ShareGPTVideo /train_raw_video ShareGPTVideo Raw ActivityNet Videos for Train data All dataset and models can be found at ShareGPTVideo. Contents: Due to our scene split, we provide our processed activityNet videos corresponding to test frames in train video frames the processing script is process_activitynet.py textquestion-answering10K<n<100K2 likes323 downloads2y agoHugging Face07DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes316 downloads2y agoHugging Face08MMInstruction /Video-T3-QATextual Temporal Understanding Dataset Temporal Reasoning Transfer from Text to Video, ICLR 2025 Project Page: https://video-t3.github.io/ In each json file, we provide LLaVA-style text QA samples, using the synthesization method described in our paper. For example: [ { "from": "human", "value": "Based on the following captions describing keyframes of a video, answer the next question.\n\nCaptions:\nThe image displays a circular emblem with a metallic appearance, conveying a… See the full description on the dataset page: https://huggingface.co/datasets/MMInstruction/Video-T3-QA.textquestion-answering100K<n<1M2 likes222 downloads2y agoHugging Face09ov015 /VideoHallucer VideoHallucer Paper: https://huggingface.co/papers/2406.16338 This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/ov015/VideoHallucer.textquestion-answering1K<n<10K0 likes178 downloads5mo agoHugging Face10facebook /minimal_video_pairs Minimal Video Pairs A shortcut-aware benchmark for spatio-temporal and intuitive physics video understanding (VideoQA) using minimally different video pairs. Github For legal reasons, we are unable to upload the videos directly to Huggingface. However, we provide scripts in this repository for downloading the videos in our github repository. Our benchmark is built on top of videos source from 9 domains: Subset Data sources Human object interactions PerceptionTest… See the full description on the dataset page: https://huggingface.co/datasets/facebook/minimal_video_pairs.textquestion-answering10K<n<100K6 likes169 downloads1y agoHugging Face11video-reasoning /morse-500 MORSE-500 Benchmark 🔥 News May 15, 2025: We release MORSE-500, 500 programmatically generated videos across six reasoning categories: abstract, mathematical, physical, planning, spatial, and temporal, to stress-test multimodal reasoning. Frontier models including OpenAI o3 and Gemini 2.5 Pro score lower than… See the full description on the dataset page: https://huggingface.co/datasets/video-reasoning/morse-500.textvideo-classificationn<1K2 likes140 downloads1y agoHugging Face12DixinChen /VideoMind 🔍VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding Dataset Description VideoMind is a large-scale video-centric multimodal dataset that can be used to learn powerful and transferable text-video representations for video understanding tasks such as video question answering and video retrieval. The VideoMind dataset contains 105K(5K test for only) video samples, each of which is accompanied by audio, as well as systematic… See the full description on the dataset page: https://huggingface.co/datasets/DixinChen/VideoMind.textquestion-answering100K<n<1M1 likes135 downloads1y agoHugging Face13MCG-NJU /VideoChatOnline-IT Overview This dataset provides a comprehensive collection for Online Spatial-Temporal Understanding tasks, covering multiple domains including Dense Video Captioning, Video Grounding, Step Localization, Spatial-Temporal Action Localization, and Object Tracking. Data Formation Our pipeline begins with 96K high-quality samples curated from 5 tasks across 12 datasets. The conversion process enhances online spatiotemporal understanding through template transformation. We… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChatOnline-IT.textvisual-question-answering100K<n<1M5 likes131 downloads2y agoHugging Face14MongoDB /cooking-videos-with-captionsDataset of cooking videos obtained from pexels.com. Captions have been generated using AI. textquestion-answeringn<1K0 likes118 downloads9mo agoHugging Face15OpenGVLab /VideoChat2-ITgated Instruction Data Annotations A comprehensive dataset of 1.9M data annotations is available in JSON format. Due to the extensive size of the full data, we provide only JSON files here. For corresponding images and videos, please follow our instructions. Source data Image For image datasets, we utilized M3IT, filtering out lower-quality data by: Correcting typos: Most sentences with incorrect punctuation usage were rectified. Rephrasing incorrect… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VideoChat2-IT.textvisual-question-answering1M<n<10M52 likes112 downloads2y agoHugging Face16beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K3 likes99 downloads1mo agoHugging Face17ShareGPTVideo /test_raw_video_data ShareGPTVideo Raw Videos for Testing data All dataset and models can be found at ShareGPTVideo. Contents: In case of need, this contains raw videos corresponding to test frames in Test video frames textquestion-answering1K<n<10K2 likes69 downloads2y agoHugging Face18shuzhig /elv-halluc-videos ELV-Halluc — videos + annotations A self-contained mirror of the ELV-Halluc benchmark (CVPR 2026), bundling the raw .mp4 files together with the annotations so the benchmark can be run without sourcing videos separately. Paper: arXiv:2508.21496 Original annotations: HLSv/ELV-Halluc (no videos) Project page: https://elv-halluc.github.io/ This is an unofficial mirror. All credit for the benchmark goes to the original authors; please cite their paper (below) rather than this… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/elv-halluc-videos.tabularvideo-text-to-text1K<n<10K0 likes61 downloads1mo agoHugging Face19video-reasoning /morse-500-view MORSE-500 Benchmark The viewing version of MORSE-500 Benchmark, which allows you to view the video in the webpage directly. Dataset Structure test/: Contains all MP4 video files test/metadata.csv: Contains the dataset metadata, including video_path, query, ground_truth, question_text, and main_category textvideo-classificationn<1K3 likes55 downloads1y agoHugging Face20lmgame /VideoScienceBench VideoScienceBench A benchmark for evaluating video understanding and scientific reasoning in vision-language models. Each example pairs a textual description of an experiment (what is shown) with the correct scientific explanation (expected phenomenon). Dataset Summary Attribute Value Examples 160 Domains Physics, Chemistry Format JSONL (prompt + expected phenomenon + vid) Data Creation Pipeline Each researcher selects two or more scientific… See the full description on the dataset page: https://huggingface.co/datasets/lmgame/VideoScienceBench.textquestion-answeringn<1K3 likes52 downloads7mo agoHugging Face21uuookk /VideoChatOnline-IT Overview This dataset provides a comprehensive collection for Online Spatial-Temporal Understanding tasks, covering multiple domains including Dense Video Captioning, Video Grounding, Step Localization, Spatial-Temporal Action Localization, and Object Tracking. Data Formation Our pipeline begins with 96K high-quality samples curated from 5 tasks across 12 datasets. The conversion process enhances online spatiotemporal understanding through template transformation.… See the full description on the dataset page: https://huggingface.co/datasets/uuookk/VideoChatOnline-IT.textvisual-question-answering100K<n<1M0 likes51 downloads6mo agoHugging Face22VideoSimpleQA /VideoSimpleQA Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models 📖 Overview Video SimpleQA is the first comprehensive benchmark specifically designed for evaluating factual grounding capabilities in Large Video Language Models (LVLMs). Unlike existing video benchmarks that often involve subjective speculation or conflate factual grounding with reasoning skills, Video SimpleQA focuses exclusively on objective factuality evaluation through multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/VideoSimpleQA/VideoSimpleQA.textquestion-answering1K<n<10K2 likes42 downloads1y agoHugging Face23shuzhig /VideoHallucer VideoHallucer (mirror) A redistribution of the VideoHallucer benchmark, bundled together with its videos so the whole benchmark comes down in a single snapshot_download. This is not the official release. All credit goes to the original authors. Official code: https://github.com/patrick-tssn/VideoHallucer · Official data: https://huggingface.co/datasets/bigai-nlco/VideoHallucer VideoHallucer is the first comprehensive benchmark for hallucination detection in large… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/VideoHallucer.textquestion-answering1K<n<10K0 likes27 downloads1mo agoHugging Face24yifangsm /VideoHallucer VideoHallucer Paper: https://huggingface.co/papers/2406.16338 This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/yifangsm/VideoHallucer.textquestion-answering1K<n<10K0 likes24 downloads5mo agoHugging Face25snfacademy /snfa-youtube-videodaten SNFA YouTube-Videodaten Ein strukturierter Datensatz mit veröffentlichten YouTube-Videos der SNF Academy und zugehörigen Inhalten aus den Bereichen Fitness, Ernährung, Coaching, Mindset, Personal Training und Ausbildung. Datensatzübersicht 1'423 eindeutige Videos 1'423 eindeutige YouTube-Video-IDs 1'065 Videos mit Beschreibung 472'190 erfasste Views Veröffentlichungszeitraum: 6. September 2013 bis 10. Mai 2026 Datenprüfung: 17. Juli 2026 Sprache: überwiegend… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/snfa-youtube-videodaten.texttext-generation1K<n<10K0 likes20 downloads2mo agoHugging Face26aintropy-ai /legal-videos-rag Legal Videos RAG Benchmark Legal Videos is a benchmark for evaluating RAG pipelines on real-world legal videos pulled from two legal proceedings video datasets. LocalView, the largest known database of local government public meetings as they are captured and uploaded online covering more than 1000 hours of video. Seattle City meetings from the Council Data Project (CDP), is the Seattle city subset of the CDP data having meeting videos and multiple metadata covering about 1200… See the full description on the dataset page: https://huggingface.co/datasets/aintropy-ai/legal-videos-rag.textquestion-answering1K<n<10K0 likes18 downloads4mo agoHugging Face27ryohu053 /LoMo_Video_Benchmark LoMo Benchmark: Longer and More Video Benchmark Introduction LoMo benchmark is an fully automatic annotated video understanding benchmark with over 14,000+ videos. Videos duration are from 1 minutes to 4 hours. For every video, we provide 6 tasks. Usage If you want to use this benchmark, you should download the original video data. You can follow the instruction here to download. Attention: Due to copyright issues, we are unable to provide the original… See the full description on the dataset page: https://huggingface.co/datasets/ryohu053/LoMo_Video_Benchmark.textquestion-answering10K<n<100K1 likes16 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.