CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-VLM-Dataset-v2 Nemotron-VLM-Dataset v2 Versions Date Commit Changes 2025-11-05 head Fix nights_cot dataset. Fix/filter broken <think> entries. Update fintabnet instructions. Update indexes. 2025-10-28 214051e Initial Release Dataset Description Following up on Llama Nemotron VLM Dataset V1 with 3 million samples, we are releasing the Nemotron VLM Dataset V2 with almost three times as many high-quality samples. This time, our focus was on three… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-VLM-Dataset-v2.textvisual-question-answering1M<n<10M97 likes9.1k downloads9mo agoHugging Face02VLM2Vec /MSR-VTTClone from "friedrichor/MSR-VTT". MSRVTT contains 10K video clips and 200K captions. We adopt the standard 1K-A split protocol, which was introduced in JSFusion and has since become the de facto benchmark split in the Text-Video Retrieval field. Train: train_7k: 7,010 videos, 140,200 captions train_9k: 9,000 videos, 180,000 captions Test: test_1k: 1,000 videos, 1,000 captions 🌟 Citation @inproceedings{xu2016msrvtt, title={Msr-vtt: A large video description dataset… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MSR-VTT.tabulartext-to-video10K<n<100K4 likes7k downloads1y agoHugging Face03VLM2Vec /DiDeMoClone from friedrichor/DiDeMo. About DiDeMo contains 10K long-form videos from Flickr. For each video, ~4 short sentences are annotated in temporal order. We follow the existing works to concatenate those short sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark. We adopt the official split: Train: 8,395 videos, 8,395 captions (concatenate from 33,005 short captions) Val: 1,065 videos, 1,065 captions (concatenate from 4,290 short captions) (We don't have… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/DiDeMo.texttext-to-video1K<n<10K0 likes4.4k downloads1y agoHugging Face04VLM2Vec /MSVDClone from "friedrichor/MSVD". MSVD contains 1,970 videos, each of which is paired with ~40 captions. We adopt the official split: Train: 1,200 videos, 48,774 captions Val: 100 videos, 4,290 captions Test: 670 videos, 27,763 captions 🌟 Citation @inproceedings{chen2011collecting, title={Collecting highly parallel data for paraphrase evaluation}, author={Chen, David and Dolan, William B}, booktitle={Proceedings of the Annual Meeting of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MSVD.texttext-to-video1K<n<10K4 likes3.6k downloads1y agoHugging Face05nvidia /Llama-Nemotron-VLM-Dataset-v1 Llama-Nemotron-VLM-Dataset v1 Versions Date Commit Changes 2025-08-11 bdb3899 Initial release 2025-08-18 5abc7df Fixes bug (ocr_1 and ocr_3 images were swapped) 2025-08-19 ef85bef Update instructions for ocr_9 2025-08-25 4e46f2b Added example for Megatron Energon 2025-09-02 head Update license headers Quickstart If you want to dive in right away and load some samples using Megatron Energon, check out this section below. Data… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-VLM-Dataset-v1.textvisual-question-answering1M<n<10M168 likes3.2k downloads11mo agoHugging Face06initiacms /XLRS-Bench-lite_VLM 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite_VLM.textvisual-question-answering1K<n<10K0 likes2.5k downloads11mo agoHugging Face07VLM2Vec /MVBench MVBench Forked from https://huggingface.co/datasets/OpenGVLab/MVBench for reproducibility. Important Update [18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit ROSE Lab to access the data. We also provide a list of the 320 videos used in MVBench for your reference. We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MVBench.imagevisual-question-answering1K<n<10K0 likes1.5k downloads1y agoHugging Face08introvoyz041 /Nemotron-VLM-Dataset-v2 Nemotron-VLM-Dataset v2 Versions Date Commit Changes 2025-11-05 head Fix nights_cot dataset. Fix/filter broken <think> entries. Update fintabnet instructions. Update indexes. 2025-10-28 214051e Initial Release Dataset Description Following up on Llama Nemotron VLM Dataset V1 with 3 million samples, we are releasing the Nemotron VLM Dataset V2 with almost three times as many high-quality samples. This time, our focus was on three main… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/Nemotron-VLM-Dataset-v2.textvisual-question-answering1M<n<10M0 likes1k downloads9mo agoHugging Face09MicroAGI-Labs /vlm-info-loss-results VLM Grounding Evaluation Results Grounding evaluation results for vision-language models on robotics manipulation datasets. Part of the vlm-info-loss project studying how VLM connectors transform visual representations. Background Our embedding-level analysis shows VLM connectors perform a compress-then-expand transformation: they sharpen dominant-object representations while compressing secondary-object category identity. All tested models converge to ~83%… See the full description on the dataset page: https://huggingface.co/datasets/MicroAGI-Labs/vlm-info-loss-results.imageobject-detectionn<1K0 likes1k downloads5mo agoHugging Face10VLM2Vec /HMDB51text10K<n<100K0 likes744 downloads1y agoHugging Face11VLM2Vec /UCF101text10K<n<100K0 likes741 downloads1y agoHugging Face12VLM2Vec /Breakfasttextn<1K0 likes738 downloads1y agoHugging Face13VLM2Vec /SmthSmthV2text100K<n<1M0 likes714 downloads1y agoHugging Face14Journey9ni /VLM-3R-DATA VLM-3R Training Data Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split). Erratum (2026-07-13): corrected camera-position ground truth A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.text100K<n<1M17 likes489 downloads2mo agoHugging Face15arcadia-impact /reward-projection-goal-generalisation-vlmtabular1K<n<10K0 likes366 downloads2mo agoHugging Face16initiacms /OmniEarth-Bench_MCQ_VLMtext1K<n<10K0 likes290 downloads1y agoHugging Face17VLM-Reasoning /VCR-Bench VCR-Bench ( A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning) 🌐 Homepage | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub Dataset Details As shown in the figure below, current video benchmarks often lack comprehensive annotations of CoT steps, focusing only on the accuracy of final answers during model evaluation while neglecting the quality of the reasoning process. This evaluation approach makes it difficult to comprehensively evaluate model’s… See the full description on the dataset page: https://huggingface.co/datasets/VLM-Reasoning/VCR-Bench.image1K<n<10K6 likes237 downloads1y agoHugging Face18zGinger /XLRS-Bench-lite_VLM 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/zGinger/XLRS-Bench-lite_VLM.textvisual-question-answering1K<n<10K0 likes231 downloads9mo agoHugging Face19henry1477 /pcbslm-static-v2-unsloth-vlm PCBSLM static-v2 Unsloth VLM Portable multimodal Unsloth dataset for PCB layout/document-grounded training. The JSONL splits use Unsloth/Gemma-style chat messages: { "messages": [ {"role": "user", "content": [ {"type": "image", "image": "assets/raw_docs/.../images/page.png"}, {"type": "text", "text": "instruction..."} ]}, {"role": "assistant", "content": [ {"type": "text", "text": "{...json answer...}"} ]} ] } Files… See the full description on the dataset page: https://huggingface.co/datasets/henry1477/pcbslm-static-v2-unsloth-vlm.documentimage-text-to-text1K<n<10K0 likes193 downloads5mo agoHugging Face20Harisundar /PALL-VLM-data PALL-VLM-data — Dental Vision-Language Dataset The training dataset for Harisundar/PALL-VLM, a dental vision-language model. It contains 32,884 records over 52,461 images, formatted as image+text conversations for LLaVA-style instruction tuning. Curated by: Harisundar R Used by: Harisundar/PALL-VLM · PALL on GitHub Language: English Layout vlm_train/ ├── images/ # 52,461 dental images ├── train.jsonl # 29,667 records ├── val.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/PALL-VLM-data.imageimage-text-to-text10K<n<100K1 likes186 downloads3mo agoHugging Face21VLM-Forgetting /vlm-forgetting-datasetstext1M<n<10M0 likes153 downloads1y agoHugging Face22VLM4Geo /oda-recon-val-4Btabularn<1K0 likes127 downloads5mo agoHugging Face23zixuanlan /VLM_Train VLM Single-turn Deduplicated Collection v3 Locally curated version, dated September 6, 2026. This collection retains 529,816 question-answer pairs associated with 276,437 unique images. Images were deduplicated by the SHA-256 hash of their file bytes, without re-encoding. Samples containing different question-answer pairs for the same image were retained. Image files total approximately 75.672 GB; the combined logical size of images and annotations is approximately 76.257 GB.… See the full description on the dataset page: https://huggingface.co/datasets/zixuanlan/VLM_Train.textvisual-question-answering10K<n<100K0 likes117 downloads15d agoHugging Face24UCSC-VLAA /VLM-CapCurriculum-Perception-Data VLM-CapCurriculum-Perception (D_perc) Stage-1 visual perception data for the staged post-training recipe in "From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models" (ICML 2026). Each sample is a 4-way multiple-choice question over an image where the question can be answered from a fine-grained image caption but is missed by a strong VLM looking only at the image — by construction, these samples isolate perception failures from… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-Perception-Data.textvisual-question-answering1K<n<10K0 likes96 downloads4mo agoHugging Face25Nhanvi282 /openvivqa-formating-vlm OpenViVQA Formatting Dataset for VLM A Vietnamese multimodal instruction-format dataset for training Vision Language Models (VLMs) on Visual Question Answering (VQA) tasks. This dataset reformats OpenViVQA-style samples into conversational instruction-tuning format compatible with modern VLM training pipelines such as: Qwen2-VL LLaVA InternVL Phi-3 Vision Idefics SmolVLM Dataset Structure Each sample contains: image: input image conversations: multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/Nhanvi282/openvivqa-formating-vlm.imagevisual-question-answering10K<n<100K1 likes83 downloads5mo agoHugging Face26aitf-its-tim3-dfk /aitf-dfk3-vlm-dataset-jsonlimage10K<n<100K0 likes55 downloads3mo agoHugging Face27VLM-Reasoner /deepscalertext10K<n<100K3 likes51 downloads2y agoHugging Face28beezza /ogiri-bokete-unsloth-vlm Japanese Bokete Ogiri — Unsloth VLM format YANS-official/ogiri-bokete を、UnslothのVision SFTで扱える会話形式に変換した非公開用データセットです。 各JSONLレコードは「1画像 + 1回答」です。 { "messages": [ {"role": "user", "content": [ {"type": "image", "image": "images/124469.jpg"}, {"type": "text", "text": "この画像のお題に対して、面白い一言を1つ返してください。"} ]}, {"role": "assistant", "content": [ {"type": "text", "text": "..."} ]} ] } Files train.jsonl: 1,678 records / 630 prompts… See the full description on the dataset page: https://huggingface.co/datasets/beezza/ogiri-bokete-unsloth-vlm.imageimage-to-text1K<n<10K0 likes50 downloads2mo agoHugging Face29Wendy-Fly /TBStar-VLM-R2image100K<n<1M0 likes44 downloads1y agoHugging Face30maticmatusek /VLM_semantics_SLO_benchmark VLM Semantics SLO Benchmark VLM Semantics SLO is a Slovenian multimodal benchmark for studying cultural and semiotic reasoning in vision-language models. It goes beyond object recognition by asking models to interpret visual hierarchy, spatial relations, colour and mood, composition, cultural symbols, metaphor, denotation and connotation, intertextuality, communicative intent, and relevance to Slovenia. The released JSON contains 4,950 image-level records. Every record has ten… See the full description on the dataset page: https://huggingface.co/datasets/maticmatusek/VLM_semantics_SLO_benchmark.imagevisual-question-answering1K<n<10K0 likes42 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.