CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01picbreeder-vlm /picbreeder-vlm-archive Picbreeder-VLM Archive Every image evolved by the swarm of vision-language-model "breeders" in In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models (GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the lineage graphs, and the analysis artifacts behind the paper and blog. The original Picbreeder (Secretan et al., 2008) let crowds of humans collaboratively evolve images from CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.imageimage-to-text100K<n<1M14 likes66k downloads3mo agoHugging Face02aialliance /GEOBench-VLM GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks Summary While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they fall short in addressing the unique demands of geospatial applications. Generic VLM benchmarks are not designed to handle the complexities of geospatial data, which is critical for applications such as environmental monitoring, urban planning, and disaster management. Some of the unique… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/GEOBench-VLM.image1K<n<10K17 likes4k downloads1y agoHugging Face03VLM2Vec /VATEXClone from lmms-lab/VATEX. text1K<n<10K1 likes3.6k downloads1y agoHugging Face04lingamvamshikrishnareddy /ramanv-image-vlm-instructiontext100K<n<1M0 likes3.6k downloads24d agoHugging Face05JosselinSom /Latex-VLMimagequestion-answering1K<n<10K9 likes3.3k downloads3y agoHugging Face06VLM2Vec /MMLongBench-docimage10K<n<100K0 likes2.6k downloads1y agoHugging Face07Vi-VLM /Vista Dataset Card for "Vista" "700.000 Vietnamese vision-language samples open-source dataset" Dataset Overview This dataset contains over 700,000 Vietnamese vision-language samples, created by Gemini Pro. We employed several prompt engineering techniques: few-shot learning, caption-based prompting and image-based prompting. For the COCO dataset, we generated data using Llava-style prompts For the ShareGPT4V dataset, we used translation prompts. Caption-based prompting:… See the full description on the dataset page: https://huggingface.co/datasets/Vi-VLM/Vista.imagevisual-question-answering100K<n<1M45 likes2.4k downloads2y agoHugging Face08VLM2Vec /NExTQAtabular10K<n<100K1 likes2.3k downloads2y agoHugging Face09VLM2Vec /Video-MMEtext1K<n<10K0 likes2.2k downloads2y agoHugging Face10VLM2Vec /EgoSchematext10K<n<100K0 likes2.2k downloads2y agoHugging Face11VLM2Vec /ViDoSeekimage10K<n<100K1 likes2.1k downloads1y agoHugging Face12qgallouedec /test-grpo-vlm-log-completions TRL Completion logs This dataset contains the completions generated during training using trl. The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument). Each file contains the following columns: step: the step of training prompt: the prompt used to generate the completion completion: the completion generated by the model <reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/test-grpo-vlm-log-completions.tabularn<1K0 likes2k downloads6mo agoHugging Face13VLM2Vec /MMLongBench-page-fixedimage1K<n<10K0 likes1.9k downloads11mo agoHugging Face14VLM2Vec /ViDoSeek-page-fixedimage1K<n<10K0 likes1.9k downloads11mo agoHugging Face15XAI /vlmsareblindArXiv - Website imagequestion-answering1K<n<10K28 likes1.8k downloads2y agoHugging Face16Flame-Code-VLM /Flame-Waterfall-React Flame-Waterfall-React: A Structured Data Synthesis Dataset for Multimodal React Code Generation Flame-Waterfall-React is a dataset synthesized using the Waterfall-Model-Based Synthesis method, Advancing Vision-Language Models in Front-End Development via Data Synthesis. This dataset is designed to train vision-language models (VLMs) for React code generation from UI design mockups and specifications. The Waterfall synthesis approach mimics real-world software development by… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Waterfall-React.textimage-to-text100K<n<1M2 likes1.8k downloads1y agoHugging Face17ko-vlm /KoLLaVA-v1.5-Instruct-581k KoLLaVA-v1.5-Instruct-581k 한국어 Vision-Language 모델을 위한 instruction tuning 데이터셋입니다. 데이터셋 정보 총 샘플 수: 435,093개 형식: ChatML 형식 (role: user/assistant, content: 텍스트) 이미지: COCO + GQA + Visual Genome 데이터셋 언어: 한국어 포함된 데이터셋 COCO 데이터: 362,953개 샘플 MS COCO 2017 이미지 기반 한국어 대화 데이터 GQA 데이터: 72,140개 샘플 GQA (Visual Question Answering) 이미지 기반 한국어 대화 데이터 Visual Genome 데이터: 포함 Visual Genome 이미지 기반 한국어 대화 데이터 제외된 데이터셋 EKVQA 데이터: AI Hub 라이선스로 인해 공개 불가… See the full description on the dataset page: https://huggingface.co/datasets/ko-vlm/KoLLaVA-v1.5-Instruct-581k.image100K<n<1M0 likes1.7k downloads1y agoHugging Face18JoyboyBrian /nano-omni-vlmimage1M<n<10M0 likes1.7k downloads2y agoHugging Face19KORMo-VL /Nemotron-VLM-Dataset-v2from nvidia/Nemotron-VLM-Dataset-v2 samples are: visual7w_telling_cot: 435299 plotqa_cot: 295354 wiki_ko: 200000 wiki_en: 200000 mulberry_cot_1: 189378 mulberry_cot_2: 102279 sparsetables: 100000 mantis_instruct_cot: 67714 llava_cot_100k: 63019 visual_web_instruct_cot: 47800 chartqa_cot: 45710 docvqa_cot: 36333 tabmwp_cot: 20305 infographicsvqa_cot: 19548 hiertext: 514 image1M<n<10M0 likes1.6k downloads7mo agoHugging Face20anvo25 /vlms-are-biased Vision Language Models are Biased by An Vo1*, Khai-Nguyen Nguyen2*, Mohammad Reza Taesiri3, Vy Tuong Dang1, Anh Totti Nguyen4†, Daeyoung Kim1† *Equal contribution    †Equal advising 1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vlms-are-biased.imagevisual-question-answering10K<n<100K29 likes1.2k downloads10mo agoHugging Face21shijiezhou /VLM4D VLM4D VLM4D is a benchmark for evaluating the spatiotemporal reasoning capabilities of Vision Language Models (VLMs). It contains real and synthetic videos paired with multiple-choice questions that require models to reason about translation, rotation, perspective, motion continuity, counting, and false-positive events. The dataset was introduced in VLM4D: Towards Spatiotemporal Awareness in Vision Language Models, accepted to ICCV 2025. Project page: https://vlm4d.github.io/… See the full description on the dataset page: https://huggingface.co/datasets/shijiezhou/VLM4D.textvideo-text-to-textn<1K4 likes1.1k downloads3mo agoHugging Face22VLM2Vec /MomentSeekertext1K<n<10K0 likes1k downloads1y agoHugging Face23Koa-Chang /TissueMNIST-224-full-gpt5nano-with-vlm-features TissueMNIST 224 Full Train Val with GPT-5-nano VLM Features The full TissueMNIST train and validation splits with categorical morphology features generated by GPT-5-nano. Test is included as the full TissueMNIST passthrough split with null vlm_model_name and placeholder vlm_feature values for schema consistency. This dataset is derived from the official MedMNIST TissueMNIST 224px data. The VLM feature labels are categorical privileged-information annotations for CS231N VLM-LUPI… See the full description on the dataset page: https://huggingface.co/datasets/Koa-Chang/TissueMNIST-224-full-gpt5nano-with-vlm-features.text100K<n<1M0 likes995 downloads4mo agoHugging Face24dgorbatov /vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10 Trajectory Ranking Dataset This dataset contains trajectory ranking results for autonomous navigation scenarios. Dataset Statistics Total examples: 39558 Chunks processed: 40 Upload date: 2025-09-13T00:44:30.335177 Features Image data with terrain analysis Trajectory rankings and reasoning Quality and diversity analysis Terrain and trajectory descriptions imageimage-classification10K<n<100K0 likes745 downloads1y agoHugging Face25VLM2Vec /Kinetics-700 Damaged Videos List The following rows are removed due to damaged video files: NNazT7dDWxA_000130_000140 5d9mIpws4cg_000130_000140 SYTMgaqGhfg_000010_000020 ixQrfusr6k8_000001_000011 BSN_nDiTwBo_000004_000014 y7cYaYX4gdw_000047_000057 A-FCzUzEd4U_000000_000010 _dbw-EJqoMY_001023_001033 zLD_q2djrYs_000030_000040 FAqHwAPZfeE_000018_000028 textimage-classification1K<n<10K1 likes610 downloads1y agoHugging Face26prism-vlm /gemini_public_mmr1 PRISM Public SFT Data Overview PRISM Public SFT Data is the public supervised fine-tuning data collection used in the PRISM project.PRISM studies the distributional drift problem in the standard SFT → RLVR post-training pipeline for large multimodal models. Before the distribution alignment and RLVR stages, we first use large-scale public multimodal demonstrations to obtain a broad SFT initialization. This dataset serves as the public SFT data source for the… See the full description on the dataset page: https://huggingface.co/datasets/prism-vlm/gemini_public_mmr1.textimage-to-text1M<n<10M2 likes609 downloads5mo agoHugging Face27Jiwon-Kang /Llama-Nemotron-VLM-Dataset-v1-OCR4image100K<n<1M1 likes579 downloads8mo agoHugging Face28VLM2Vec /QVHighlighttext1K<n<10K0 likes572 downloads1y agoHugging Face29chuuhtetnaing /myanmar-ocr-dataset-for-vlm Myanmar OCR Dataset A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts. Subsets Subset Description Details single_font Rendered with Pyidaungsu font only 437 books multi_font Rendered with 76 Myanmar fonts 3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.image100K<n<1M2 likes566 downloads5mo agoHugging Face30alay2shah /example-vlm-sft-dataset Sample VLM Reference Dataset Dataset Description This is a reference dataset demonstrating the proper VLM SFT (Supervised Fine-Tuning) format for vision-language model training. It contains 10 minimal conversational examples that show the exact structure and formatting required for VLM training pipelines. ⚠️ Important: This dataset is for format reference only - not intended for actual model training. Use this as a template to understand the required data structure for… See the full description on the dataset page: https://huggingface.co/datasets/alay2shah/example-vlm-sft-dataset.textvisual-question-answeringn<1K0 likes466 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.