CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-visionx /Cambrian-10M Cambrian-10M Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-10M is a comprehensive dataset designed for instruction tuning, particularly in multimodal settings involving visual interaction data. The dataset is crafted to address the scarcity of high-quality multimodal instruction-tuning data and to maintain the language abilities of multimodal large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-10M.visual-question-answering1M<n<10M131 likes17k downloads2y agoHugging Face02nyu-visionx /Cambrian-Alignment Cambrian-Alignment Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V. Getting Started with Cambrian Alignment Data Before you start, ensure you have sufficient storage space to download and process the data. Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.imagevisual-question-answering100K<n<1M38 likes6.5k downloads2y agoHugging Face03twinkle-ai /tw-drug-labels-vision Dataset Card for tw-drug-labels-vision 💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。 Dataset Details Dataset Description 本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段: 下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。 頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。 OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.imageimage-to-text10K<n<100K4 likes454 downloads5mo agoHugging Face04visionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes209 downloads9mo agoHugging Face05translorentz /vision-token-compression-bench OPTIC-Bench Optical Text In-Context Benchmark: how reliably do LLMs consume text delivered as rendered images versus plain text tokens? In summary, the evaluation reported here finds that optical text compression is effective only within a narrow and specific envelope. Delivering content as rendered images genuinely reduces input tokens, by thirteen to fifty-four per cent depending on the model and the language, but only when the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.imagevisual-question-answering1K<n<10K0 likes170 downloads2mo agoHugging Face06twinkle-ai /Formosa-Vision Dataset Card for Formosa-Vision Formosa Vision 是一份以台灣在地文化為核心的開源視覺語言資料集,從國家文化記憶庫 2.0中精選兩千餘張資料,文字描述採用 OGDL 1.0 授權、及圖片為 CC By SA(及更開放的授權條款)授權的影像,內容涵蓋景點、建築、生活場景與歷史脈絡。資料集以模型生成與人工審核並行的方式建立,透過視覺語言模型產生影像對話,再由參與者逐一檢查與修訂,確保描述的正確性、文化脈絡的一致性與語句的自然性。專案由 Twinkle AI 社群發起,結合社群協作與開放文化精神,期待成為訓練繁體中文視覺語言模型的重要基礎,幫助研究者與開發者打造能真正理解台灣文化細節的 VLM 模型。 Dataset Details Dataset Description Formosa Vision(又稱 台灣視覺資料集)是一個以台灣在地視覺文化為核心、集結社群力量共創的開源資料集。這項專案源自近年視覺語言模型(Vision Language Model… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/Formosa-Vision.imagequestion-answering1K<n<10K13 likes122 downloads10mo agoHugging Face07beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K3 likes107 downloads1mo agoHugging Face08sujet-ai /Sujet-Finance-QA-Vision-100k Dataset Description 📊🔍 The Sujet-Finance-QA-Vision-100k is a comprehensive dataset containing over 100,000 question-answer pairs derived from more than 9,800 financial document images. This dataset is designed to support research and development in the field of financial document analysis and visual question answering. Key Features: 🖼️ 9,801 unique financial document images ❓ 107,050 question-answer pairs 🇬🇧 English language 📄 Diverse financial document types… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-QA-Vision-100k.imagequestion-answering1K<n<10K39 likes93 downloads2y agoHugging Face09Haonian /unified-math-vision-dataset Unified Math Vision Dataset Bundle Generated at: 2025-09-19 17:12:44 This is a unified dataset bundle containing multiple math and vision reasoning datasets. Dataset Statistics Total samples: 15858 mathvision: 3344 samples wemath: 500 samples mmmu: 415 samples mathvista: 6141 samples logicvista: 448 samples dynamath: 5010 samples Contents manifest.jsonl: Complete dataset in JSONL format (1 JSON per line) manifest.csv: Summary in CSV format images/: Directory… See the full description on the dataset page: https://huggingface.co/datasets/Haonian/unified-math-vision-dataset.textquestion-answering10K<n<100K0 likes62 downloads1y agoHugging Face10HAERAE-HUB /HAERAE-VISION HAERAE-VISION A Korean visual QA benchmark featuring real-world, under-specified questions. Dataset Description This dataset includes two question types: original: Under-specified, authentic user queries explicit: Clarified queries with full context Both share the same images and reference answers, allowing controlled evaluation of query under-specification. Evaluation Code See our GitHub repository for evaluation scripts. Citation… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAERAE-VISION.imagevisual-question-answeringn<1K17 likes61 downloads2mo agoHugging Face11LoadingBFX /GeoQA-train-Vision-R1-cot-rewrite Dataset Card for GeoQA-train-Vision-R1-cot-rewrite This dataset provides a rewritten version of the CoT (Chain-of-Thought) annotations for the GeoQA subset of the Vision-R1-cold dataset. It is designed to support efficient and structured multimodal reasoning with large language models. Dataset Details Dataset Description The original Vision-R1 dataset, introduced in the paper Vision-R1: Reflective Multimodal Reasoning with Aha Moments, features detailed and… See the full description on the dataset page: https://huggingface.co/datasets/LoadingBFX/GeoQA-train-Vision-R1-cot-rewrite.imagequestion-answering1K<n<10K0 likes53 downloads1y agoHugging Face12DinoDS /dino-data-vision-tooling-preview Dino Data Vision Tooling Preview What This Dataset Is This dataset is a focused vision-tooling preview built from two Dino Data capability slices: image context understanding image tooling The goal is to train or inspect assistant behavior for image-related tasks where visual context, multimodal interpretation, or tool-aware image handling is relevant. Included Capability Slices Source lane Public task name What it teaches… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/dino-data-vision-tooling-preview.tabulartext-generationn<1K0 likes37 downloads5mo agoHugging Face13amalia-llm /MATH-Vision-PT MATH-Vision-PT European Portuguese (pt-PT) machine translation of MATH-Vision, a benchmark of competition-level mathematics problems presented in visual contexts. Translated from the original English test split using gemini-3.1-pro. Original Dataset: https://huggingface.co/datasets/MathLLMs/MathVision Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MATH-Vision-PT.imagequestion-answering1K<n<10K0 likes35 downloads3mo agoHugging Face14LoadingBFX /GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite Dataset Card for GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite This dataset provides a rewritten version of the CoT (Chain-of-Thought) annotations for the GeoQA-PLUS subset of the Vision-R1-cold dataset. It is designed to support efficient and structured multimodal reasoning with large language models. Dataset Details Dataset Description The original Vision-R1 dataset, introduced in the paper Vision-R1: Reflective Multimodal Reasoning with Aha Moments, features… See the full description on the dataset page: https://huggingface.co/datasets/LoadingBFX/GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite.imagequestion-answering1K<n<10K0 likes34 downloads1y agoHugging Face151-800-SHARED-TASKS /Sujet-Vision-QA Dataset Description 📊🔍 The Sujet-Finance-QA-Vision-100k is a comprehensive dataset containing over 100,000 question-answer pairs derived from more than 9,800 financial document images. This dataset is designed to support research and development in the field of financial document analysis and visual question answering. Key Features: 🖼️ 9,801 unique financial document images ❓ 107,050 question-answer pairs 🇬🇧 English language 📄 Diverse financial document types… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/Sujet-Vision-QA.imagequestion-answering1K<n<10K0 likes31 downloads2y agoHugging Face16HPAI-BSC /CareQA-Vision CareQA-Vision Dataset Dataset Summary CareQA-Vision is a vision-based healthcare QA dataset derived from the Spanish Specialized Healthcare Training (FSE) exams. All questions are curated by medical experts and cover the specialties of nursing and medicine. This dataset extends the original CareQA dataset by including image-based questions from exams conducted between… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/CareQA-Vision.question-answeringn<1K2 likes27 downloads3mo agoHugging Face17Gene829 /gene-computer-vision-instruct computer-vision-instruct v4 Gate-passed instruction data for computer-vision — published when 50 fresh examples cleared the quality bar Kind: synthetic Domain: computer-vision Records: 177 Created: 2026-06-25T13:44:26+00:00 SHA-256: 60c3ff90c0668c54e691ed4f1ff2840584283f475e26358fa294a97882fb7767 Pipeline: v2.0.0 Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7} Generated by: Qwen3-4B-Instruct-2507-Q4_K_M.gguf (backend:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-computer-vision-instruct.text-generationn<1K0 likes25 downloads3mo agoHugging Face18initiacms /Text-Before-VisionGitHub: https://github.com/MiliLab/Text-Before-Vision textquestion-answering100K<n<1M2 likes23 downloads7mo agoHugging Face19artificialguybr /brazilian-math-physics-qa-vision Brazilian Math & Physics QA — Image Dependent English | Português do Brasil English Summary Brazilian Portuguese educational question-answer pairs whose problem statement or solution depends on one or more images. Examples: 3,808 Referenced image URLs: 5,094 unique Language: Brazilian Portuguese (pt-BR) Schema {"id":"vqa_...","subject":"matematica","category":"geometria","title":"...","messages":[{"role":"user","content":"...… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa-vision.textvisual-question-answering1K<n<10K0 likes22 downloads1mo agoHugging Face20lxm123123 /Text-Before-VisionGitHub: https://github.com/MiliLab/Text-Before-Vision textquestion-answering100K<n<1M0 likes14 downloads7mo agoHugging Face21Sivanuja /Legal_vision_finetuning_data Sri Lankan Property Law Fine-Tuning Dataset Dataset Summary This dataset is a domain-specific legal instruction-tuning dataset designed for fine-tuning large language models for Sri Lankan property law reasoning and legal assistance. It focuses on core areas of Sri Lankan property law, including: Property transfer and conveyancing Title registration (Bim Saviya) Prescription and adverse possession Partition of co-owned property Mortgage and securities Lease and tenancy… See the full description on the dataset page: https://huggingface.co/datasets/Sivanuja/Legal_vision_finetuning_data.texttext-generation1K<n<10K0 likes13 downloads7mo agoHugging Face22recursal /OKReddit-Visionary Dataset Summary OKReddit Visionary is a collection of 50 GiB (~74K pairs) of image Question & Answers. This dataset has been prepared for research or archival purposes. Curated by: KaraKaraWitch Funded by: Recursal.ai Shared by: KaraKaraWitch Special Thanks: harrison (Suggestion) Language(s) (NLP): Mainly English. License: Refer to Licensing Information for data license. Dataset Sources Source Data: Academic Torrents by (stuck_in_the_matrix, Watchful1… See the full description on the dataset page: https://huggingface.co/datasets/recursal/OKReddit-Visionary.imagequestion-answering100K<n<1M1 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.