CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-visionx /Cambrian-Alignment Cambrian-Alignment Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V. Getting Started with Cambrian Alignment Data Before you start, ensure you have sufficient storage space to download and process the data. Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.imagevisual-question-answering100K<n<1M38 likes6.4k downloads2y agoHugging Face02twinkle-ai /tw-drug-labels-vision Dataset Card for tw-drug-labels-vision 💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。 Dataset Details Dataset Description 本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段: 下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。 頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。 OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.imageimage-to-text10K<n<100K4 likes455 downloads5mo agoHugging Face03translorentz /vision-token-compression-bench OPTIC-Bench Optical Text In-Context Benchmark: how reliably do LLMs consume text delivered as rendered images versus plain text tokens? In summary, the evaluation reported here finds that optical text compression is effective only within a narrow and specific envelope. Delivering content as rendered images genuinely reduces input tokens, by thirteen to fifty-four per cent depending on the model and the language, but only when the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.imagevisual-question-answering1K<n<10K0 likes170 downloads2mo agoHugging Face04twinkle-ai /Formosa-Vision Dataset Card for Formosa-Vision Formosa Vision 是一份以台灣在地文化為核心的開源視覺語言資料集,從國家文化記憶庫 2.0中精選兩千餘張資料,文字描述採用 OGDL 1.0 授權、及圖片為 CC By SA(及更開放的授權條款)授權的影像,內容涵蓋景點、建築、生活場景與歷史脈絡。資料集以模型生成與人工審核並行的方式建立,透過視覺語言模型產生影像對話,再由參與者逐一檢查與修訂,確保描述的正確性、文化脈絡的一致性與語句的自然性。專案由 Twinkle AI 社群發起,結合社群協作與開放文化精神,期待成為訓練繁體中文視覺語言模型的重要基礎,幫助研究者與開發者打造能真正理解台灣文化細節的 VLM 模型。 Dataset Details Dataset Description Formosa Vision(又稱 台灣視覺資料集)是一個以台灣在地視覺文化為核心、集結社群力量共創的開源資料集。這項專案源自近年視覺語言模型(Vision Language Model… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/Formosa-Vision.imagequestion-answering1K<n<10K13 likes123 downloads10mo agoHugging Face05sujet-ai /Sujet-Finance-QA-Vision-100k Dataset Description 📊🔍 The Sujet-Finance-QA-Vision-100k is a comprehensive dataset containing over 100,000 question-answer pairs derived from more than 9,800 financial document images. This dataset is designed to support research and development in the field of financial document analysis and visual question answering. Key Features: 🖼️ 9,801 unique financial document images ❓ 107,050 question-answer pairs 🇬🇧 English language 📄 Diverse financial document types… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-QA-Vision-100k.imagequestion-answering1K<n<10K39 likes100 downloads2y agoHugging Face06HAERAE-HUB /HAERAE-VISION HAERAE-VISION A Korean visual QA benchmark featuring real-world, under-specified questions. Dataset Description This dataset includes two question types: original: Under-specified, authentic user queries explicit: Clarified queries with full context Both share the same images and reference answers, allowing controlled evaluation of query under-specification. Evaluation Code See our GitHub repository for evaluation scripts. Citation… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAERAE-VISION.imagevisual-question-answeringn<1K17 likes62 downloads2mo agoHugging Face07LoadingBFX /GeoQA-train-Vision-R1-cot-rewrite Dataset Card for GeoQA-train-Vision-R1-cot-rewrite This dataset provides a rewritten version of the CoT (Chain-of-Thought) annotations for the GeoQA subset of the Vision-R1-cold dataset. It is designed to support efficient and structured multimodal reasoning with large language models. Dataset Details Dataset Description The original Vision-R1 dataset, introduced in the paper Vision-R1: Reflective Multimodal Reasoning with Aha Moments, features detailed and… See the full description on the dataset page: https://huggingface.co/datasets/LoadingBFX/GeoQA-train-Vision-R1-cot-rewrite.imagequestion-answering1K<n<10K0 likes54 downloads1y agoHugging Face08DinoDS /dino-data-vision-tooling-preview Dino Data Vision Tooling Preview What This Dataset Is This dataset is a focused vision-tooling preview built from two Dino Data capability slices: image context understanding image tooling The goal is to train or inspect assistant behavior for image-related tasks where visual context, multimodal interpretation, or tool-aware image handling is relevant. Included Capability Slices Source lane Public task name What it teaches… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/dino-data-vision-tooling-preview.tabulartext-generationn<1K0 likes38 downloads5mo agoHugging Face09amalia-llm /MATH-Vision-PT MATH-Vision-PT European Portuguese (pt-PT) machine translation of MATH-Vision, a benchmark of competition-level mathematics problems presented in visual contexts. Translated from the original English test split using gemini-3.1-pro. Original Dataset: https://huggingface.co/datasets/MathLLMs/MathVision Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MATH-Vision-PT.imagequestion-answering1K<n<10K0 likes34 downloads3mo agoHugging Face10LoadingBFX /GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite Dataset Card for GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite This dataset provides a rewritten version of the CoT (Chain-of-Thought) annotations for the GeoQA-PLUS subset of the Vision-R1-cold dataset. It is designed to support efficient and structured multimodal reasoning with large language models. Dataset Details Dataset Description The original Vision-R1 dataset, introduced in the paper Vision-R1: Reflective Multimodal Reasoning with Aha Moments, features… See the full description on the dataset page: https://huggingface.co/datasets/LoadingBFX/GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite.imagequestion-answering1K<n<10K0 likes32 downloads1y agoHugging Face111-800-SHARED-TASKS /Sujet-Vision-QA Dataset Description 📊🔍 The Sujet-Finance-QA-Vision-100k is a comprehensive dataset containing over 100,000 question-answer pairs derived from more than 9,800 financial document images. This dataset is designed to support research and development in the field of financial document analysis and visual question answering. Key Features: 🖼️ 9,801 unique financial document images ❓ 107,050 question-answer pairs 🇬🇧 English language 📄 Diverse financial document types… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/Sujet-Vision-QA.imagequestion-answering1K<n<10K0 likes31 downloads2y agoHugging Face12artificialguybr /brazilian-math-physics-qa-vision Brazilian Math & Physics QA — Image Dependent English | Português do Brasil English Summary Brazilian Portuguese educational question-answer pairs whose problem statement or solution depends on one or more images. Examples: 3,808 Referenced image URLs: 5,094 unique Language: Brazilian Portuguese (pt-BR) Schema {"id":"vqa_...","subject":"matematica","category":"geometria","title":"...","messages":[{"role":"user","content":"...… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa-vision.textvisual-question-answering1K<n<10K0 likes20 downloads1mo agoHugging Face13recursal /OKReddit-Visionary Dataset Summary OKReddit Visionary is a collection of 50 GiB (~74K pairs) of image Question & Answers. This dataset has been prepared for research or archival purposes. Curated by: KaraKaraWitch Funded by: Recursal.ai Shared by: KaraKaraWitch Special Thanks: harrison (Suggestion) Language(s) (NLP): Mainly English. License: Refer to Licensing Information for data license. Dataset Sources Source Data: Academic Torrents by (stuck_in_the_matrix, Watchful1… See the full description on the dataset page: https://huggingface.co/datasets/recursal/OKReddit-Visionary.imagequestion-answering100K<n<1M1 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.