CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zai-org /Vision2Web Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification [🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results] Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle. This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.imagetext-generationn<1K23 likes1.5k downloads6mo agoHugging Face02twinkle-ai /tw-drug-labels-vision Dataset Card for tw-drug-labels-vision 💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。 Dataset Details Dataset Description 本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段: 下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。 頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。 OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.imageimage-to-text10K<n<100K4 likes455 downloads5mo agoHugging Face03gimmy256 /afri-aya-vision Afri-Aya Vision: Restructured for Multimodal & Adaption Fine-Tuning This dataset is a restructured, multimodal Vision-Language (VLM) adaptation of CohereLabsCommunity/afri-aya (Giving Sight to African LLMs). Why This Restructured Version? The original Afri-Aya dataset stores multiple question-and-answer pairs per image inside a nested list column (qa_pairs). Fine-tuning platforms (such as Adaption, Unsloth, LLaVA, and standard VLM training harnesses) require: 1… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/afri-aya-vision.imagevisual-question-answering1K<n<10K0 likes84 downloads3d agoHugging Face04renhehuang /formosa-vision-finegrained Formosa Vision Fine-grained (Expanded) Dataset Summary 此資料集以台灣在地文化與地景為核心,提供具細節的中文描述,並保留原始圖像。 擴充版本針對每張圖像生成更長、更密集的語義描述,以強化模型在細節理解上的表現。 Motivation 『資料合成』FLAIR 的核心在於訓練模型「聽得懂細節」。這意味著「長文本」越具體、包含越多方位詞 (左上角、紅色物體旁...),模型學到的局部特徵就越好。因為在此階段會透過大型多模態模型生成豐富且長的中文描述夠「碎唸」(包含大量方位、顏色、材質等細節)。相較於網路爬蟲數據,此資料庫具備高品質的本土文化實體 (Entity) 標註,是訓練台灣在地化 AI 的最佳基石。 Source Data 原始資料集:twinkle-ai/Formosa-Vision(Hugging Face Datasets) 擴充流程:以本地 VLM 產生更細緻的中文長描述 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/formosa-vision-finegrained.imageimage-to-text1K<n<10K0 likes43 downloads8mo agoHugging Face05DinoDS /dino-data-vision-tooling-preview Dino Data Vision Tooling Preview What This Dataset Is This dataset is a focused vision-tooling preview built from two Dino Data capability slices: image context understanding image tooling The goal is to train or inspect assistant behavior for image-related tasks where visual context, multimodal interpretation, or tool-aware image handling is relevant. Included Capability Slices Source lane Public task name What it teaches… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/dino-data-vision-tooling-preview.tabulartext-generationn<1K0 likes38 downloads5mo agoHugging Face06amalia-llm /MATH-Vision-PT MATH-Vision-PT European Portuguese (pt-PT) machine translation of MATH-Vision, a benchmark of competition-level mathematics problems presented in visual contexts. Translated from the original English test split using gemini-3.1-pro. Original Dataset: https://huggingface.co/datasets/MathLLMs/MathVision Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MATH-Vision-PT.imagequestion-answering1K<n<10K0 likes34 downloads3mo agoHugging Face07yahelr1 /beecare-vision-public-bilingual-train BeeCare Vision Public Bilingual Train Default train split for Unsloth vision fine-tuning. Each row has image, messages, and metadata. Use in Unsloth Studio as: yahelr1/beecare-vision-public-bilingual-train. imagevisual-question-answeringn<1K0 likes10 downloads4mo agoHugging Face08yahelr1 /beecare-vision-rich-qa-bilingual BeeCare Vision Rich QA Bilingual Unsloth-friendly rich image dataset. Default split is train. Columns include image, text, question, answer, condition_label, task, severity, and safety/provenance fields. In Unsloth, select this repo and map image to image, text to text if asked. imagevisual-question-answeringn<1K0 likes7 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.