CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01twinkle-ai /tw-drug-labels-vision Dataset Card for tw-drug-labels-vision 💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。 Dataset Details Dataset Description 本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段: 下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。 頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。 OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.imageimage-to-text10K<n<100K4 likes430 downloads5mo agoHugging Face02twinkle-ai /Formosa-Vision Dataset Card for Formosa-Vision Formosa Vision 是一份以台灣在地文化為核心的開源視覺語言資料集,從國家文化記憶庫 2.0中精選兩千餘張資料,文字描述採用 OGDL 1.0 授權、及圖片為 CC By SA(及更開放的授權條款)授權的影像,內容涵蓋景點、建築、生活場景與歷史脈絡。資料集以模型生成與人工審核並行的方式建立,透過視覺語言模型產生影像對話,再由參與者逐一檢查與修訂,確保描述的正確性、文化脈絡的一致性與語句的自然性。專案由 Twinkle AI 社群發起,結合社群協作與開放文化精神,期待成為訓練繁體中文視覺語言模型的重要基礎,幫助研究者與開發者打造能真正理解台灣文化細節的 VLM 模型。 Dataset Details Dataset Description Formosa Vision(又稱 台灣視覺資料集)是一個以台灣在地視覺文化為核心、集結社群力量共創的開源資料集。這項專案源自近年視覺語言模型(Vision Language Model… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/Formosa-Vision.imagequestion-answering1K<n<10K13 likes123 downloads10mo agoHugging Face03twinkle-ai /finevision-zhtw Dataset Card for finevisions-zhtw Brief Summary (Preliminary) This section describes the initial design intent of the project. It serves as a first-edition summary and may be revised or expanded as the dataset scope, governance, and contribution process become more clearly defined. 如何幫忙本專案 本專案正在進行中,歡迎任何夥伴加入幫忙,一起讓繁體中文的預訓練語料更多元及完善。 question-answering1M<n<10M2 likes13 downloads6mo agoHugging Face04twinkle-ai /F1-identitygated Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/F1-identity.textquestion-answering1K<n<10K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.