CoolFace
20 results

vision-ai

lmarena-ai /vision-arena-bench-v0.1 VisionArena-Bench: An automatic eval pipeline to estimate model preference rankings An automatic benchmark of 500 diverse user prompts that can be used to cheaply approximate Chatbot Arena model rankings via automatic benchmarking with VLM as a judge. Dataset Sources Repository: https://github.com/lm-sys/FastChat Paper: https://arxiv.org/abs/2412.08687 Automatic Evaluation Code: Coming Soon! Dataset Structure question_id: The unique hash representing the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/vision-arena-bench-v0.1.imagevisual-question-answeringn<1K4 likes1.3k downloads2y agoHugging Facetwinkle-ai /tw-drug-labels-vision Dataset Card for tw-drug-labels-vision 💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。 Dataset Details Dataset Description 本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段: 下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。 頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。 OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.imageimage-to-text10K<n<100K4 likes453 downloads5mo agoHugging Facemultimodal-vision-ai /mvai-doctag-r3-v1 MVAI DocTag R3 v1 Basic Information Field Value Dataset ID multimodal-vision-ai/mvai-doctag-r3-v1 Version v1 Owner Dizzar, Hohai University / multimodal-vision-ai Dataset type Image + DocTags text annotations Intended use OCR and document layout research, especially image-to-DocTags SFT / GRPO This dataset is the formal closure version of the existing HohaiR3 DocTags data asset. It contains the HohaiR3 gold, human-labeled, HTML-rendered… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-vision-ai/mvai-doctag-r3-v1.image-to-text100K<n<1M0 likes403 downloads28d agoHugging FaceDavidsv /airline-vision-dataset Airline Industry VQA Dataset ⚠️ Note: This dataset currently contains only the text data (questions). The images are being processed and will be added in a future update. This dataset contains a comprehensive collection of visual question-answering (VQA) pairs generated from official documentation of 18 major airline companies. About the Creator I'm David Soeiro-Vuong, an engineering student specializing in Computer Science, Big Data, and AI, currently working as an… See the full description on the dataset page: https://huggingface.co/datasets/Davidsv/airline-vision-dataset.image10K<n<100K0 likes350 downloads1y agoHugging FaceAIGym /harmony-visionimage100K<n<1M2 likes306 downloads1y agoHugging Facebzcasper /ai-tool-pool-jewelry-vision AI Tool Pool Jewelry Vision Dataset Dataset Description This dataset contains 5,130 jewelry images organized into 5 categories for computer vision tasks. The dataset was originally created and hosted on Roboflow Universe. Categories Bracelet: Bracelet jewelry images Earrings: Earring jewelry images Necklace: Necklace jewelry images Pendant: Pendant jewelry images Ring: Ring jewelry images Dataset Structure AI-Tool-Pool-Jewelry-Vision/ ├── train/… See the full description on the dataset page: https://huggingface.co/datasets/bzcasper/ai-tool-pool-jewelry-vision.imageimage-classification1K<n<10K1 likes193 downloads1y agoHugging Face