vision-ai
vision-arena-bench-v0.1
VisionArena-Bench: An automatic eval pipeline to estimate model preference rankings
An automatic benchmark of 500 diverse user prompts that can be used to cheaply approximate Chatbot Arena model rankings via automatic benchmarking with VLM as a judge.
Dataset Sources
Repository: https://github.com/lm-sys/FastChat
Paper: https://arxiv.org/abs/2412.08687
Automatic Evaluation Code: Coming Soon!
Dataset Structure
question_id: The unique hash representing the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/vision-arena-bench-v0.1.tw-drug-labels-vision
Dataset Card for tw-drug-labels-vision
💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。
Dataset Details
Dataset Description
本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段:
下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。
頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。
OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.mvai-doctag-r3-v1
MVAI DocTag R3 v1
Basic Information
Field
Value
Dataset ID
multimodal-vision-ai/mvai-doctag-r3-v1
Version
v1
Owner
Dizzar, Hohai University / multimodal-vision-ai
Dataset type
Image + DocTags text annotations
Intended use
OCR and document layout research, especially image-to-DocTags SFT / GRPO
This dataset is the formal closure version of the existing HohaiR3 DocTags data
asset. It contains the HohaiR3 gold, human-labeled, HTML-rendered… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-vision-ai/mvai-doctag-r3-v1.airline-vision-dataset
Airline Industry VQA Dataset
⚠️ Note: This dataset currently contains only the text data (questions). The images are being processed and will be added in a future update.
This dataset contains a comprehensive collection of visual question-answering (VQA) pairs generated from official documentation of 18 major airline companies.
About the Creator
I'm David Soeiro-Vuong, an engineering student specializing in Computer Science, Big Data, and AI, currently working as an… See the full description on the dataset page: https://huggingface.co/datasets/Davidsv/airline-vision-dataset.harmony-visionai-tool-pool-jewelry-vision
AI Tool Pool Jewelry Vision Dataset
Dataset Description
This dataset contains 5,130 jewelry images organized into 5 categories for computer vision tasks. The dataset was originally created and hosted on Roboflow Universe.
Categories
Bracelet: Bracelet jewelry images
Earrings: Earring jewelry images
Necklace: Necklace jewelry images
Pendant: Pendant jewelry images
Ring: Ring jewelry images
Dataset Structure
AI-Tool-Pool-Jewelry-Vision/
├── train/… See the full description on the dataset page: https://huggingface.co/datasets/bzcasper/ai-tool-pool-jewelry-vision.
