datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Gutenberg-Arabic-OCR-HTML-Pages
Gutenberg Arabic HTML-Page Dataset
📖 Dataset Description
The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language.
The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.buat-task-2SurgCOTBench
SurgCOTBench
SurgCOTBench is a reasoning-focused vision-language benchmark for robotic-assisted surgery. It contains frame-level question-answer pairs across robotic surgical procedures, covering five surgical scene-understanding tasks: action recognition, instrument recognition, action prediction, surgical outcome, and patient detail.
This release follows the dataset description and Table I statistics from:
Paper: SurgRAW: Multi-Agent Workflow with Chain of Thought Reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/FatNinjaaaaa/SurgCOTBench.amazon-product-reviews
