datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mint-1t-html-images-gte6-sample
Size: 6769158 images sampled from Mint-1t-html
Criteria: Data entries with greater than or equal to 6 images (gte6)
pubtabnet-htmlwebui-react-htmlcssjs-8740
WebUI React + HTML/CSS/JS 8,740
Curated export from ronantakizawa/webui containing every row where framework = react, plus 4,000 additional rows where framework = vanilla.
Screenshots: 8,740
React / vanilla HTML-CSS-JS rows: 4,740 / 4,000
Unique sample IDs: 2,914
Train / validation / test: 7,501 / 456 / 783
Viewports: 2,914 desktop / 2,913 mobile / 2,913 tablet
Images are stored as real image files and verified with Pillow.
viewer.html is a self-contained, paginated local… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/webui-react-htmlcssjs-8740.fintabnet-htmlhtml-sampletable-image-html-pairshtml-gen-assetsGutenberg-Arabic-OCR-HTML-Pages
Gutenberg Arabic HTML-Page Dataset
📖 Dataset Description
The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language.
The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.pubtables-htmlmc_html_screenshot50k-HTML-PRETRAIN
50k-HTML-PRETRAIN
Pretrain-style pairs: an English site assignment and a complete HTML document that implements it.
Each row is one assignment (user) and one HTML page (assistant). A handful of rows include a PNG screenshot of the page; the rest leave screenshot empty.
Split
split
n
train
57,617
10 rows have a PNG in screenshot / images/. The other rows have a null screenshot.
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/domofon/50k-HTML-PRETRAIN.pubtabnet-with-htmlHTMLDocumentPipeline_manual_claude_2
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
codet5-finetuned-htmlimages1htmls10kd.HTML
d.HTML
Overview
d.HTML is a lightweight dataset designed for Image-to-Text OCR and structured HTML reconstruction tasks. The dataset pairs document page images with corresponding markup outputs, primarily in HTML (and occasionally Markdown-like structures). It is intended for evaluating and training multimodal models that convert visual documents into structured, machine-readable formats. The dataset focuses on preserving document structure, including headings, paragraphs… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/d.HTML.fintabnet-html-testphilosophy-culture-translations-html-csv
AI-Culture Philosophy and Culture Translations CSV + HTML Corpus
The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind.
This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.HTMLDocumentPipeline_form_claude_2
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
financial-statement-table-htmlscifi_html
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
Multimodal-Mind2Web-HTML-WM-messagesarticle_htmlhtml_document_point
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
HTMLDocumentPipeline_manual_agenda_0
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
html_document
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
MCFTable-HTMLMultimodal-Mind2Web-HTML-WM-messages-testhtml_chart
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
image-html
