datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-1T-HTML
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.pubtabnet-htmlllm-jp-corpus-v4-ja_sip_comprehensive_html
llm-jp-corpus-v4 — ja_sip_comprehensive_html
Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_html
Files: 181 × jsonl.gz (23.4 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.llm-jp-corpus-v4-ja_warp_html
llm-jp-corpus-v4 — ja_warp_html
Mirror of the ja/ja_warp_html sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_warp_html
Files: 45 × jsonl.gz (1.6 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_html.html-table-reconstruction-benchmark
HTML Table Reconstruction Benchmark
This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation.
The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.geocities-prompt-html
GeoCities prompt → HTML — fine-tune
Fine-tunes Gemma-4-E2B-it (LoRA) to generate a full, vintage-style HTML page
from a plain-language description. This repo holds the dataset and the
training scripts so the whole thing runs from one place.
What's in here
dataset.jsonl — the training data: one {"prompt": ..., "html": ...} per line.
train_geocities.py — training entrypoint (loads this JSONL format).
train-geocities-5090.sh — launch tuned for a 32 GB card (bf16… See the full description on the dataset page: https://huggingface.co/datasets/thekevinscott/geocities-prompt-html.50k-HTML-PRETRAIN
50k-HTML-PRETRAIN
Pretrain-style pairs: an English site assignment and a complete HTML document that implements it.
Each row is one assignment (user) and one HTML page (assistant). A handful of rows include a PNG screenshot of the page; the rest leave screenshot empty.
Split
split
n
train
57,617
10 rows have a PNG in screenshot / images/. The other rows have a null screenshot.
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/domofon/50k-HTML-PRETRAIN.dataviz-html-dataset
DataViz HTML Dashboard Dataset
100 HTML dashboard files built with the dash-viz-kit library — 10 themes, 12 chart types (ApexCharts + ECharts), zero config, declarative HTML.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
data/*.csv — CSV data files used by 15 dashboards
README.md — Dataset card
Columns
Column
Type
Description
filename
string
File name of the dashboard
title
string
Human-readable title… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-html-dataset.crawlerlm-html-to-json
CrawlerLM: HTML to JSON Extraction
A synthetic instruction-tuning dataset for training language models to extract structured JSON from HTML.
Dataset Description
This dataset contains HTML paired with structured JSON extraction tasks in chat format. It's designed for fine-tuning small language models to perform structured data extraction from messy, real-world HTML across multiple domains.
Key Features
447 examples in instruction-tuning chat format
Real HTML… See the full description on the dataset page: https://huggingface.co/datasets/espsluar/crawlerlm-html-to-json.ui-distill-html-648
ui-distill-html-648
648 single-file HTML UI components, generated by
Ornith-1.0-35B on a single
RTX 3060 12GB, paired with the build request that produced each one.
Built to fine-tune a 3B model into writing UI
(DogukanUrker/ui-distill-3b), but
it stands on its own — distill your own student from it.
The interesting part
The instructions in this dataset are not the prompts that generated the HTML.
The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.philosophy-culture-translations-html-csv
AI-Culture Philosophy and Culture Translations CSV + HTML Corpus
The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind.
This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.full-html-stying-dataset-generated-css-from-style-plan
Generated CSS From Style Plan
kogai/full-html-stying-dataset-generated-css-from-style-plan contains generated_css_from_style_plan.jsonl, a JSONL dataset with 44458 synthetic examples. Model-generated CSS outputs conditioned on source HTML, user style requests, and structured style plans.
Schema
chat_template_overhead_tokens: field present in the JSONL records.
created_at: field present in the JSONL records.
input_html: source HTML before Tailwind classes are… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-generated-css-from-style-plan.ID_Supreme_Court_HTML
⚖️ Indonesian Supreme Court HTML Dataset (ID_Supreme_Court_HTML)
This repository contains a massive collection of official court decisions from the Supreme Court of the Republic of Indonesia (Mahkamah Agung RI), preserved in their original HTML structure. 🏛️
Unlike plain text datasets, this collection retains the underlying web formatting from the official Direktori Putusan, making it an invaluable resource for developers building legal tech tools that require structural context… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Supreme_Court_HTML.html-compression
HTML Compression Dataset
This dataset is designed for fine-tuning language models to understand and generate compressed HTML notation.
Dataset Description
The dataset contains pairs of standard HTML and compressed HTML using a custom compression algorithm that:
Abbreviates common HTML tags (div → d, span → s, etc.)
Shortens attribute names (class → c, style → st, etc.)
Compresses CSS patterns (display: flex → d:f, etc.)
Removes unnecessary whitespace and comments… See the full description on the dataset page: https://huggingface.co/datasets/tkattkat101/html-compression.dataviz-charts-html-dataset
DataViz Individual Charts HTML Dataset
120 individual chart HTML files — 12 chart types × 10 themes — for teaching models to generate themed single-chart visualizations.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
README.md — Dataset card
Columns
Column
Type
Description
filename
string
treemap-tech-innovation.html
title
string
e.g. "Bar Chart — Midnight Galaxy"
chart_type
string
bar, line, area, pie… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-charts-html-dataset.full-html-stying-dataset-motion-geometry-styles
Full HTML Tailwind Motion Geometry Styles
kogai/full-html-stying-dataset-motion-geometry-styles contains phase1_dataset_10k_motion_geometry_style.jsonl, a JSONL dataset with 10000 synthetic examples. Tailwind styling transforms emphasizing motion, geometry, and expressive visual direction for full HTML pages.
Schema
input_html: source HTML before Tailwind classes are added.
output_html: transformed HTML with Tailwind utility classes applied.
style_phrase:… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-motion-geometry-styles.full-html-stying-dataset-nested-element-styles
Full HTML Tailwind Nested Element Styles
kogai/full-html-stying-dataset-nested-element-styles contains phase1_dataset_10k_1_elem_in_1_elem_style.jsonl, a JSONL dataset with 10000 synthetic examples. Tailwind styling transforms for simple single-element inputs that often include nested inline structure.
Schema
input_html: source HTML before Tailwind classes are added.
output_html: transformed HTML with Tailwind utility classes applied.
style_phrase: normalized… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-nested-element-styles.deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript.
geo_html_200
GEO HTML 200 Dataset
A curated dataset of 200 web documents for Generative Engine Optimization (GEO) research.
Features
Column
Description
doc_id
Unique document identifier
url
Source URL
cleaned_text
Parsed plain text content
cleaned_text_length
Character count
query
Associated search query
title
Document title
topic_tags
Topic classification
Usage
from datasets import load_dataset
ds = load_dataset("erv1n/geo_html_200")
full-html-stying-dataset-complex-layouts
Full HTML Tailwind Complex Layouts
kogai/full-html-stying-dataset-complex-layouts contains phase1_dataset_10k_complex.jsonl, a JSONL dataset with 10000 synthetic examples. Tailwind styling transforms for larger multi-section HTML layouts such as dashboards, marketing pages, and shells.
Schema
input_html: source HTML before Tailwind classes are added.
output_html: transformed HTML with Tailwind utility classes applied.
style_phrase: normalized style request or… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-complex-layouts.full-html-stying-dataset-style-planning-dataset
HTML Styling Plan Dataset
kogai/full-html-stying-dataset-style-planning-dataset contains style_planning_dataset.jsonl, a JSONL dataset with 46374 synthetic examples. Structured style-planning records that pair source HTML and style requests with a generated styling plan.
Schema
created_at: field present in the JSONL records.
input_html: source HTML before Tailwind classes are added.
input_style: user-facing styling request or directive.
model: field present in… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-style-planning-dataset.
