datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
c4-en-html-with-metadataaudio-htmlc4-en-html-with-metadata-ppl-cleanFile list:
"c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.c4-en-html-with-training_metadata_allMind2Web-HTML-cleaned-lite-with-desc_w_taoenwiki-articles-html-2410html-css-js-cot
🌐 HTML/CSS/JS Reasoning Traces Dataset
A high-quality, large-scale dataset of complex HTML, CSS, and JavaScript programming questions and model reasoning traces.
📊 Dataset Overview
This repository contains a comprehensively structured dataset of reasoning traces for frontend web development tasks. The data maps intricate, multi-step prompts to step-by-step reasoning solutions generated by advanced Language Models.
It is designed for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/html-css-js-cot.Mind2Web-HTML-cleaned-lite-with-desc_w_tao_value_rationaledataviz-html-dataset
DataViz HTML Dashboard Dataset
100 HTML dashboard files built with the dash-viz-kit library — 10 themes, 12 chart types (ApexCharts + ECharts), zero config, declarative HTML.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
data/*.csv — CSV data files used by 15 dashboards
README.md — Dataset card
Columns
Column
Type
Description
filename
string
File name of the dashboard
title
string
Human-readable title… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-html-dataset.context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-balanced-10k-prompt-1context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-20000-prompt-3context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-10000-prompt-1context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-30000-prompt-1context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-30000-prompt-1context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-30000-prompt-3context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-10000-prompt-0dataviz-charts-html-dataset
DataViz Individual Charts HTML Dataset
120 individual chart HTML files — 12 chart types × 10 themes — for teaching models to generate themed single-chart visualizations.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
README.md — Dataset card
Columns
Column
Type
Description
filename
string
treemap-tech-innovation.html
title
string
e.g. "Bar Chart — Midnight Galaxy"
chart_type
string
bar, line, area, pie… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-charts-html-dataset.context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-1000-prompt-1context-5-from-finmath-time4j-html-mavendoxia-portion-0.4-prompt-1context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-10000-prompt-1html-ai-battle-experiment-tracker
HTML AI Battle Experiment Tracker
This dataset contains the experiment tracker for the paper:
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
Author: Diego Cabezas Palacios
arXiv: 2605.06707Code and materials: https://github.com/diegocp01/html_ai_battle
Dataset Summary
This dataset supports a longitudinal observational comparison of first-output LLM web generation across public chat interfaces.… See the full description on the dataset page: https://huggingface.co/datasets/diegocp01/html-ai-battle-experiment-tracker.context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-20000-prompt-1MIMIQ-Htmlgeo_html_200_fullcontext-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-1000-prompt-2context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-10000-prompt-3geo_html_200
GEO HTML 200 Dataset
A curated dataset of 200 web documents for Generative Engine Optimization (GEO) research.
Features
Column
Description
doc_id
Unique document identifier
url
Source URL
cleaned_text
Parsed plain text content
cleaned_text_length
Character count
query
Associated search query
title
Document title
topic_tags
Topic classification
Usage
from datasets import load_dataset
ds = load_dataset("erv1n/geo_html_200")
context-5-from-finmath-time4j-html-mavendoxia-portion-0.1-prompt-1chunked-enwiki-ns0-20250301-enterprise-html20250301.jsonl contains preprocessed text chunks from the wikipedia 20250301 html dump.
dead_letter_queue.jsonl contains articles that couldn't be processed due to formatitng or missing tags.
en_redirection_map.pkl can be deserialize to avoid building redirection_map everytime preprocessing_wikipedia_html_dump runs.
html-only
