datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mind2Web-HTML-cleaned-lite-with-desc_w_taoenwiki-articles-html-2410Mind2Web-HTML-cleaned-lite-with-desc_w_tao_value_rationalecontext-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-balanced-10k-prompt-1dataviz-html-dataset
DataViz HTML Dashboard Dataset
100 HTML dashboard files built with the dash-viz-kit library — 10 themes, 12 chart types (ApexCharts + ECharts), zero config, declarative HTML.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
data/*.csv — CSV data files used by 15 dashboards
README.md — Dataset card
Columns
Column
Type
Description
filename
string
File name of the dashboard
title
string
Human-readable title… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-html-dataset.context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-10000-prompt-1context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-20000-prompt-3dataviz-charts-html-dataset
DataViz Individual Charts HTML Dataset
120 individual chart HTML files — 12 chart types × 10 themes — for teaching models to generate themed single-chart visualizations.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
README.md — Dataset card
Columns
Column
Type
Description
filename
string
treemap-tech-innovation.html
title
string
e.g. "Bar Chart — Midnight Galaxy"
chart_type
string
bar, line, area, pie… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-charts-html-dataset.context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-30000-prompt-1context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-30000-prompt-1context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-10000-prompt-0context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-30000-prompt-3context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-1000-prompt-1context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-10000-prompt-1html-filteredcontext-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-20000-prompt-1context-5-from-finmath-time4j-html-mavendoxia-portion-0.4-prompt-1geo_html_200_fullcontext-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-1000-prompt-2context-5-from-finmath-time4j-html-mavendoxia-portion-0.1-prompt-1context-5-rhino-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-len-10000-prompt-3geo_html_200
GEO HTML 200 Dataset
A curated dataset of 200 web documents for Generative Engine Optimization (GEO) research.
Features
Column
Description
doc_id
Unique document identifier
url
Source URL
cleaned_text
Parsed plain text content
cleaned_text_length
Character count
query
Associated search query
title
Document title
topic_tags
Topic classification
Usage
from datasets import load_dataset
ds = load_dataset("erv1n/geo_html_200")
MIMIQ-Htmlhtml-onlytest_split_12dec_htmltrain_split_12dec_htmlhtml-annotated-1k
