CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-HTML 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.textimage-to-text100M<n<1B98 likes691k downloads2y agoHugging Face02apoidea /pubtabnet-htmlimagevisual-question-answering100K<n<1M23 likes946 downloads2y agoHugging Face03Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_html llm-jp-corpus-v4 — ja_sip_comprehensive_html Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_html Files: 181 × jsonl.gz (23.4 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.texttext-generation1M<n<10M0 likes659 downloads2mo agoHugging Face04Podtech /llm-jp-corpus-v4-ja_warp_html llm-jp-corpus-v4 — ja_warp_html Mirror of the ja/ja_warp_html sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_warp_html Files: 45 × jsonl.gz (1.6 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_html.texttext-generation1M<n<10M0 likes207 downloads2mo agoHugging Face05sfd-anonymous /html-table-reconstruction-benchmark HTML Table Reconstruction Benchmark This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation. The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.table-question-answeringn<1K0 likes194 downloads5mo agoHugging Face06thekevinscott /geocities-prompt-html GeoCities prompt → HTML — fine-tune Fine-tunes Gemma-4-E2B-it (LoRA) to generate a full, vintage-style HTML page from a plain-language description. This repo holds the dataset and the training scripts so the whole thing runs from one place. What's in here dataset.jsonl — the training data: one {"prompt": ..., "html": ...} per line. train_geocities.py — training entrypoint (loads this JSONL format). train-geocities-5090.sh — launch tuned for a 32 GB card (bf16… See the full description on the dataset page: https://huggingface.co/datasets/thekevinscott/geocities-prompt-html.texttext-generation1K<n<10K2 likes91 downloads3mo agoHugging Face07domofon /50k-HTML-PRETRAIN 50k-HTML-PRETRAIN Pretrain-style pairs: an English site assignment and a complete HTML document that implements it. Each row is one assignment (user) and one HTML page (assistant). A handful of rows include a PNG screenshot of the page; the rest leave screenshot empty. Split split n train 57,617 10 rows have a PNG in screenshot / images/. The other rows have a null screenshot. from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/domofon/50k-HTML-PRETRAIN.imagetext-generation10K<n<100K0 likes80 downloads16d agoHugging Face08sanjaymalladi /dataviz-html-dataset DataViz HTML Dashboard Dataset 100 HTML dashboard files built with the dash-viz-kit library — 10 themes, 12 chart types (ApexCharts + ECharts), zero config, declarative HTML. Structure data/train-00000-of-00001.parquet — Main dataset in Parquet format data/*.csv — CSV data files used by 15 dashboards README.md — Dataset card Columns Column Type Description filename string File name of the dashboard title string Human-readable title… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-html-dataset.tabulartext-generationn<1K0 likes46 downloads2mo agoHugging Face09espsluar /crawlerlm-html-to-json CrawlerLM: HTML to JSON Extraction A synthetic instruction-tuning dataset for training language models to extract structured JSON from HTML. Dataset Description This dataset contains HTML paired with structured JSON extraction tasks in chat format. It's designed for fine-tuning small language models to perform structured data extraction from messy, real-world HTML across multiple domains. Key Features 447 examples in instruction-tuning chat format Real HTML… See the full description on the dataset page: https://huggingface.co/datasets/espsluar/crawlerlm-html-to-json.texttext-generationn<1K2 likes36 downloads9mo agoHugging Face10DogukanUrker /ui-distill-html-648 ui-distill-html-648 648 single-file HTML UI components, generated by Ornith-1.0-35B on a single RTX 3060 12GB, paired with the build request that produced each one. Built to fine-tune a 3B model into writing UI (DogukanUrker/ui-distill-3b), but it stands on its own — distill your own student from it. The interesting part The instructions in this dataset are not the prompts that generated the HTML. The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.texttext-generationn<1K0 likes32 downloads2mo agoHugging Face11AI-Culture-Commons /philosophy-culture-translations-html-csv AI-Culture Philosophy and Culture Translations CSV + HTML Corpus The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind. This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.imagetranslation1K<n<10K2 likes30 downloads1y agoHugging Face12kogai /full-html-stying-dataset-generated-css-from-style-plan Generated CSS From Style Plan kogai/full-html-stying-dataset-generated-css-from-style-plan contains generated_css_from_style_plan.jsonl, a JSONL dataset with 44458 synthetic examples. Model-generated CSS outputs conditioned on source HTML, user style requests, and structured style plans. Schema chat_template_overhead_tokens: field present in the JSONL records. created_at: field present in the JSONL records. input_html: source HTML before Tailwind classes are… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-generated-css-from-style-plan.text-generation10K<n<100K0 likes29 downloads2mo agoHugging Face13Azzindani /ID_Supreme_Court_HTML ⚖️ Indonesian Supreme Court HTML Dataset (ID_Supreme_Court_HTML) This repository contains a massive collection of official court decisions from the Supreme Court of the Republic of Indonesia (Mahkamah Agung RI), preserved in their original HTML structure. 🏛️ Unlike plain text datasets, this collection retains the underlying web formatting from the official Direktori Putusan, making it an invaluable resource for developers building legal tech tools that require structural context… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Supreme_Court_HTML.text-generation0 likes27 downloads7mo agoHugging Face14tkattkat101 /html-compression HTML Compression Dataset This dataset is designed for fine-tuning language models to understand and generate compressed HTML notation. Dataset Description The dataset contains pairs of standard HTML and compressed HTML using a custom compression algorithm that: Abbreviates common HTML tags (div → d, span → s, etc.) Shortens attribute names (class → c, style → st, etc.) Compresses CSS patterns (display: flex → d:f, etc.) Removes unnecessary whitespace and comments… See the full description on the dataset page: https://huggingface.co/datasets/tkattkat101/html-compression.texttext-generation1K<n<10K0 likes23 downloads1y agoHugging Face15sanjaymalladi /dataviz-charts-html-dataset DataViz Individual Charts HTML Dataset 120 individual chart HTML files — 12 chart types × 10 themes — for teaching models to generate themed single-chart visualizations. Structure data/train-00000-of-00001.parquet — Main dataset in Parquet format README.md — Dataset card Columns Column Type Description filename string treemap-tech-innovation.html title string e.g. "Bar Chart — Midnight Galaxy" chart_type string bar, line, area, pie… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-charts-html-dataset.tabulartext-generationn<1K0 likes21 downloads2mo agoHugging Face16kogai /full-html-stying-dataset-motion-geometry-styles Full HTML Tailwind Motion Geometry Styles kogai/full-html-stying-dataset-motion-geometry-styles contains phase1_dataset_10k_motion_geometry_style.jsonl, a JSONL dataset with 10000 synthetic examples. Tailwind styling transforms emphasizing motion, geometry, and expressive visual direction for full HTML pages. Schema input_html: source HTML before Tailwind classes are added. output_html: transformed HTML with Tailwind utility classes applied. style_phrase:… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-motion-geometry-styles.text-generation10K<n<100K0 likes18 downloads2mo agoHugging Face17kogai /full-html-stying-dataset-nested-element-styles Full HTML Tailwind Nested Element Styles kogai/full-html-stying-dataset-nested-element-styles contains phase1_dataset_10k_1_elem_in_1_elem_style.jsonl, a JSONL dataset with 10000 synthetic examples. Tailwind styling transforms for simple single-element inputs that often include nested inline structure. Schema input_html: source HTML before Tailwind classes are added. output_html: transformed HTML with Tailwind utility classes applied. style_phrase: normalized… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-nested-element-styles.text-generation10K<n<100K0 likes17 downloads2mo agoHugging Face18Boakpe /deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript. texttext-generationn<1K1 likes14 downloads8mo agoHugging Face19erv1n /geo_html_200 GEO HTML 200 Dataset A curated dataset of 200 web documents for Generative Engine Optimization (GEO) research. Features Column Description doc_id Unique document identifier url Source URL cleaned_text Parsed plain text content cleaned_text_length Character count query Associated search query title Document title topic_tags Topic classification Usage from datasets import load_dataset ds = load_dataset("erv1n/geo_html_200") tabulartext-generationn<1K0 likes11 downloads8mo agoHugging Face20kogai /full-html-stying-dataset-complex-layouts Full HTML Tailwind Complex Layouts kogai/full-html-stying-dataset-complex-layouts contains phase1_dataset_10k_complex.jsonl, a JSONL dataset with 10000 synthetic examples. Tailwind styling transforms for larger multi-section HTML layouts such as dashboards, marketing pages, and shells. Schema input_html: source HTML before Tailwind classes are added. output_html: transformed HTML with Tailwind utility classes applied. style_phrase: normalized style request or… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-complex-layouts.text-generation10K<n<100K0 likes10 downloads2mo agoHugging Face21kogai /full-html-stying-dataset-style-planning-dataset HTML Styling Plan Dataset kogai/full-html-stying-dataset-style-planning-dataset contains style_planning_dataset.jsonl, a JSONL dataset with 46374 synthetic examples. Structured style-planning records that pair source HTML and style requests with a generated styling plan. Schema created_at: field present in the JSONL records. input_html: source HTML before Tailwind classes are added. input_style: user-facing styling request or directive. model: field present in… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-style-planning-dataset.text-generation10K<n<100K0 likes9 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.