datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-1T-HTML
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.pubtabnet-htmlUN_Sitemap_Multilingual_HTML_Corpus
Dataset Card for "UN Sitemap Multilingual HTML Corpus"
Update Time +8:00 2023-3-23 17:25:20
Dataset Summary
此数据集是从联合国网站提供的的sitemap中爬取的,包含了各种语言的HTML文件,并按照语言进行分类。数据集包含了不同语言的文章、新闻等联合国文本。数据集旨在为研究人员、学者和语言技术开发人员提供一个多语言文本集,可用于各种自然语言处理任务和应用。
数据集包括以下语言:汉语(zh)、英语(en)、西班牙语(ar)、俄语(ru)、西班牙语(es)、法语(fr)。
Dataset Structure
Data Instances
数据集文件大小: 约 14 GB
一个 'zh' 的例子如下:
{
'uuid': 'a154688c-b385-4d2a-bec7-f239f1397d21',
'url':… See the full description on the dataset page: https://huggingface.co/datasets/ranWang/UN_Sitemap_Multilingual_HTML_Corpus.HTML-CSS-Website# Dataset
This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant.
# Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com)
table-image-html-pairsfintabnet-htmlgithub-code-html-css-1Mind2Web-HTML-cleaned-lite-with-desc_w_taogithub-code-html-css-2html-query-text-HtmlRAG
html-query-text-HtmlRAG
Warning: This dataset is under development and its content is subject to change!
This dataset is a processed and cleaned version of the zstanjj/HtmlRAG-train dataset. It has been specifically prepared for task of HTML cleaning.
🚀 Supported Tasks
This dataset is primarily designed for:
HTML Cleaning: Training models to take the messy html as input and generate the cleaned_html or cleaned_text as output.
Question Answering: Training models to… See the full description on the dataset page: https://huggingface.co/datasets/williambrach/html-query-text-HtmlRAG.pubtables-htmlenwiki-articles-html-2410html_game_gengithub-code-html-css50k-HTML-PRETRAIN
50k-HTML-PRETRAIN
Pretrain-style pairs: an English site assignment and a complete HTML document that implements it.
Each row is one assignment (user) and one HTML page (assistant). A handful of rows include a PNG screenshot of the page; the rest leave screenshot empty.
Split
split
n
train
57,617
10 rows have a PNG in screenshot / images/. The other rows have a null screenshot.
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/domofon/50k-HTML-PRETRAIN.github-code-html-css-split-3website-html-2kThis dataset was a subset code tasks 33k filtered exclusively for high-quality websites.
The filtering includes:
Minimum of 100 lines
Includes CSS or Javascript content
The filtering was fairly light, as for stricter cleaning I only had 1,100 examples, which I considered to be low. Having 900 more examples is worth it in my opinion.
I have also included a python file which grabs the HTML content from random responses and saves those as files 1.html ==> 5.html to view them.
job-html-to-jsonMind2Web-HTML-cleaned-lite-with-desc_w_tao_value_rationalemind2web_html_observationphish_html
Phishing Dataset (Parquet)
Dataset Summary
This dataset contains HTML pages labeled as benign or malicious, collected from confirmed phishing URLs reported on PhishTank across 2018-2020. The data here is a direct conversion of the Kaggle dataset "phishing-dataset" by asifejazitu into a Hugging Face-compatible Parquet file.
Data Files
train.parquet
eval.parquet
test.parquet
Data Fields
text (string): HTML content of the page.
label (class): benign… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/phish_html.context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-balanced-10k-prompt-1Mind2Web-HTML-cleaned-lite-with-refined-tao-formerllmjp-warp-htmlllm-jp-corpus-v3のwarp_htmlのうちlevel2フィルタリングされたデータをHFフォーマットに変換し、各データに付与されたURLから元記事のタイトルを取得可能なものについては取得して付与したデータセットです。
ライセンスは元ページに従いCC-BY 4.0とします。
thanhnien_raw_htmlpubtabnet-with-htmlHTMLDocumentPipeline_manual_claude_2
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
crawlerlm-html-to-json
CrawlerLM: HTML to JSON Extraction
A synthetic instruction-tuning dataset for training language models to extract structured JSON from HTML.
Dataset Description
This dataset contains HTML paired with structured JSON extraction tasks in chat format. It's designed for fine-tuning small language models to perform structured data extraction from messy, real-world HTML across multiple domains.
Key Features
447 examples in instruction-tuning chat format
Real HTML… See the full description on the dataset page: https://huggingface.co/datasets/espsluar/crawlerlm-html-to-json.dataviz-html-dataset
DataViz HTML Dashboard Dataset
100 HTML dashboard files built with the dash-viz-kit library — 10 themes, 12 chart types (ApexCharts + ECharts), zero config, declarative HTML.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
data/*.csv — CSV data files used by 15 dashboards
README.md — Dataset card
Columns
Column
Type
Description
filename
string
File name of the dashboard
title
string
Human-readable title… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-html-dataset.html-visual-design-preference-dataset-1k
