CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-HTML 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.textimage-to-text100M<n<1B100 likes739k downloads2y agoHugging Face02apoidea /pubtabnet-htmlimagevisual-question-answering100K<n<1M24 likes900 downloads2y agoHugging Face03ranWang /UN_Sitemap_Multilingual_HTML_Corpus Dataset Card for "UN Sitemap Multilingual HTML Corpus" Update Time +8:00 2023-3-23 17:25:20 Dataset Summary 此数据集是从联合国网站提供的的sitemap中爬取的,包含了各种语言的HTML文件,并按照语言进行分类。数据集包含了不同语言的文章、新闻等联合国文本。数据集旨在为研究人员、学者和语言技术开发人员提供一个多语言文本集,可用于各种自然语言处理任务和应用。 数据集包括以下语言:汉语(zh)、英语(en)、西班牙语(ar)、俄语(ru)、西班牙语(es)、法语(fr)。 Dataset Structure Data Instances 数据集文件大小: 约 14 GB 一个 'zh' 的例子如下: { 'uuid': 'a154688c-b385-4d2a-bec7-f239f1397d21', 'url':… See the full description on the dataset page: https://huggingface.co/datasets/ranWang/UN_Sitemap_Multilingual_HTML_Corpus.text100K<n<1M3 likes533 downloads3y agoHugging Face04MAsad789565 /HTML-CSS-Website# Dataset This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant. # Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com) textn<1K4 likes499 downloads3y agoHugging Face05cognaize /table-image-html-pairsimage10K<n<100K0 likes244 downloads6mo agoHugging Face06apoidea /fintabnet-htmlimage10K<n<100K9 likes211 downloads2y agoHugging Face07hardikg2907 /github-code-html-css-1text1M<n<10M2 likes137 downloads2y agoHugging Face08LangAGI-Lab /Mind2Web-HTML-cleaned-lite-with-desc_w_taotabular1K<n<10K1 likes128 downloads2y agoHugging Face09hardikg2907 /github-code-html-css-2text1M<n<10M0 likes108 downloads2y agoHugging Face10williambrach /html-query-text-HtmlRAG html-query-text-HtmlRAG Warning: This dataset is under development and its content is subject to change! This dataset is a processed and cleaned version of the zstanjj/HtmlRAG-train dataset. It has been specifically prepared for task of HTML cleaning. 🚀 Supported Tasks This dataset is primarily designed for: HTML Cleaning: Training models to take the messy html as input and generate the cleaned_html or cleaned_text as output. Question Answering: Training models to… See the full description on the dataset page: https://huggingface.co/datasets/williambrach/html-query-text-HtmlRAG.textfeature-extraction10K<n<100K0 likes106 downloads11mo agoHugging Face11katphlab /pubtables-htmlimage100K<n<1M0 likes102 downloads2y agoHugging Face12tom-010 /enwiki-articles-html-2410tabular1M<n<10M0 likes101 downloads2y agoHugging Face132packer /html_game_gentextn<1K2 likes96 downloads1y agoHugging Face14hardikg2907 /github-code-html-csstext100K<n<1M4 likes91 downloads2y agoHugging Face15domofon /50k-HTML-PRETRAIN 50k-HTML-PRETRAIN Pretrain-style pairs: an English site assignment and a complete HTML document that implements it. Each row is one assignment (user) and one HTML page (assistant). A handful of rows include a PNG screenshot of the page; the rest leave screenshot empty. Split split n train 57,617 10 rows have a PNG in screenshot / images/. The other rows have a null screenshot. from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/domofon/50k-HTML-PRETRAIN.imagetext-generation10K<n<100K0 likes80 downloads20d agoHugging Face16hardikg2907 /github-code-html-css-split-3text1M<n<10M1 likes79 downloads2y agoHugging Face17Sweaterdog /website-html-2kThis dataset was a subset code tasks 33k filtered exclusively for high-quality websites. The filtering includes: Minimum of 100 lines Includes CSS or Javascript content The filtering was fairly light, as for stricter cleaning I only had 1,100 examples, which I considered to be low. Having 900 more examples is worth it in my opinion. I have also included a python file which grabs the HTML content from random responses and saves those as files 1.html ==> 5.html to view them. text1K<n<10K0 likes73 downloads8mo agoHugging Face18mdhasnainali /job-html-to-jsontext10K<n<100K2 likes70 downloads1y agoHugging Face19LangAGI-Lab /Mind2Web-HTML-cleaned-lite-with-desc_w_tao_value_rationaletabular1K<n<10K0 likes67 downloads2y agoHugging Face20LangAGI-Lab /mind2web_html_observationtext1K<n<10K1 likes60 downloads2y agoHugging Face21puyang2025 /phish_html Phishing Dataset (Parquet) Dataset Summary This dataset contains HTML pages labeled as benign or malicious, collected from confirmed phishing URLs reported on PhishTank across 2018-2020. The data here is a direct conversion of the Kaggle dataset "phishing-dataset" by asifejazitu into a Hugging Face-compatible Parquet file. Data Files train.parquet eval.parquet test.parquet Data Fields text (string): HTML content of the page. label (class): benign… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/phish_html.texttext-classification100K<n<1M1 likes57 downloads8mo agoHugging Face22chathuranga-jayanath /context-5-finmath-times4j-html-mavendoxia-wro4j-guava-supercsv-balanced-10k-prompt-1tabular100K<n<1M0 likes45 downloads3y agoHugging Face23LangAGI-Lab /Mind2Web-HTML-cleaned-lite-with-refined-tao-formertext1K<n<10K1 likes45 downloads2y agoHugging Face24hpprc /llmjp-warp-htmlllm-jp-corpus-v3のwarp_htmlのうちlevel2フィルタリングされたデータをHFフォーマットに変換し、各データに付与されたURLから元記事のタイトルを取得可能なものについては取得して付与したデータセットです。 ライセンスは元ページに従いCC-BY 4.0とします。 text100K<n<1M3 likes42 downloads2y agoHugging Face25tinhpx2911 /thanhnien_raw_htmltext1M<n<10M0 likes41 downloads3y agoHugging Face26nhhsag12 /pubtabnet-with-htmlimage100K<n<1M0 likes40 downloads11mo agoHugging Face27yyupenn /HTMLDocumentPipeline_manual_claude_2 Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K0 likes39 downloads2y agoHugging Face28espsluar /crawlerlm-html-to-json CrawlerLM: HTML to JSON Extraction A synthetic instruction-tuning dataset for training language models to extract structured JSON from HTML. Dataset Description This dataset contains HTML paired with structured JSON extraction tasks in chat format. It's designed for fine-tuning small language models to perform structured data extraction from messy, real-world HTML across multiple domains. Key Features 447 examples in instruction-tuning chat format Real HTML… See the full description on the dataset page: https://huggingface.co/datasets/espsluar/crawlerlm-html-to-json.texttext-generationn<1K2 likes38 downloads9mo agoHugging Face29sanjaymalladi /dataviz-html-dataset DataViz HTML Dashboard Dataset 100 HTML dashboard files built with the dash-viz-kit library — 10 themes, 12 chart types (ApexCharts + ECharts), zero config, declarative HTML. Structure data/train-00000-of-00001.parquet — Main dataset in Parquet format data/*.csv — CSV data files used by 15 dashboards README.md — Dataset card Columns Column Type Description filename string File name of the dashboard title string Human-readable title… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-html-dataset.tabulartext-generationn<1K0 likes37 downloads2mo agoHugging Face30Harveyaot /html-visual-design-preference-dataset-1ktextn<1K1 likes35 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.