CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01paralym /mint-1t-html-images-gte6-sample Size: 6769158 images sampled from Mint-1t-html Criteria: Data entries with greater than or equal to 6 images (gte6) image1M<n<10M0 likes1.1k downloads2y agoHugging Face02apoidea /pubtabnet-htmlimagevisual-question-answering100K<n<1M23 likes946 downloads2y agoHugging Face03Reubencf /webui-react-htmlcssjs-8740 WebUI React + HTML/CSS/JS 8,740 Curated export from ronantakizawa/webui containing every row where framework = react, plus 4,000 additional rows where framework = vanilla. Screenshots: 8,740 React / vanilla HTML-CSS-JS rows: 4,740 / 4,000 Unique sample IDs: 2,914 Train / validation / test: 7,501 / 456 / 783 Viewports: 2,914 desktop / 2,913 mobile / 2,913 tablet Images are stored as real image files and verified with Pillow. viewer.html is a self-contained, paginated local… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/webui-react-htmlcssjs-8740.imageimage-to-text1K<n<10K0 likes603 downloads2mo agoHugging Face04apoidea /fintabnet-htmlimage10K<n<100K9 likes219 downloads2y agoHugging Face05big-computer /html-sampleimagen<1K0 likes203 downloads2y agoHugging Face06cognaize /table-image-html-pairsimage10K<n<100K0 likes191 downloads6mo agoHugging Face07kaisenkang /html-gen-assetsimagen<1K0 likes153 downloads1mo agoHugging Face08FatimahEmadEldin /Gutenberg-Arabic-OCR-HTML-Pages Gutenberg Arabic HTML-Page Dataset 📖 Dataset Description The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language. The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.image1K<n<10K3 likes129 downloads1y agoHugging Face09katphlab /pubtables-htmlimage100K<n<1M0 likes105 downloads2y agoHugging Face10irotem98 /mc_html_screenshotimage1K<n<10K0 likes90 downloads1y agoHugging Face11domofon /50k-HTML-PRETRAIN 50k-HTML-PRETRAIN Pretrain-style pairs: an English site assignment and a complete HTML document that implements it. Each row is one assignment (user) and one HTML page (assistant). A handful of rows include a PNG screenshot of the page; the rest leave screenshot empty. Split split n train 57,617 10 rows have a PNG in screenshot / images/. The other rows have a null screenshot. from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/domofon/50k-HTML-PRETRAIN.imagetext-generation10K<n<100K0 likes80 downloads16d agoHugging Face12nhhsag12 /pubtabnet-with-htmlimage100K<n<1M0 likes38 downloads11mo agoHugging Face13yyupenn /HTMLDocumentPipeline_manual_claude_2 Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K0 likes36 downloads2y agoHugging Face14hadir1 /codet5-finetuned-htmlimages1image0 likes33 downloads1y agoHugging Face15irotem98 /htmls10kimage1 likes33 downloads1y agoHugging Face16prithivMLmods /d.HTML d.HTML Overview d.HTML is a lightweight dataset designed for Image-to-Text OCR and structured HTML reconstruction tasks. The dataset pairs document page images with corresponding markup outputs, primarily in HTML (and occasionally Markdown-like structures). It is intended for evaluating and training multimodal models that convert visual documents into structured, machine-readable formats. The dataset focuses on preserving document structure, including headings, paragraphs… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/d.HTML.imageimage-text-to-textn<1K1 likes32 downloads7mo agoHugging Face17tarun-menta /fintabnet-html-testimage1K<n<10K0 likes28 downloads2y agoHugging Face18AI-Culture-Commons /philosophy-culture-translations-html-csv AI-Culture Philosophy and Culture Translations CSV + HTML Corpus The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind. This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.imagetranslation1K<n<10K2 likes28 downloads1y agoHugging Face19yyupenn /HTMLDocumentPipeline_form_claude_2 Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K1 likes27 downloads2y agoHugging Face20apoidea /financial-statement-table-htmlimagen<1K6 likes26 downloads2y agoHugging Face21AjayP13 /scifi_html Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K0 likes19 downloads2y agoHugging Face22LangAGI-Lab /Multimodal-Mind2Web-HTML-WM-messagesimage1K<n<10K0 likes17 downloads2y agoHugging Face23megamattc /article_htmlimage10K<n<100K0 likes17 downloads1y agoHugging Face24nglebm19 /html_document_point Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K0 likes14 downloads1y agoHugging Face25yyupenn /HTMLDocumentPipeline_manual_agenda_0 Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K0 likes13 downloads2y agoHugging Face26nglebm19 /html_document Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K0 likes13 downloads1y agoHugging Face27mengluoye /MCFTable-HTMLimage100K<n<1M0 likes13 downloads1y agoHugging Face28LangAGI-Lab /Multimodal-Mind2Web-HTML-WM-messages-testimagen<1K0 likes12 downloads2y agoHugging Face29nglebm19 /html_chart Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K0 likes10 downloads1y agoHugging Face30d-manuardi /image-htmlimagen<1K1 likes9 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.