CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01isp-uv-es /Web_site_legacy0 likes3.9k downloads2y agoHugging Face02malaysia-ai /crawl-my-website6 likes3.4k downloads2y agoHugging Face03NoeFlandre /osm-polygon-website-tag OSM Polygon Website Dataset OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files. At a glance Polygons 1,726,474 With extracted text 1,192,980 Words of text 407,685,655 Languages 397 Regional sources 386 / 386 Duplicate objects removed 104,927 Candidates rejected 868,905,743 Status In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.tabular1M<n<10M2 likes2.4k downloads2d agoHugging Face04tianweiy /causvid_websitevideon<1K0 likes2.1k downloads2y agoHugging Face05silatus /1k_Website_Screenshots_and_Metadata Dataset Card for 1000 Website Screenshots with Metadata Dataset Summary Silatus is sharing, for free, a segment of a dataset that we are using to train a generative AI model for text-to-mockup conversions. This dataset was collected in December 2022 and early January 2023, so it contains very recent data from 1,000 of the world's most popular websites. You can get our larger 10,000 website dataset for free at: https://silatus.com/datasets This dataset includes: High-res… See the full description on the dataset page: https://huggingface.co/datasets/silatus/1k_Website_Screenshots_and_Metadata.imagetext-to-image1K<n<10K20 likes1.3k downloads4y agoHugging Face06Voxel51 /mind2web_multimodal_test_website Dataset Card for Multimodal Mind2Web "Cross-Website" Test Split Note: This dataset is the test split of the Cross-Website dataset introduced in the paper. This is a FiftyOne dataset with 1019 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_website.imageimage-classification1K<n<10K1 likes978 downloads1y agoHugging Face07FrancophonIA /Translations_Hungarian_public_websites [!NOTE] Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/18982 Description A webcrawl of 14 different websites covering parallel corpora of Hungarian with Polish, Czech, Swedish, Finnish, French, German, Italian, English and Slovenian Citation Translations of Hungarian from public websites (2022). Version 1.0. [Dataset (Text corpus)]. Source: European Language Grid. https://live.european-language-grid.eu/catalogue/corpus/18982 translation0 likes953 downloads1y agoHugging Face08shresthsamyak /phishing-website-screenshots Phishing Website Screenshots A dataset of 8,370 full-page website screenshots labelled as legitimate or phishing, intended for training and evaluating visual phishing-detection models. Contents Label label Images legitimate 0 7,924 phishing 1 446 Total 8,370 Screenshots were captured at a desktop viewport (1920×1080) as PNG images. Structure legitimate/<brand>/<page>.png phishing/<source>/<page>.png metadata.csv metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/shresthsamyak/phishing-website-screenshots.imageimage-classificationn<1K0 likes931 downloads2mo agoHugging Face09OpenRAL /website-media OpenRAL — website media Video clips shown in the "See it run" section of openral.com (benchmarks, simulation and on-hardware deployment runs). Each clip lives under <category>/<benchmark>_<rskill>_<success|fail>/ with three web-optimised assets: poster.jpg — first-frame thumbnail preview.mp4 — square 640px, muted (the autoscroll strip) full.mp4 — native aspect, ≤1080p, with audio (the expand modal) Generated and published by scripts/build-media.mjs in the website repo.… See the full description on the dataset page: https://huggingface.co/datasets/OpenRAL/website-media.imagen<1K0 likes824 downloads2mo agoHugging Face10APProjects /us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered Rebuilt 2026-09-24. Five state labor agencies — Connecticut, Michigan, New York, North Carolina and Pennsylvania — retired the web pages their older WARN Act layoff notices lived on. Their current pages start years later. This dataset is every notice in our file that came from one of those retired pages and is not on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.tabulartabular-classification1K<n<10K0 likes751 downloads46m agoHugging Face11FredZhang7 /malicious-website-features-2.4MImportant Notice: A subset of the URL dataset is from Kaggle, and the Kaggle datasets contained 10%-15% mislabelled data. See this dicussion I opened for some false positives. I have contacted Kaggle regarding their erroneous "Usability" score calculation for these unreliable datasets. The feature extraction methods shown here are not robust at all in 2023, and there're even silly mistakes in 3 functions: not_indexed_by_google, domain_registration_length, and age_of_domain. The features… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/malicious-website-features-2.4M.text-classification1M<n<10M6 likes541 downloads3y agoHugging Face12MAsad789565 /HTML-CSS-Website# Dataset This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant. # Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com) textn<1K4 likes512 downloads3y agoHugging Face13TheBaldDudeCo /website-gallery The Bald Dude Co. Website Gallery Web-ready portfolio media for The Bald Dude Co. The album structure and public file URLs are published in archive.json for use by the official website. Copyright The Bald Dude Co. All rights reserved. These files are not licensed for redistribution, resale, dataset compilation, or model training. image1K<n<10K0 likes414 downloads1d agoHugging Face14selfforcing-pp /websitevideon<1K0 likes384 downloads1y agoHugging Face15Philopater-Luka /Coptic-websites-scrappingtext10K<n<100K1 likes305 downloads20d agoHugging Face16snncn /shopify-websites Shopify Websites Dataset Dataset Description The Shopify Websites Dataset is a curated collection of 10,000 verified Shopify-powered e-commerce store URLs, providing researchers and analysts with a comprehensive resource for studying the Shopify e-commerce ecosystem. This dataset offers a diverse snapshot of real-world Shopify stores across various industries and geographies, making it valuable for market research, web scraping projects, and e-commerce platform analysis.… See the full description on the dataset page: https://huggingface.co/datasets/snncn/shopify-websites.text10K<n<100K1 likes215 downloads10mo agoHugging Face17huggingface /figma-Playground-Inference-for-PRO-s-Website1 likes192 downloads3y agoHugging Face18RelitLRM /website_assets3dn<1K0 likes172 downloads2y agoHugging Face19bs-modeling-metadata /website_metadata_c4The dataset is in the form of a json lines file with 1,20,000 examples, where an example consists of text (extracted from C4 English dataset) and metadata fields (website description extracted from Wikipedia). Example: { "text": "US10289222B2 - Handling of touch events in a browser environment - Google Patents\nHandling of touch events in a browser environment Download PDF\nUS10289222B2\nUS10289222B2 US13/857,848 US201313857848A US10289222B2 US 10289222 B2 US10289222 B2 US 10289222B2 US… See the full description on the dataset page: https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4.text10K<n<100K4 likes166 downloads5y agoHugging Face20Sadeeshkumargopalan /blog-website-dataimagen<1K0 likes151 downloads8h agoHugging Face21kwakrhkr59 /WebsiteFingerprinting-Raw0 likes138 downloads11mo agoHugging Face22Zexanima /website_screenshots_image_dataset Website Screenshots Image Dataset This dataset is obtainable here from roboflow.. Dataset Details Dataset Description Language(s) (NLP): [English] License: [MIT] Dataset Sources Source: [https://universe.roboflow.com/roboflow-gw7yv/website-screenshots/dataset/1] Uses From the roboflow website: Annotated screenshots are very useful in Robotic Process Automation. But they can be expensive to label. This dataset would cost over… See the full description on the dataset page: https://huggingface.co/datasets/Zexanima/website_screenshots_image_dataset.imageobject-detection1K<n<10K22 likes133 downloads3y agoHugging Face23naorm /website-screenshots-blip-largeimage1K<n<10K1 likes113 downloads3y agoHugging Face24G1nOnly /MoRE_Website_Mesh0 likes112 downloads8mo agoHugging Face25arcadia1991 /top-1M-websitetext100K<n<1M1 likes97 downloads3y agoHugging Face26Symato /website-datadocument0 likes93 downloads3y agoHugging Face27i2ebuddy /website_data0 likes91 downloads2y agoHugging Face28Shivlovejyotibd /websiteimagen<1K0 likes86 downloads3d agoHugging Face29raphael0202 /amenity-website-images-wds0 likes83 downloads7d agoHugging Face30intertwine-expel /expel-website Expel.com Website Pages textn<1K1 likes79 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.