datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Web_site_legacycrawl-my-websiteosm-polygon-website-tag
OSM Polygon Website Dataset
OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files.
At a glance
Polygons
1,726,474
With extracted text
1,192,980
Words of text
407,685,655
Languages
397
Regional sources
386 / 386
Duplicate objects removed
104,927
Candidates rejected
868,905,743
Status
In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.causvid_website1k_Website_Screenshots_and_Metadata
Dataset Card for 1000 Website Screenshots with Metadata
Dataset Summary
Silatus is sharing, for free, a segment of a dataset that we are using to train a generative AI model for text-to-mockup conversions. This dataset was collected in December 2022 and early January 2023, so it contains very recent data from 1,000 of the world's most popular websites. You can get our larger 10,000 website dataset for free at: https://silatus.com/datasets
This dataset includes:
High-res… See the full description on the dataset page: https://huggingface.co/datasets/silatus/1k_Website_Screenshots_and_Metadata.mind2web_multimodal_test_website
Dataset Card for Multimodal Mind2Web "Cross-Website" Test Split
Note: This dataset is the test split of the Cross-Website dataset introduced in the paper.
This is a FiftyOne dataset with 1019 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_website.Translations_Hungarian_public_websites
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/18982
Description
A webcrawl of 14 different websites covering parallel corpora of Hungarian with Polish, Czech, Swedish, Finnish, French, German, Italian, English and Slovenian
Citation
Translations of Hungarian from public websites (2022). Version 1.0. [Dataset (Text corpus)]. Source: European Language Grid. https://live.european-language-grid.eu/catalogue/corpus/18982
phishing-website-screenshots
Phishing Website Screenshots
A dataset of 8,370 full-page website screenshots labelled as legitimate or phishing, intended for training and evaluating visual phishing-detection models.
Contents
Label
label
Images
legitimate
0
7,924
phishing
1
446
Total
8,370
Screenshots were captured at a desktop viewport (1920×1080) as PNG images.
Structure
legitimate/<brand>/<page>.png
phishing/<source>/<page>.png
metadata.csv
metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/shresthsamyak/phishing-website-screenshots.website-media
OpenRAL — website media
Video clips shown in the "See it run" section of openral.com
(benchmarks, simulation and on-hardware deployment runs).
Each clip lives under <category>/<benchmark>_<rskill>_<success|fail>/ with three
web-optimised assets:
poster.jpg — first-frame thumbnail
preview.mp4 — square 640px, muted (the autoscroll strip)
full.mp4 — native aspect, ≤1080p, with audio (the expand modal)
Generated and published by scripts/build-media.mjs in the
website repo.… See the full description on the dataset page: https://huggingface.co/datasets/OpenRAL/website-media.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-24. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.malicious-website-features-2.4MImportant Notice:
A subset of the URL dataset is from Kaggle, and the Kaggle datasets contained 10%-15% mislabelled data. See this dicussion I opened for some false positives. I have contacted Kaggle regarding their erroneous "Usability" score calculation for these unreliable datasets.
The feature extraction methods shown here are not robust at all in 2023, and there're even silly mistakes in 3 functions: not_indexed_by_google, domain_registration_length, and age_of_domain.
The features… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/malicious-website-features-2.4M.HTML-CSS-Website# Dataset
This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant.
# Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com)
website-gallery
The Bald Dude Co. Website Gallery
Web-ready portfolio media for The Bald Dude Co. The album structure and public file URLs are published in archive.json for use by the official website.
Copyright The Bald Dude Co. All rights reserved. These files are not licensed for redistribution, resale, dataset compilation, or model training.
websiteCoptic-websites-scrappingshopify-websites
Shopify Websites Dataset
Dataset Description
The Shopify Websites Dataset is a curated collection of 10,000 verified Shopify-powered e-commerce store URLs, providing researchers and analysts with a comprehensive resource for studying the Shopify e-commerce ecosystem.
This dataset offers a diverse snapshot of real-world Shopify stores across various industries and geographies, making it valuable for market research, web scraping projects, and e-commerce platform analysis.… See the full description on the dataset page: https://huggingface.co/datasets/snncn/shopify-websites.figma-Playground-Inference-for-PRO-s-Websitewebsite_assetswebsite_metadata_c4The dataset is in the form of a json lines file with 1,20,000 examples, where an example consists of text (extracted from C4 English dataset) and metadata fields (website description extracted from Wikipedia).
Example:
{
"text": "US10289222B2 - Handling of touch events in a browser environment - Google Patents\nHandling of touch events in a browser environment Download PDF\nUS10289222B2\nUS10289222B2 US13/857,848 US201313857848A US10289222B2 US 10289222 B2 US10289222 B2 US 10289222B2 US… See the full description on the dataset page: https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4.blog-website-dataWebsiteFingerprinting-Rawwebsite_screenshots_image_dataset
Website Screenshots Image Dataset
This dataset is obtainable here from roboflow..
Dataset Details
Dataset Description
Language(s) (NLP): [English]
License: [MIT]
Dataset Sources
Source: [https://universe.roboflow.com/roboflow-gw7yv/website-screenshots/dataset/1]
Uses
From the roboflow website:
Annotated screenshots are very useful in Robotic Process Automation. But they can be expensive to label. This dataset would cost over… See the full description on the dataset page: https://huggingface.co/datasets/Zexanima/website_screenshots_image_dataset.website-screenshots-blip-largeMoRE_Website_Meshtop-1M-websitewebsite-datawebsite_datawebsiteamenity-website-images-wdsexpel-website
Expel.com Website Pages
