datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
osm-polygon-website-tag
OSM Polygon Website Dataset
OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files.
At a glance
Polygons
1,726,474
With extracted text
1,192,980
Words of text
407,685,655
Languages
397
Regional sources
386 / 386
Duplicate objects removed
104,927
Candidates rejected
868,905,743
Status
In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-24. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.HTML-CSS-Website# Dataset
This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant.
# Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com)
Coptic-websites-scrappingshopify-websites
Shopify Websites Dataset
Dataset Description
The Shopify Websites Dataset is a curated collection of 10,000 verified Shopify-powered e-commerce store URLs, providing researchers and analysts with a comprehensive resource for studying the Shopify e-commerce ecosystem.
This dataset offers a diverse snapshot of real-world Shopify stores across various industries and geographies, making it valuable for market research, web scraping projects, and e-commerce platform analysis.… See the full description on the dataset page: https://huggingface.co/datasets/snncn/shopify-websites.website_metadata_c4The dataset is in the form of a json lines file with 1,20,000 examples, where an example consists of text (extracted from C4 English dataset) and metadata fields (website description extracted from Wikipedia).
Example:
{
"text": "US10289222B2 - Handling of touch events in a browser environment - Google Patents\nHandling of touch events in a browser environment Download PDF\nUS10289222B2\nUS10289222B2 US13/857,848 US201313857848A US10289222B2 US 10289222 B2 US10289222 B2 US 10289222B2 US… See the full description on the dataset page: https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4.blog-website-datawebsite_screenshots_image_dataset
Website Screenshots Image Dataset
This dataset is obtainable here from roboflow..
Dataset Details
Dataset Description
Language(s) (NLP): [English]
License: [MIT]
Dataset Sources
Source: [https://universe.roboflow.com/roboflow-gw7yv/website-screenshots/dataset/1]
Uses
From the roboflow website:
Annotated screenshots are very useful in Robotic Process Automation. But they can be expensive to label. This dataset would cost over… See the full description on the dataset page: https://huggingface.co/datasets/Zexanima/website_screenshots_image_dataset.website-screenshots-blip-largetop-1M-websiteexpel-website
Expel.com Website Pages
website-html-2kThis dataset was a subset code tasks 33k filtered exclusively for high-quality websites.
The filtering includes:
Minimum of 100 lines
Includes CSS or Javascript content
The filtering was fairly light, as for stricter cleaning I only had 1,100 examples, which I considered to be low. Having 900 more examples is worth it in my opinion.
I have also included a python file which grabs the HTML content from random responses and saves those as files 1.html ==> 5.html to view them.
jusci-website
Dataset Card for "jusci-website"
This dataset was scrape jusci.net. JuSci is science website news and is part of Blognone Content Network. It use CC-BY 3.0 license.
45_Million_Websitescode-website-setgoethe-website
Dataset Card for "goethe-website"
This dataset collect the article from https://www.goethe.de/ins/th. (CC BY-SA only)
osm-polygon-website-tag-eunis
OSM Polygon Website Dataset
OpenStreetMap closed ways and polygon relations carrying a non-empty website OR contact:website tag, with full main-page text extracted using Trafilatura. Every statistic below is regenerated from the current upload-acknowledged Parquet artifacts.
Snapshot
Metric
Value
What it means
Snapshot status
In progress
Current published snapshot
Regional PBFs included
386 / 386
Published source shards / expected source PBFs… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag-eunis.website-technology-evidence-dataset
Website Technology Evidence Dataset
A balanced synthetic dataset for classifying observable website fingerprint evidence into a likely web technology.
Provenance
All records are synthetically generated from documented, recognizable public fingerprints. No claim is made that these records were collected from real websites.
Dataset
38 technologies
6,840 examples
180 examples per technology
train: 5,472
validation: 684
test: 684
balanced classes… See the full description on the dataset page: https://huggingface.co/datasets/newazhala/website-technology-evidence-dataset.inquiringmind-website
Dataset Card for "inquiringmind-website"
This dataset collect web page from Inquiring Mind (ครูไทย หัวใจสืบเสาะ). Ihe license is cc-by-sa-3.0.
Website_Traffic_and_Engagementwebsite_metadata_c4_toyA smaller version (100 samples) of https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4
merged-coder-website-reasoningsearch-results-websites
Dataset Card for General
This instruction tuning dataset has been created with SPInO and has 2720 rows of textual data related to General.
from datasets import load_dataset
dataset = load_dataset("1rsh/search-results-websites")
print(dataset)
website_categorieswebsite-industry-13m
Dataset Card for 13M+ Website Domains with Industry Labels
This dataset contains over 12 million website domain URLs mapped to their associated industry categories.It is useful for domain classification, industry prediction, NLP preprocessing, clustering, and machine learning research.
This dataset card has been generated based on the Hugging Face dataset card template.
Dataset Details
The dataset has two columns:
website: Domain URL (cleaned, no DNS tags)… See the full description on the dataset page: https://huggingface.co/datasets/bedead/website-industry-13m.crawl-malaysian-websitejina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564
jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 Dataset
Dataset Description
jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 model.
How to Use
To use this dataset for model training or evaluation… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564.website-screenshots-git-largeUSCIS-knowledge-base-full-website
A comprehensive dataset of 99,489 content chunks from 4,666 pages on the USCIS website, with pre-computed OpenAI text-embedding-ada-002 embeddings (1536 dimensions).
Built for RAG (Retrieval-Augmented Generation), semantic search, and GraphRAG applications focused on U.S. immigration law and policy.
🔗 GitHub: github.com/0xrphl/USCIS-knowledge-base-full-website🔥 Scraped with: Firecrawl — The open-source web scraping API for AI🍎 Visualized with: Embedding Atlas — Interactive embedding… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/USCIS-knowledge-base-full-website.ml6-website-rag
