datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
osm-polygon-website-tag
OSM Polygon Website Dataset
OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files.
At a glance
Polygons
1,726,474
With extracted text
1,192,980
Words of text
407,685,655
Languages
397
Regional sources
386 / 386
Duplicate objects removed
104,927
Candidates rejected
868,905,743
Status
In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.45_Million_Websitesosm-polygon-website-tag-eunis
OSM Polygon Website Dataset
OpenStreetMap closed ways and polygon relations carrying a non-empty website OR contact:website tag, with full main-page text extracted using Trafilatura. Every statistic below is regenerated from the current upload-acknowledged Parquet artifacts.
Snapshot
Metric
Value
What it means
Snapshot status
In progress
Current published snapshot
Regional PBFs included
386 / 386
Published source shards / expected source PBFs… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag-eunis.Website_Traffic_and_Engagementwebsite_categoriesUSCIS-knowledge-base-full-website
A comprehensive dataset of 99,489 content chunks from 4,666 pages on the USCIS website, with pre-computed OpenAI text-embedding-ada-002 embeddings (1536 dimensions).
Built for RAG (Retrieval-Augmented Generation), semantic search, and GraphRAG applications focused on U.S. immigration law and policy.
🔗 GitHub: github.com/0xrphl/USCIS-knowledge-base-full-website🔥 Scraped with: Firecrawl — The open-source web scraping API for AI🍎 Visualized with: Embedding Atlas — Interactive embedding… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/USCIS-knowledge-base-full-website.website-project-cost-benchmarks
Website & App Project Cost Benchmarks 2026
Website & app project cost benchmarks - calibrated on 600+ project quotes and public rate benchmarks. CC-BY 4.0. Source and methodology: https://projectcostestimator.com
This is the dataset behind Project Cost Estimator, an independent website cost estimator. The canonical machine-readable source is the live endpoint https://projectcostestimator.com/api/cost-data (no auth, CORS open). The files here are a published snapshot of that… See the full description on the dataset page: https://huggingface.co/datasets/zimzum1984/website-project-cost-benchmarks.state-of-cpa-firm-websites-2026
The State of CPA Firm Websites 2026: Structured-Data and AI-Legibility Dataset (n=556)
This dataset accompanies the study "The State of CPA Firm Websites 2026" by Axion Deep Digital. It measures the structured-data legibility of United States accounting-firm websites: whether their sites make named experts, identity, and answer-shaped content machine-readable for search engines and AI answer engines.
Sample: 556 unique accounting-firm domains sampled from OpenStreetMap… See the full description on the dataset page: https://huggingface.co/datasets/joshuarg007/state-of-cpa-firm-websites-2026.ai-visibility-small-business-websites-2026
AI Crawler Visibility and JavaScript Rendering Gaps in Small Business Websites: Dual-Capture Dataset (n=368)
This dataset accompanies the study by Axion Deep Digital measuring how much answer-critical content on small business websites is available to a JavaScript-rendering engine (Google) but absent from a non-rendering AI crawler's direct fetch (GPTBot, ClaudeBot, PerplexityBot).
Method: Each site was captured twice in a single pass, once rendered in headless Chromium with… See the full description on the dataset page: https://huggingface.co/datasets/joshuarg007/ai-visibility-small-business-websites-2026.fw-darija-websitesstate-of-small-business-websites-2026
The State of Small Business Websites 2026: Core Web Vitals and Technical SEO Dataset (n=292)
This dataset accompanies the study by Axion Deep Digital measuring Core Web Vitals and technical SEO health across small business websites, audited with the DeepAudit AI engine (real-browser Lighthouse plus rule-based checks) rather than raw-HTML parsing.
Sample: 292 small business domains. Core Web Vitals and Lighthouse statistics are computed over the 191 sites that returned valid… See the full description on the dataset page: https://huggingface.co/datasets/joshuarg007/state-of-small-business-websites-2026.websitepopularitygc_websitesASU_website
