datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
web-server-logs
Web Server Access Logs (Synthetic) (Free Sample)
This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables.
Realistic HTTP access logs from a simulated SaaS company running an
e-commerce API and marketing website. 50,000 requests across 3 servers
over 12 months.
Includes realistic patterns: weekday/weekend traffic variation, peak hours,
seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and
database outage) for anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/web-server-logs.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.webis-touche2020-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/webis-touche2020-qrels.WebVid-CoVRarxiv.org/abs/2308.14746
casia-webfaceagent-web-index
Agent Web Index — how much of the web can AI assistants actually read?
49,253 domains measured live. 26% of them cannot be read by at least one of
ChatGPT, Claude, Perplexity or Gemini. Updated daily. Live index: https://shop.lumnika.com/ai-readiness/
Every row here is the result of real HTTP requests, not an estimate and not a re-publication of
someone else's crawl: each domain's homepage is requested once as a browser and once as each of the
published AI crawler user-agents… See the full description on the dataset page: https://huggingface.co/datasets/DeusHorizon/agent-web-index.nexus-fpv
NEXUS FPV Physics Dataset Sampler by webXOS
Auto-generated small FPV drone flight-telemetry and reinforcement-learning for experience sample dataset captured directly in the browser
by NEXUS FPV. Each row is one physics tick recorded while a drone flew through a waypoint course, either under manual control or the built-in PID auto-pilot.
Play the game and make your own datasets: https://webxos.itch.io/nexus-fpv or download it from the /gym/ folder of this repo.
Generator: NEXUS… See the full description on the dataset page: https://huggingface.co/datasets/webxos/nexus-fpv.web-attacks-longweb-attacksagent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.45_Million_Websitesweb-server-logs
Web Server Access Logs (Synthetic) (Free Sample)
This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables.
Realistic HTTP access logs from a simulated SaaS company running an
e-commerce API and marketing website. 50,000 requests across 3 servers
over 12 months.
Includes realistic patterns: weekday/weekend traffic variation, peak hours,
seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and
database outage) for anomaly… See the full description on the dataset page: https://huggingface.co/datasets/linaaaaaaaaaaaaa/web-server-logs.enterprise-financial-crime-ai-datasetTransactions → Risk Analysis → Alerts → Investigation → SAR Reports
Dataset Statistics
Total records: 310,396Dataset size: 339 MBAuto-converted parquet size: 65 MB
Languages:
English
French
Spanish
Main fields:
email_id
thread_id
timestamp
language
bank
department
country
risk_level
Enterprise Financial Crime AI Dataset
The Enterprise Financial Crime AI Dataset is a high-fidelity dataset built from real-world operational patterns and enterprise data structures… See the full description on the dataset page: https://huggingface.co/datasets/Webopen2026/enterprise-financial-crime-ai-dataset.Website_Traffic_and_EngagementCOVID-19-vaccine-attitude-tweets
Dataset Card for COVID-19-vaccine-attitude-tweets
Dataset Summary
The dataset consists of 2564 manually annotated tweets related to COVID-19 vaccines. The dataset can be used to discover the attitude expressed in the tweet towards the subject of COVID-19 vaccines. Tweets are in English. The dataset was curated in such a way as to maximize the likelihood of tweets with a strong emotional tone. We have assumed the existence of three classes:
PRO (label 0): positive, the… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-vaccine-attitude-tweets.barometre-prix-creation-site-web-france-2026
Website Creation Pricing in France 2026 — Barometer
103 real, publicly available price points for website creation services in France
(collected 2026-06-11), by service type and provider category. License CC-BY 4.0.
Companion study (FR): https://lescreavores.fr/prix-creation-site-internet/
Maintainer: Les Créavores — https://lescreavores.fr
DOI (Zenodo): https://doi.org/10.5281/zenodo.20690911
See METHODOLOGY.md for the collection method and README/columns dictionary below.
web-attack-detectionThe dataset contains 625,904 attack payload samples, with 294,771 labeled as 1 and 331,129 labeled as 0, including SQL injection, XSS, command injection, and other vulnerabilities.
website-project-cost-benchmarks
Website & App Project Cost Benchmarks 2026
Website & app project cost benchmarks - calibrated on 600+ project quotes and public rate benchmarks. CC-BY 4.0. Source and methodology: https://projectcostestimator.com
This is the dataset behind Project Cost Estimator, an independent website cost estimator. The canonical machine-readable source is the live endpoint https://projectcostestimator.com/api/cost-data (no auth, CORS open). The files here are a published snapshot of that… See the full description on the dataset page: https://huggingface.co/datasets/zimzum1984/website-project-cost-benchmarks.Weather_Underground_WebscrapeDEplain-web-doc
DEplain-web-doc: A corpus for German Document Simplification
DEplain-web-doc is a subcorpus of DEplain Stodden et al., 2023 for document simplification.
The corpus consists of 396 (199/50/147) parallel documents crawled from the web in standard German and plain German (or easy-to-read German). All documents are either published under an open license or the copyright holders gave us the permission to share the data.
If you are interested in a larger corpus, please check our paper… See the full description on the dataset page: https://huggingface.co/datasets/DEplain/DEplain-web-doc.ai-search-visibility-romania-electronics-market
AI Search Visibility — Romania's Electronics & IT Market (August 2026)
18 brand-free purchase questions × 5 AI engines = 87 answers. 86 of them name a major retailer. Position, not presence, decides the market. Raw data CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
Eighteen real purchase questions were put to ChatGPT, Google Gemini, Perplexity, Google AI Mode and Google AI Overviews, in Romanian, from Romania, in… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-electronics-market.webgpu-compute-benchmarks
WebGPU compute benchmarks across browsers and GPU vendors
Every run submitted to gpubench.dev between 2026-03-30
and 2026-08-14, exported from the live table. WebGPU compute-shader throughput
measured in real browsers on whatever hardware visitors happened to have.
This exists because two preprints cite cross-vendor results from this table and
the table is live and mutable, so those claims could not be checked by a reader.
This snapshot is the checkable version.… See the full description on the dataset page: https://huggingface.co/datasets/abgunaydin/webgpu-compute-benchmarks.indonesia-affordable-housing
Indonesia Affordable Housing Dataset
A comprehensive dataset of affordable housing projects across Indonesia, containing detailed information about residential properties, specifications, locations, pricing, and developer data.
Dataset Creator: web3hungry
Dataset ID: web3hungry/indonesia-affordable-housing
License: CC0 1.0
Dataset Overview
This dataset provides extensive data on housing developments throughout Indonesia, covering both subsidized and commercial… See the full description on the dataset page: https://huggingface.co/datasets/web3hungry/indonesia-affordable-housing.state-of-cpa-firm-websites-2026
The State of CPA Firm Websites 2026: Structured-Data and AI-Legibility Dataset (n=556)
This dataset accompanies the study "The State of CPA Firm Websites 2026" by Axion Deep Digital. It measures the structured-data legibility of United States accounting-firm websites: whether their sites make named experts, identity, and answer-shaped content machine-readable for search engines and AI answer engines.
Sample: 556 unique accounting-firm domains sampled from OpenStreetMap… See the full description on the dataset page: https://huggingface.co/datasets/joshuarg007/state-of-cpa-firm-websites-2026.ai-visibility-small-business-websites-2026
AI Crawler Visibility and JavaScript Rendering Gaps in Small Business Websites: Dual-Capture Dataset (n=368)
This dataset accompanies the study by Axion Deep Digital measuring how much answer-critical content on small business websites is available to a JavaScript-rendering engine (Google) but absent from a non-rendering AI crawler's direct fetch (GPTBot, ClaudeBot, PerplexityBot).
Method: Each site was captured twice in a single pass, once rendered in headless Chromium with… See the full description on the dataset page: https://huggingface.co/datasets/joshuarg007/ai-visibility-small-business-websites-2026.web-attacks-ab2ai-visibility-romania-luxury-jewelry
AI Visibility — Romania's Luxury Jewelry (July 2026)
We asked ten large language models the same question, word for word. We got
29 different brands across 50 available positions, and no brand appeared in
all ten lists.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
The prompt
Which luxury jewelry brands from Romania do you recommend for wedding bands
and engagement rings? Give me a top 5, with a short argument for each and the
sources you… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-visibility-romania-luxury-jewelry.DEplain-web-sent
DEplain-web-sent: A corpus for German Sentence Simplification
DEplain-web-sent is a subcorpus of DEplain Stodden et al., 2023 for evaluation of sentence simplification.
The corpus consists of 1846 sentence pairs of 147 parallel documents crawled from the web in standard German and plain German (or easy-to-read German). All documents are either published under an open license, or the copyright holders gave us permission to share the data.
Human annotators sentence-wise aligned the… See the full description on the dataset page: https://huggingface.co/datasets/DEplain/DEplain-web-sent.webvid10m_motionrki-grippe-web-wochenbericht-2026-02-06
RKI GrippeWeb – Daten des Wochenberichts (re-hosted on Hugging Face)
This Hugging Face dataset repository re-hosts a snapshot (2026-02-06) of the public weekly dataset “GrippeWeb – Daten des Wochenberichts” published by the Robert Koch-Institut (RKI) on GitHub under the Creative Commons Attribution 4.0 license:
https://github.com/robert-koch-institut/GrippeWeb_Daten_des_Wochenberichts
No endorsement: This is an independent re-hosting of RKI’s open data. Robert Koch-Institut (RKI)… See the full description on the dataset page: https://huggingface.co/datasets/Ohja-ai/rki-grippe-web-wochenbericht-2026-02-06.
