CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mindweave /web-server-logs Web Server Access Logs (Synthetic) (Free Sample) This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables. Realistic HTTP access logs from a simulated SaaS company running an e-commerce API and marketing website. 50,000 requests across 3 servers over 12 months. Includes realistic patterns: weekday/weekend traffic variation, peak hours, seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and database outage) for anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/web-server-logs.tabulartabular-classification1K<n<10K0 likes1.4k downloads6mo agoHugging Face02APProjects /us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York, North Carolina and Pennsylvania — retired the web pages their older WARN Act layoff notices lived on. Their current pages start years later. This dataset is every notice in our file that came from one of those retired pages and is not on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.tabulartabular-classification1K<n<10K0 likes821 downloads1d agoHugging Face03BeIR /webis-touche2020-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/webis-touche2020-qrels.tabulartext-retrieval1K<n<10K0 likes399 downloads4y agoHugging Face04lucas-ventura /WebVid-CoVRarxiv.org/abs/2308.14746 tabular1M<n<10M5 likes133 downloads2y agoHugging Face05Pijush22049 /casia-webfaceimage100K<n<1M0 likes118 downloads5mo agoHugging Face06DeusHorizon /agent-web-index Agent Web Index — how much of the web can AI assistants actually read? 49,253 domains measured live. 26% of them cannot be read by at least one of ChatGPT, Claude, Perplexity or Gemini. Updated daily. Live index: https://shop.lumnika.com/ai-readiness/ Every row here is the result of real HTTP requests, not an estimate and not a re-publication of someone else's crawl: each domain's homepage is requested once as a browser and once as each of the published AI crawler user-agents… See the full description on the dataset page: https://huggingface.co/datasets/DeusHorizon/agent-web-index.tabular10K<n<100K0 likes118 downloads2h agoHugging Face07webxos /nexus-fpv NEXUS FPV Physics Dataset Sampler by webXOS Auto-generated small FPV drone flight-telemetry and reinforcement-learning for experience sample dataset captured directly in the browser by NEXUS FPV. Each row is one physics tick recorded while a drone flew through a waypoint course, either under manual control or the built-in PID auto-pilot. Play the game and make your own datasets: https://webxos.itch.io/nexus-fpv or download it from the /gym/ folder of this repo. Generator: NEXUS… See the full description on the dataset page: https://huggingface.co/datasets/webxos/nexus-fpv.tabularreinforcement-learningn<1K1 likes100 downloads2mo agoHugging Face08shengqin /web-attacks-longtabular10K<n<100K4 likes99 downloads3y agoHugging Face09shengqin /web-attackstabular10K<n<100K12 likes76 downloads3y agoHugging Face10WebSEM-ai /agent-discoverability-ado-score-romania Agent Discoverability (ADO Score) — Romania, September 2026 130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0. Canonical study (analysis, charts, interpretation): Romanian · English What this is On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.tabularn<1K0 likes69 downloads17d agoHugging Face11Plugiloinc /45_Million_Websitestabular1M<n<10M1 likes63 downloads1y agoHugging Face12linaaaaaaaaaaaaa /web-server-logs Web Server Access Logs (Synthetic) (Free Sample) This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables. Realistic HTTP access logs from a simulated SaaS company running an e-commerce API and marketing website. 50,000 requests across 3 servers over 12 months. Includes realistic patterns: weekday/weekend traffic variation, peak hours, seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and database outage) for anomaly… See the full description on the dataset page: https://huggingface.co/datasets/linaaaaaaaaaaaaa/web-server-logs.tabulartabular-classification1K<n<10K1 likes56 downloads3mo agoHugging Face13Webopen2026 /enterprise-financial-crime-ai-datasetTransactions → Risk Analysis → Alerts → Investigation → SAR Reports Dataset Statistics Total records: 310,396Dataset size: 339 MBAuto-converted parquet size: 65 MB Languages: English French Spanish Main fields: email_id thread_id timestamp language bank department country risk_level Enterprise Financial Crime AI Dataset The Enterprise Financial Crime AI Dataset is a high-fidelity dataset built from real-world operational patterns and enterprise data structures… See the full description on the dataset page: https://huggingface.co/datasets/Webopen2026/enterprise-financial-crime-ai-dataset.tabulartext-classification100K<n<1M0 likes50 downloads7mo agoHugging Face14OmriShtayer /Website_Traffic_and_Engagementtabulartable-question-answeringn<1K0 likes45 downloads1y agoHugging Face15webimmunization /COVID-19-vaccine-attitude-tweets Dataset Card for COVID-19-vaccine-attitude-tweets Dataset Summary The dataset consists of 2564 manually annotated tweets related to COVID-19 vaccines. The dataset can be used to discover the attitude expressed in the tweet towards the subject of COVID-19 vaccines. Tweets are in English. The dataset was curated in such a way as to maximize the likelihood of tweets with a strong emotional tone. We have assumed the existence of three classes: PRO (label 0): positive, the… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-vaccine-attitude-tweets.tabulartext-classification1K<n<10K2 likes38 downloads4y agoHugging Face16tresor2k /barometre-prix-creation-site-web-france-2026 Website Creation Pricing in France 2026 — Barometer 103 real, publicly available price points for website creation services in France (collected 2026-06-11), by service type and provider category. License CC-BY 4.0. Companion study (FR): https://lescreavores.fr/prix-creation-site-internet/ Maintainer: Les Créavores — https://lescreavores.fr DOI (Zenodo): https://doi.org/10.5281/zenodo.20690911 See METHODOLOGY.md for the collection method and README/columns dictionary below. tabularn<1K0 likes36 downloads1mo agoHugging Face17YangYang-Research /web-attack-detectionThe dataset contains 625,904 attack payload samples, with 294,771 labeled as 1 and 331,129 labeled as 0, including SQL injection, XSS, command injection, and other vulnerabilities. tabulartext-classification100K<n<1M1 likes33 downloads2y agoHugging Face18zimzum1984 /website-project-cost-benchmarks Website & App Project Cost Benchmarks 2026 Website & app project cost benchmarks - calibrated on 600+ project quotes and public rate benchmarks. CC-BY 4.0. Source and methodology: https://projectcostestimator.com This is the dataset behind Project Cost Estimator, an independent website cost estimator. The canonical machine-readable source is the live endpoint https://projectcostestimator.com/api/cost-data (no auth, CORS open). The files here are a published snapshot of that… See the full description on the dataset page: https://huggingface.co/datasets/zimzum1984/website-project-cost-benchmarks.tabularn<1K0 likes26 downloads1mo agoHugging Face19zeyu8800 /Weather_Underground_Webscrapetabular1M<n<10M0 likes25 downloads11mo agoHugging Face20DEplain /DEplain-web-doc DEplain-web-doc: A corpus for German Document Simplification DEplain-web-doc is a subcorpus of DEplain Stodden et al., 2023 for document simplification. The corpus consists of 396 (199/50/147) parallel documents crawled from the web in standard German and plain German (or easy-to-read German). All documents are either published under an open license or the copyright holders gave us the permission to share the data. If you are interested in a larger corpus, please check our paper… See the full description on the dataset page: https://huggingface.co/datasets/DEplain/DEplain-web-doc.tabularn<1K1 likes24 downloads3y agoHugging Face21WebSEM-ai /ai-search-visibility-romania-electronics-market AI Search Visibility — Romania's Electronics & IT Market (August 2026) 18 brand-free purchase questions × 5 AI engines = 87 answers. 86 of them name a major retailer. Position, not presence, decides the market. Raw data CC BY 4.0. Canonical study (analysis, charts, interpretation): Romanian · English What this is Eighteen real purchase questions were put to ChatGPT, Google Gemini, Perplexity, Google AI Mode and Google AI Overviews, in Romanian, from Romania, in… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-electronics-market.tabularn<1K0 likes24 downloads2mo agoHugging Face22abgunaydin /webgpu-compute-benchmarks WebGPU compute benchmarks across browsers and GPU vendors Every run submitted to gpubench.dev between 2026-03-30 and 2026-08-14, exported from the live table. WebGPU compute-shader throughput measured in real browsers on whatever hardware visitors happened to have. This exists because two preprints cite cross-vendor results from this table and the table is live and mutable, so those claims could not be checked by a reader. This snapshot is the checkable version.… See the full description on the dataset page: https://huggingface.co/datasets/abgunaydin/webgpu-compute-benchmarks.tabulartabular-regression1K<n<10K0 likes24 downloads1mo agoHugging Face23web3hungry /indonesia-affordable-housing Indonesia Affordable Housing Dataset A comprehensive dataset of affordable housing projects across Indonesia, containing detailed information about residential properties, specifications, locations, pricing, and developer data. Dataset Creator: web3hungry Dataset ID: web3hungry/indonesia-affordable-housing License: CC0 1.0 Dataset Overview This dataset provides extensive data on housing developments throughout Indonesia, covering both subsidized and commercial… See the full description on the dataset page: https://huggingface.co/datasets/web3hungry/indonesia-affordable-housing.tabulartabular-classification10K<n<100K0 likes23 downloads7mo agoHugging Face24joshuarg007 /state-of-cpa-firm-websites-2026 The State of CPA Firm Websites 2026: Structured-Data and AI-Legibility Dataset (n=556) This dataset accompanies the study "The State of CPA Firm Websites 2026" by Axion Deep Digital. It measures the structured-data legibility of United States accounting-firm websites: whether their sites make named experts, identity, and answer-shaped content machine-readable for search engines and AI answer engines. Sample: 556 unique accounting-firm domains sampled from OpenStreetMap… See the full description on the dataset page: https://huggingface.co/datasets/joshuarg007/state-of-cpa-firm-websites-2026.tabularn<1K0 likes23 downloads3mo agoHugging Face25joshuarg007 /ai-visibility-small-business-websites-2026 AI Crawler Visibility and JavaScript Rendering Gaps in Small Business Websites: Dual-Capture Dataset (n=368) This dataset accompanies the study by Axion Deep Digital measuring how much answer-critical content on small business websites is available to a JavaScript-rendering engine (Google) but absent from a non-rendering AI crawler's direct fetch (GPTBot, ClaudeBot, PerplexityBot). Method: Each site was captured twice in a single pass, once rendered in headless Chromium with… See the full description on the dataset page: https://huggingface.co/datasets/joshuarg007/ai-visibility-small-business-websites-2026.tabularn<1K0 likes23 downloads3mo agoHugging Face26shengqin /web-attacks-ab2tabular10K<n<100K6 likes22 downloads3y agoHugging Face27WebSEM-ai /ai-visibility-romania-luxury-jewelry AI Visibility — Romania's Luxury Jewelry (July 2026) We asked ten large language models the same question, word for word. We got 29 different brands across 50 available positions, and no brand appeared in all ten lists. Canonical study (analysis, charts, interpretation): Romanian · English The prompt Which luxury jewelry brands from Romania do you recommend for wedding bands and engagement rings? Give me a top 5, with a short argument for each and the sources you… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-visibility-romania-luxury-jewelry.tabularn<1K1 likes22 downloads2mo agoHugging Face28DEplain /DEplain-web-sent DEplain-web-sent: A corpus for German Sentence Simplification DEplain-web-sent is a subcorpus of DEplain Stodden et al., 2023 for evaluation of sentence simplification. The corpus consists of 1846 sentence pairs of 147 parallel documents crawled from the web in standard German and plain German (or easy-to-read German). All documents are either published under an open license, or the copyright holders gave us permission to share the data. Human annotators sentence-wise aligned the… See the full description on the dataset page: https://huggingface.co/datasets/DEplain/DEplain-web-sent.tabular1K<n<10K1 likes20 downloads3y agoHugging Face29Doubiiu /webvid10m_motiontabular1M<n<10M16 likes18 downloads2y agoHugging Face30Ohja-ai /rki-grippe-web-wochenbericht-2026-02-06 RKI GrippeWeb – Daten des Wochenberichts (re-hosted on Hugging Face) This Hugging Face dataset repository re-hosts a snapshot (2026-02-06) of the public weekly dataset “GrippeWeb – Daten des Wochenberichts” published by the Robert Koch-Institut (RKI) on GitHub under the Creative Commons Attribution 4.0 license: https://github.com/robert-koch-institut/GrippeWeb_Daten_des_Wochenberichts No endorsement: This is an independent re-hosting of RKI’s open data. Robert Koch-Institut (RKI)… See the full description on the dataset page: https://huggingface.co/datasets/Ohja-ai/rki-grippe-web-wochenbericht-2026-02-06.tabulartime-series-forecasting10K<n<100K0 likes18 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.