CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NoeFlandre /osm-polygon-website-tag OSM Polygon Website Dataset OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files. At a glance Polygons 1,726,474 With extracted text 1,192,980 Words of text 407,685,655 Languages 397 Regional sources 386 / 386 Duplicate objects removed 104,927 Candidates rejected 868,905,743 Status In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.tabular1M<n<10M2 likes2.4k downloads2d agoHugging Face02APProjects /us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered Rebuilt 2026-09-24. Five state labor agencies — Connecticut, Michigan, New York, North Carolina and Pennsylvania — retired the web pages their older WARN Act layoff notices lived on. Their current pages start years later. This dataset is every notice in our file that came from one of those retired pages and is not on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.tabulartabular-classification1K<n<10K0 likes751 downloads42m agoHugging Face03MAsad789565 /HTML-CSS-Website# Dataset This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant. # Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com) textn<1K4 likes512 downloads3y agoHugging Face04Philopater-Luka /Coptic-websites-scrappingtext10K<n<100K1 likes305 downloads20d agoHugging Face05snncn /shopify-websites Shopify Websites Dataset Dataset Description The Shopify Websites Dataset is a curated collection of 10,000 verified Shopify-powered e-commerce store URLs, providing researchers and analysts with a comprehensive resource for studying the Shopify e-commerce ecosystem. This dataset offers a diverse snapshot of real-world Shopify stores across various industries and geographies, making it valuable for market research, web scraping projects, and e-commerce platform analysis.… See the full description on the dataset page: https://huggingface.co/datasets/snncn/shopify-websites.text10K<n<100K1 likes215 downloads10mo agoHugging Face06bs-modeling-metadata /website_metadata_c4The dataset is in the form of a json lines file with 1,20,000 examples, where an example consists of text (extracted from C4 English dataset) and metadata fields (website description extracted from Wikipedia). Example: { "text": "US10289222B2 - Handling of touch events in a browser environment - Google Patents\nHandling of touch events in a browser environment Download PDF\nUS10289222B2\nUS10289222B2 US13/857,848 US201313857848A US10289222B2 US 10289222 B2 US10289222 B2 US 10289222B2 US… See the full description on the dataset page: https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4.text10K<n<100K4 likes166 downloads5y agoHugging Face07Sadeeshkumargopalan /blog-website-dataimagen<1K0 likes151 downloads10h agoHugging Face08Zexanima /website_screenshots_image_dataset Website Screenshots Image Dataset This dataset is obtainable here from roboflow.. Dataset Details Dataset Description Language(s) (NLP): [English] License: [MIT] Dataset Sources Source: [https://universe.roboflow.com/roboflow-gw7yv/website-screenshots/dataset/1] Uses From the roboflow website: Annotated screenshots are very useful in Robotic Process Automation. But they can be expensive to label. This dataset would cost over… See the full description on the dataset page: https://huggingface.co/datasets/Zexanima/website_screenshots_image_dataset.imageobject-detection1K<n<10K22 likes133 downloads3y agoHugging Face09naorm /website-screenshots-blip-largeimage1K<n<10K1 likes113 downloads3y agoHugging Face10arcadia1991 /top-1M-websitetext100K<n<1M1 likes97 downloads3y agoHugging Face11intertwine-expel /expel-website Expel.com Website Pages textn<1K1 likes79 downloads3y agoHugging Face12Sweaterdog /website-html-2kThis dataset was a subset code tasks 33k filtered exclusively for high-quality websites. The filtering includes: Minimum of 100 lines Includes CSS or Javascript content The filtering was fairly light, as for stricter cleaning I only had 1,100 examples, which I considered to be low. Having 900 more examples is worth it in my opinion. I have also included a python file which grabs the HTML content from random responses and saves those as files 1.html ==> 5.html to view them. text1K<n<10K0 likes74 downloads8mo agoHugging Face13pythainlp /jusci-website Dataset Card for "jusci-website" This dataset was scrape jusci.net. JuSci is science website news and is part of Blognone Content Network. It use CC-BY 3.0 license. texttext-generation1K<n<10K4 likes62 downloads2y agoHugging Face14Plugiloinc /45_Million_Websitestabular1M<n<10M1 likes62 downloads1y agoHugging Face15Bhargavtz /code-website-settext100K<n<1M0 likes60 downloads1y agoHugging Face16pythainlp /goethe-website Dataset Card for "goethe-website" This dataset collect the article from https://www.goethe.de/ins/th. (CC BY-SA only) texttext-generationn<1K0 likes55 downloads3y agoHugging Face17NoeFlandre /osm-polygon-website-tag-eunis OSM Polygon Website Dataset OpenStreetMap closed ways and polygon relations carrying a non-empty website OR contact:website tag, with full main-page text extracted using Trafilatura. Every statistic below is regenerated from the current upload-acknowledged Parquet artifacts. Snapshot Metric Value What it means Snapshot status In progress Current published snapshot Regional PBFs included 386 / 386 Published source shards / expected source PBFs… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag-eunis.tabular1M<n<10M0 likes53 downloads5d agoHugging Face18newazhala /website-technology-evidence-dataset Website Technology Evidence Dataset A balanced synthetic dataset for classifying observable website fingerprint evidence into a likely web technology. Provenance All records are synthetically generated from documented, recognizable public fingerprints. No claim is made that these records were collected from real websites. Dataset 38 technologies 6,840 examples 180 examples per technology train: 5,472 validation: 684 test: 684 balanced classes… See the full description on the dataset page: https://huggingface.co/datasets/newazhala/website-technology-evidence-dataset.texttext-classification1K<n<10K0 likes51 downloads25d agoHugging Face19pythainlp /inquiringmind-website Dataset Card for "inquiringmind-website" This dataset collect web page from Inquiring Mind (ครูไทย หัวใจสืบเสาะ). Ihe license is cc-by-sa-3.0. texttext-generationn<1K1 likes49 downloads2y agoHugging Face20OmriShtayer /Website_Traffic_and_Engagementtabulartable-question-answeringn<1K0 likes44 downloads1y agoHugging Face21shanya /website_metadata_c4_toyA smaller version (100 samples) of https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4 textn<1K1 likes41 downloads5y agoHugging Face22usernamebetter /merged-coder-website-reasoningtext10K<n<100K0 likes38 downloads2mo agoHugging Face231rsh /search-results-websites Dataset Card for General This instruction tuning dataset has been created with SPInO and has 2720 rows of textual data related to General. from datasets import load_dataset dataset = load_dataset("1rsh/search-results-websites") print(dataset) text1K<n<10K0 likes37 downloads2y agoHugging Face24massimilianowosz /website_categoriestabular10K<n<100K3 likes36 downloads3y agoHugging Face25bedead /website-industry-13m Dataset Card for 13M+ Website Domains with Industry Labels This dataset contains over 12 million website domain URLs mapped to their associated industry categories.It is useful for domain classification, industry prediction, NLP preprocessing, clustering, and machine learning research. This dataset card has been generated based on the Hugging Face dataset card template. Dataset Details The dataset has two columns: website: Domain URL (cleaned, no DNS tags)… See the full description on the dataset page: https://huggingface.co/datasets/bedead/website-industry-13m.texttext-classification10M<n<100M2 likes34 downloads1y agoHugging Face26aisyahhrazak /crawl-malaysian-websitetext10K<n<100K0 likes33 downloads3y agoHugging Face27florianhoenicke /jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 Dataset Dataset Description jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 model. How to Use To use this dataset for model training or evaluation… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564.textn<1K0 likes31 downloads2y agoHugging Face28naorm /website-screenshots-git-largeimage1K<n<10K5 likes29 downloads3y agoHugging Face290xrphl /USCIS-knowledge-base-full-website A comprehensive dataset of 99,489 content chunks from 4,666 pages on the USCIS website, with pre-computed OpenAI text-embedding-ada-002 embeddings (1536 dimensions). Built for RAG (Retrieval-Augmented Generation), semantic search, and GraphRAG applications focused on U.S. immigration law and policy. 🔗 GitHub: github.com/0xrphl/USCIS-knowledge-base-full-website🔥 Scraped with: Firecrawl — The open-source web scraping API for AI🍎 Visualized with: Embedding Atlas — Interactive embedding… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/USCIS-knowledge-base-full-website.tabulartext-retrieval100K<n<1M0 likes29 downloads3mo agoHugging Face30nielsr /ml6-website-ragtext1K<n<10K0 likes27 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.