CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01firecrawl /scrape-content-dataset-v1 Scrape Content Dataset v1 A human-curated benchmark dataset for evaluating web scraping engines on content quality. Overview This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time. Dataset Structure CSV format with columns: id: Sequential identifier url:… See the full description on the dataset page: https://huggingface.co/datasets/firecrawl/scrape-content-dataset-v1.text1K<n<10K0 likes82 downloads11mo agoHugging Face02IrishHotelReviews /scraped-hotel-reviewstabular10K<n<100K0 likes71 downloads3mo agoHugging Face03bensblueprints /atomic-scraper-leads Atomic Scraper Leads Public business listings scraped from Google Maps for leads.benjaminboyce.com. This dataset is an export of the leads table from the Atomic / gmaps-scraper-suite pipeline. Each row is a local business listing with contact and location fields plus scrape metadata. Freshness Use scraped_at (UTC timestamps) as the freshness signal. This snapshot was exported on 2026-08-25. Oldest scraped_at: 2026-08-09 Newest scraped_at: 2026-08-25 Listings… See the full description on the dataset page: https://huggingface.co/datasets/bensblueprints/atomic-scraper-leads.tabulartabular-to-text10K<n<100K0 likes58 downloads28d agoHugging Face04Hasnatkhan010 /zameen-com-scrape-datatabularn<1K0 likes54 downloads26d agoHugging Face05falan42 /healifyLLM-QA-scraped-datasettext1K<n<10K1 likes49 downloads2y agoHugging Face06crackedvibe /agent-scraper-price-index-2026-09 Agent scraper price index, September 2026 The rows behind the report Agent scraper price index, September 2026 on cracked.ai: what 1,000 results cost for 18 of the most requested scraping and search capabilities, tool by tool, through three routes. Method For each capability, every live candidate tool on Cracked is listed with three prices for 1,000 results: the price billed on Cracked's own provider account (cost_1000_cracked_usd, provider price plus the $0.001… See the full description on the dataset page: https://huggingface.co/datasets/crackedvibe/agent-scraper-price-index-2026-09.tabulartabular-regressionn<1K0 likes45 downloads19d agoHugging Face07zhoubinghong /scrape-content-dataset-v1 Scrape Content Dataset v1 A human-curated benchmark dataset for evaluating web scraping engines on content quality. Overview This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time. Dataset Structure CSV format with columns: id: Sequential… See the full description on the dataset page: https://huggingface.co/datasets/zhoubinghong/scrape-content-dataset-v1.text1K<n<10K0 likes43 downloads6d agoHugging Face08metaltiger775 /Scraped_Dataset_Articlestextn<1K0 likes37 downloads3y agoHugging Face09dev-X-dev /quotes-scrapedtextn<1K0 likes34 downloads1y agoHugging Face10Ebullioscopic /Raw-Web-Scraped-to-JSONtextn<1K0 likes26 downloads2y agoHugging Face11timashan /amazon-scrape-4-llm Amazon Scrape 4 llm Purpose? Feed an LLM raw html to identify products from an ecommerce platform.These datasets contain the extracted innerTexts of all HTML nodes from different ecommerce product pages.The cleaning process significantly reduces the token size from ex: 450k -> 6k Quickstart from datasets import load_dataset data_train = load_dataset("timashan/amazon-scrape-4-llm", "phones") data_test = load_dataset("timashan/amazon-scrape-4-llm", "laptops")… See the full description on the dataset page: https://huggingface.co/datasets/timashan/amazon-scrape-4-llm.texttext-classificationn<1K0 likes26 downloads2y agoHugging Face12Arsalan8 /medquad_scraped_contexttextn<1K0 likes24 downloads2y agoHugging Face13harsh-lodha /lc-scrapedtext1K<n<10K0 likes22 downloads3y agoHugging Face14d-e-c-d /AI_Articles_Scraped_from_arXiv-Semantic_Scholar 📘 AI Articles Scraped from arXiv & Semantic Scholar 🧩 Description This dataset contains information on articles related to major AI conferences such as AAAI, NeurIPS, IJCAI, ICML, ICLR, collected through scraping from ArXiv and Semantic Scholar.It is intended to be used as a training dataset for various model training tasks and other desired uses. 📂 File Structure File Description AI_Titles_v2025.csv Main dataset README.md This file… See the full description on the dataset page: https://huggingface.co/datasets/d-e-c-d/AI_Articles_Scraped_from_arXiv-Semantic_Scholar.tabular100K<n<1M1 likes22 downloads1y agoHugging Face15Arconte /league_of_legends_wiki_scrapeThis dataset is a scrape from the League of Legends wiki, which contains the most up-to-date version with 166 champions. The data consists of: champion name, champion icon URL, champion wiki URL, stats, biography, passive ability, ability 1, ability 2, ability 3, ability 4, and curiosities. textn<1K2 likes19 downloads3y agoHugging Face16Aswinabd17 /scraped_cleaned_automobile.csvtextn<1K0 likes16 downloads2y agoHugging Face17grunemax /instagram-scrapertextn<1K0 likes10 downloads2y agoHugging Face18Pamzyy /Scrape_finaltext10K<n<100K0 likes9 downloads2y agoHugging Face19wwgd /scrapetabular10K<n<100K0 likes9 downloads2y agoHugging Face20TiaDay /books-to-scrape-page1 Books to Scrape – Page 1 Dataset Summary Book records scraped from the first page of the Books to Scrape demo site.I created this dataset for a class assignment to practise web scraping, pandas, and publishing a dataset to the Hugging Face Hub. Data Collection Source: https://books.toscrape.com/ (public test site for scraping practice) Method: requests.get("https://books.toscrape.com/catalogue/page-1.html") Parsed with BeautifulSoup, selecting each <article… See the full description on the dataset page: https://huggingface.co/datasets/TiaDay/books-to-scrape-page1.tabularn<1K0 likes3 downloads11mo agoHugging Face21PhuongMai0410 /Small_Scrape_datatextn<1K0 likes1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.