datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scrape-content-dataset-v1
Scrape Content Dataset v1
A human-curated benchmark dataset for evaluating web scraping engines on content quality.
Overview
This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time.
Dataset Structure
CSV format with columns:
id: Sequential identifier
url:… See the full description on the dataset page: https://huggingface.co/datasets/firecrawl/scrape-content-dataset-v1.scraped-hotel-reviewsatomic-scraper-leads
Atomic Scraper Leads
Public business listings scraped from Google Maps for leads.benjaminboyce.com.
This dataset is an export of the leads table from the Atomic / gmaps-scraper-suite pipeline. Each row is a local business listing with contact and location fields plus scrape metadata.
Freshness
Use scraped_at (UTC timestamps) as the freshness signal. This snapshot was exported on 2026-08-25.
Oldest scraped_at: 2026-08-09
Newest scraped_at: 2026-08-25
Listings… See the full description on the dataset page: https://huggingface.co/datasets/bensblueprints/atomic-scraper-leads.zameen-com-scrape-datahealifyLLM-QA-scraped-datasetagent-scraper-price-index-2026-09
Agent scraper price index, September 2026
The rows behind the report Agent scraper price index, September 2026 on cracked.ai: what 1,000 results cost for 18 of the most requested scraping and search capabilities, tool by tool, through three routes.
Method
For each capability, every live candidate tool on Cracked is listed with three prices for 1,000 results: the price billed on Cracked's own provider account (cost_1000_cracked_usd, provider price plus the $0.001… See the full description on the dataset page: https://huggingface.co/datasets/crackedvibe/agent-scraper-price-index-2026-09.scrape-content-dataset-v1
Scrape Content Dataset v1
A human-curated benchmark dataset for evaluating web scraping engines on content quality.
Overview
This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time.
Dataset Structure
CSV format with columns:
id: Sequential… See the full description on the dataset page: https://huggingface.co/datasets/zhoubinghong/scrape-content-dataset-v1.Scraped_Dataset_Articlesquotes-scrapedRaw-Web-Scraped-to-JSONamazon-scrape-4-llm
Amazon Scrape 4 llm
Purpose?
Feed an LLM raw html to identify products from an ecommerce platform.These datasets contain the extracted innerTexts of all HTML nodes from different ecommerce product pages.The cleaning process significantly reduces the token size from ex: 450k -> 6k
Quickstart
from datasets import load_dataset
data_train = load_dataset("timashan/amazon-scrape-4-llm", "phones")
data_test = load_dataset("timashan/amazon-scrape-4-llm", "laptops")… See the full description on the dataset page: https://huggingface.co/datasets/timashan/amazon-scrape-4-llm.medquad_scraped_contextlc-scrapedAI_Articles_Scraped_from_arXiv-Semantic_Scholar
📘 AI Articles Scraped from arXiv & Semantic Scholar
🧩 Description
This dataset contains information on articles related to major AI conferences such as AAAI, NeurIPS, IJCAI, ICML, ICLR, collected through scraping from ArXiv and Semantic Scholar.It is intended to be used as a training dataset for various model training tasks and other desired uses.
📂 File Structure
File
Description
AI_Titles_v2025.csv
Main dataset
README.md
This file… See the full description on the dataset page: https://huggingface.co/datasets/d-e-c-d/AI_Articles_Scraped_from_arXiv-Semantic_Scholar.league_of_legends_wiki_scrapeThis dataset is a scrape from the League of Legends wiki, which contains the most up-to-date version with 166 champions. The data consists of: champion name, champion icon URL, champion wiki URL, stats, biography, passive ability, ability 1, ability 2, ability 3, ability 4, and curiosities.
scraped_cleaned_automobile.csvinstagram-scraperScrape_finalscrapebooks-to-scrape-page1
Books to Scrape – Page 1
Dataset Summary
Book records scraped from the first page of the Books to Scrape demo site.I created this dataset for a class assignment to practise web scraping, pandas,
and publishing a dataset to the Hugging Face Hub.
Data Collection
Source: https://books.toscrape.com/ (public test site for scraping practice)
Method: requests.get("https://books.toscrape.com/catalogue/page-1.html")
Parsed with BeautifulSoup, selecting each <article… See the full description on the dataset page: https://huggingface.co/datasets/TiaDay/books-to-scrape-page1.Small_Scrape_data
