datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-crawler-index
AI Crawler Index
150 web crawlers and AI user agents from 74 operators — what each one is for,
what blocking it costs you, and the IP ranges its operator publishes.
Plus a compiled user-agent regex and the union of 1997 IPv4 and 1062 IPv6
prefixes from 15 operator-published range files.
Home: https://www.pathwren.workers.dev/c/huggingface-datasets/ · CC0 · no signup, no key.
What this is, plainly
This is an independent, non-commercial automated project. It is run… See the full description on the dataset page: https://huggingface.co/datasets/pathwren/ai-crawler-index.Booking-Hotel-Reviews-Dataset
Booking.com Hotel Reviews Dataset – 4.3K Sample
A rich, structured dataset of hotel reviews collected from Booking.com, featuring a unique split of positive and negative review text, reviewer country, stay dates, traveler tags, and hotel location data. Ideal for sentiment analysis, aspect-based opinion mining, travel AI, hospitality recommendation systems, and LLM fine-tuning on real-world review data.
Dataset Overview
Field
Details
Source… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Booking-Hotel-Reviews-Dataset.Curated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.IKEA-Home-Decor-Furniture-Dataset
IKEA Home Decor & Furniture Product Dataset
A rich, structured dataset of IKEA home decor and furniture products, featuring deep category taxonomies, full product descriptions, measurements, features, and image URLs. Ideal for training product recommendation models, interior design AI applications, multimodal models, and e-commerce search systems.
Dataset Overview
Field
Details
Source
IKEA (multi-country)
Total Records
400+
Category Focus
Home… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/IKEA-Home-Decor-Furniture-Dataset.Trustpilot-Reviews-Dataset-20K-Sample
Trustpilot Reviews Dataset – 20K Sample
This dataset contains a curated sample of 20,000 English-language user reviews sourced exclusively from Trustpilot.com. It is a representative subset of our larger collection containing over 1 million Trustpilot reviews across various industries and companies.
🗂️ Dataset Overview
Source: Trustpilot
Total Records: 20,000
Language: English
Industries: E-commerce, SaaS, Travel, Finance, Education, and more
Use Case: NLP tasks… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Trustpilot-Reviews-Dataset-20K-Sample.para_crawl_cscsshelfglance-ai-crawler-blocking
133 of 10,099 Shopify stores block an AI crawler. 6 block the one ChatGPT shops with.
The robots.txt of 10,099 Shopify storefronts, read for twelve AI crawlers. 133 block at least one, almost all of them training crawlers by a copied whole-site rule; 13 block a crawler that fetches at answer time, and no store blocks the shopping crawler while allowing the training one.
One row per Shopify storefront, out of 10,099 read between 29 August 2026 and 2 September 2026, whose… See the full description on the dataset page: https://huggingface.co/datasets/bpm2007/shelfglance-ai-crawler-blocking.HomeDepot-Smart-Home-Dataset
Home Depot Smart Home Product Dataset
A structured dataset of Smart Home products from Home Depot, featuring detailed product specifications, pricing, category taxonomy, highlights, color variants, dimensions, and brand data. Ideal for training product recommendation models, smart home AI assistants, price intelligence systems, and e-commerce search engines.
Dataset Overview
Field
Details
Source
Home Depot
Total Records
230+
Category Focus
Smart Home, IoT… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/HomeDepot-Smart-Home-Dataset.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.para_crawl_multi_smallsearch-vs-store
Search vs. Store — AI app search demand vs App Store rank vs web traffic (US, 2026)
One row per major AI app joining three independent popularity signals. Two
cross-sectional US snapshots (2026-07-01 and 2026-07-26), 40 curated AI apps each.
Aggregate/derived only — no raw records, no PII. License: CC BY 4.0.
Study + methodology: https://crawlora.net/blog/search-vs-store-2026?utm_source=huggingface&utm_medium=referral&utm_campaign=search-vs-store
Dataset page:… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/search-vs-store.para_crawl_multi_allpara_crawl_dedepara_crawl_enenpara_crawl_slslCrawList
CrawList
CrawList is a dataset of URLs to crawl for extra information in multiple tasks, such as retrieval of public documentation or grounding LLMs on the latest data existing on the internet.
Dataset Details
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More Information Needed]
Out-of-Scope Use… See the full description on the dataset page: https://huggingface.co/datasets/dalitics/CrawList.crawlee-commits
