CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pathwren /ai-crawler-index AI Crawler Index 150 web crawlers and AI user agents from 74 operators — what each one is for, what blocking it costs you, and the IP ranges its operator publishes. Plus a compiled user-agent regex and the union of 1997 IPv4 and 1062 IPv6 prefixes from 15 operator-published range files. Home: https://www.pathwren.workers.dev/c/huggingface-datasets/ · CC0 · no signup, no key. What this is, plainly This is an independent, non-commercial automated project. It is run… See the full description on the dataset page: https://huggingface.co/datasets/pathwren/ai-crawler-index.tabular1K<n<10K0 likes622 downloads8d agoHugging Face02crawlfeeds /Booking-Hotel-Reviews-Dataset Booking.com Hotel Reviews Dataset – 4.3K Sample A rich, structured dataset of hotel reviews collected from Booking.com, featuring a unique split of positive and negative review text, reviewer country, stay dates, traveler tags, and hotel location data. Ideal for sentiment analysis, aspect-based opinion mining, travel AI, hospitality recommendation systems, and LLM fine-tuning on real-world review data. Dataset Overview Field Details Source… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Booking-Hotel-Reviews-Dataset.tabulartext-classification1K<n<10K0 likes126 downloads6mo agoHugging Face03crawlfeeds /Curated-Fox-News-Headlines-and-Full-Text Curated Fox News Headlines and Full Text This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis. 📁 Dataset Format Format: CSV Encoding: UTF-8 Fields: headline: The article title or headline publish_date: Date the article was published (YYYY-MM-DD) content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.imagetext-classification1K<n<10K2 likes91 downloads1y agoHugging Face04crawlfeeds /IKEA-Home-Decor-Furniture-Dataset IKEA Home Decor & Furniture Product Dataset A rich, structured dataset of IKEA home decor and furniture products, featuring deep category taxonomies, full product descriptions, measurements, features, and image URLs. Ideal for training product recommendation models, interior design AI applications, multimodal models, and e-commerce search systems. Dataset Overview Field Details Source IKEA (multi-country) Total Records 400+ Category Focus Home… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/IKEA-Home-Decor-Furniture-Dataset.imagetext-classificationn<1K1 likes70 downloads6mo agoHugging Face05crawlfeeds /Trustpilot-Reviews-Dataset-20K-Sample Trustpilot Reviews Dataset – 20K Sample This dataset contains a curated sample of 20,000 English-language user reviews sourced exclusively from Trustpilot.com. It is a representative subset of our larger collection containing over 1 million Trustpilot reviews across various industries and companies. 🗂️ Dataset Overview Source: Trustpilot Total Records: 20,000 Language: English Industries: E-commerce, SaaS, Travel, Finance, Education, and more Use Case: NLP tasks… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Trustpilot-Reviews-Dataset-20K-Sample.texttext-classification10K<n<100K0 likes62 downloads6mo agoHugging Face06yawnick /para_crawl_cscstext10K<n<100K0 likes58 downloads3y agoHugging Face07bpm2007 /shelfglance-ai-crawler-blocking 133 of 10,099 Shopify stores block an AI crawler. 6 block the one ChatGPT shops with. The robots.txt of 10,099 Shopify storefronts, read for twelve AI crawlers. 133 block at least one, almost all of them training crawlers by a copied whole-site rule; 13 block a crawler that fetches at answer time, and no store blocks the shopping crawler while allowing the training one. One row per Shopify storefront, out of 10,099 read between 29 August 2026 and 2 September 2026, whose… See the full description on the dataset page: https://huggingface.co/datasets/bpm2007/shelfglance-ai-crawler-blocking.textn<1K0 likes45 downloads17d agoHugging Face08crawlfeeds /HomeDepot-Smart-Home-Dataset Home Depot Smart Home Product Dataset A structured dataset of Smart Home products from Home Depot, featuring detailed product specifications, pricing, category taxonomy, highlights, color variants, dimensions, and brand data. Ideal for training product recommendation models, smart home AI assistants, price intelligence systems, and e-commerce search engines. Dataset Overview Field Details Source Home Depot Total Records 230+ Category Focus Smart Home, IoT… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/HomeDepot-Smart-Home-Dataset.tabulartext-classificationn<1K0 likes41 downloads6mo agoHugging Face09crawlfeeds /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes36 downloads5mo agoHugging Face10crawlfeeds /Medium-Articles-Corpus Medium Articles Corpus (10K Sample) The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers. This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles Dataset Features This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.imagetext-classification10K<n<100K2 likes33 downloads1y agoHugging Face11yawnick /para_crawl_multi_smalltext10K<n<100K0 likes22 downloads3y agoHugging Face12crawlora-net /search-vs-store Search vs. Store — AI app search demand vs App Store rank vs web traffic (US, 2026) One row per major AI app joining three independent popularity signals. Two cross-sectional US snapshots (2026-07-01 and 2026-07-26), 40 curated AI apps each. Aggregate/derived only — no raw records, no PII. License: CC BY 4.0. Study + methodology: https://crawlora.net/blog/search-vs-store-2026?utm_source=huggingface&utm_medium=referral&utm_campaign=search-vs-store Dataset page:… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/search-vs-store.tabularn<1K0 likes22 downloads1mo agoHugging Face13yawnick /para_crawl_multi_alltext100K<n<1M0 likes21 downloads3y agoHugging Face14yawnick /para_crawl_dedetext10K<n<100K0 likes20 downloads3y agoHugging Face15yawnick /para_crawl_enentext10K<n<100K0 likes12 downloads3y agoHugging Face16yawnick /para_crawl_slsltext10K<n<100K0 likes12 downloads3y agoHugging Face17dalitics /CrawListgated CrawList CrawList is a dataset of URLs to crawl for extra information in multiple tasks, such as retrieval of public documentation or grounding LLMs on the latest data existing on the internet. Dataset Details Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More Information Needed] Out-of-Scope Use… See the full description on the dataset page: https://huggingface.co/datasets/dalitics/CrawList.tabular1K<n<10K0 likes7 downloads2y agoHugging Face18rsh-raj /crawlee-commitstext1K<n<10K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.