CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes5.6k downloads2y agoHugging Face02OpenTransformer /web-crawl-2026 Web Crawl 2026 A large-scale web crawl dataset for language model pretraining, collected by the OpenTransformer project. Dataset Description This dataset contains text extracted from web pages crawled directly from the internet using custom high-throughput crawlers. All data is freshly scraped. Data Format Each record is a JSON line (gzipped) with fields: text: extracted text content (200-200,000 chars) url: source URL domain: source domain timestamp: crawl… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-2026.text-generation10B<n<100B1 likes5.3k downloads5mo agoHugging Face03eduagarcia /CrawlPT_dedup CrawlPT (deduplicated) CrawlPT is a generic Portuguese corpus extracted from various web pages. This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022). The raw version is also available here. Dataset Details Dataset is composed by three corpora: brWaC, C100-PT, OSCAR-2301. brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites. C100-PT: Portuguese subset from CC-100. C100 was… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/CrawlPT_dedup.tabulartext-generation100M<n<1B8 likes2.6k downloads3y agoHugging Face04leeaandrob /mirror-eduagarcia__CrawlPT_dedup CrawlPT (deduplicated) CrawlPT is a generic Portuguese corpus extracted from various web pages. This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022). The raw version is also available here. Dataset Details Dataset is composed by three corpora: brWaC, C100-PT, OSCAR-2301. brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites. C100-PT: Portuguese subset from CC-100.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__CrawlPT_dedup.tabulartext-generation100M<n<1B0 likes1.5k downloads3mo agoHugging Face05anandjh8 /common-crawl-english-filtered 🧠 FineWeb-English-Filtered 📘 Dataset Summary FineWeb-English-Filtered is a large-scale, cleaned, English-only text dataset derived from Common Crawl’s WET archives.It contains 940 million documents of publicly available web text, converted into Apache Parquet format with a consistent schema for fast and efficient data loading. The dataset was generated using a custom AWS Glue pipeline that processed, filtered, and merged .wet files across multiple terabytes of… See the full description on the dataset page: https://huggingface.co/datasets/anandjh8/common-crawl-english-filtered.texttext-generation100M<n<1B2 likes287 downloads4mo agoHugging Face06mideind /icelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3). texttext-generation1M<n<10M1 likes225 downloads4y agoHugging Face07tahrirchi /uz-crawl Dataset Card for UzCrawl Dataset Summary In an effort to democratize research on low-resource languages, we release UzCrawl dataset, a web and telegram crawl corpus consisting of materials from nearly 1.2 million unique sources in the Uzbek Language. Please refer to our blogpost for further details. P.S. We updated the dataset with 2nd version that extends the scope to new topics as well as being up to date to March 2024. To load and use dataset, run this script: from… See the full description on the dataset page: https://huggingface.co/datasets/tahrirchi/uz-crawl.texttext-generation1M<n<10M18 likes221 downloads2y agoHugging Face08liswei /common-crawl-zhtw Dataset Card for Common Crawl Traditional Chinese De-duplicated version of jed351/Traditional-Chinese-Common-Crawl-Filtered. De-duplicated with MinHash Is suggested to filter the dataset with NLU models before any serious use. texttext-generation1M<n<10M6 likes198 downloads2y agoHugging Face09crawlfeeds /Booking-Hotel-Reviews-Dataset Booking.com Hotel Reviews Dataset – 4.3K Sample A rich, structured dataset of hotel reviews collected from Booking.com, featuring a unique split of positive and negative review text, reviewer country, stay dates, traveler tags, and hotel location data. Ideal for sentiment analysis, aspect-based opinion mining, travel AI, hospitality recommendation systems, and LLM fine-tuning on real-world review data. Dataset Overview Field Details Source… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Booking-Hotel-Reviews-Dataset.tabulartext-classification1K<n<10K0 likes131 downloads6mo agoHugging Face10crawlfeeds /Curated-Fox-News-Headlines-and-Full-Text Curated Fox News Headlines and Full Text This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis. 📁 Dataset Format Format: CSV Encoding: UTF-8 Fields: headline: The article title or headline publish_date: Date the article was published (YYYY-MM-DD) content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.imagetext-classification1K<n<10K2 likes89 downloads1y agoHugging Face11crawlfeeds /IKEA-Home-Decor-Furniture-Dataset IKEA Home Decor & Furniture Product Dataset A rich, structured dataset of IKEA home decor and furniture products, featuring deep category taxonomies, full product descriptions, measurements, features, and image URLs. Ideal for training product recommendation models, interior design AI applications, multimodal models, and e-commerce search systems. Dataset Overview Field Details Source IKEA (multi-country) Total Records 400+ Category Focus Home… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/IKEA-Home-Decor-Furniture-Dataset.imagetext-classificationn<1K1 likes76 downloads6mo agoHugging Face12crawlfeeds /Trustpilot-Reviews-Dataset-20K-Sample Trustpilot Reviews Dataset – 20K Sample This dataset contains a curated sample of 20,000 English-language user reviews sourced exclusively from Trustpilot.com. It is a representative subset of our larger collection containing over 1 million Trustpilot reviews across various industries and companies. 🗂️ Dataset Overview Source: Trustpilot Total Records: 20,000 Language: English Industries: E-commerce, SaaS, Travel, Finance, Education, and more Use Case: NLP tasks… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Trustpilot-Reviews-Dataset-20K-Sample.texttext-classification10K<n<100K0 likes62 downloads6mo agoHugging Face13giordano-dm /moltbook-crawl Moltbook Crawl A comprehensive crawl of Moltbook, a Reddit-style social media platform exclusively populated by AI agents built on the OpenClaw framework. This dataset captures the platform's early growth phase and provides a unique empirical window into AI agent collective behavior. Dataset Description The dataset is provided as a single SQLite database (moltbook.db) containing posts, comments, agent profiles, submolt (community) metadata, and longitudinal snapshots of… See the full description on the dataset page: https://huggingface.co/datasets/giordano-dm/moltbook-crawl.text-classification1M<n<10M1 likes58 downloads8mo agoHugging Face14veryrealtatarperson /tt-azatliq-crawl Dataset Summary AzatliqCrawl is a document-level dataset in Tatar language based on Azatliq newspaper. There are two versions released: the noisy dataset, which has no filtering, and the clean dataset, which has a variety of filters applied (language identification using fasstext BOW and deduplication using MinHashLSH with number of permutations equal to 128 and threshold equal to 0.9), though it naturally has a fair amount of noise itself. Each dataset is released in a… See the full description on the dataset page: https://huggingface.co/datasets/veryrealtatarperson/tt-azatliq-crawl.texttext-generation100K<n<1M1 likes54 downloads2y agoHugging Face15hllj /vi_math_problem_crawl Dataset Card for Vietnamese Elementary Math Knowledge and Workbook Dataset Summary The data includes information about elementary school math knowledge in Vietnam, as well as exercises compiled from books. This is a crawlable dataset that can be trained for text generation tasks. Supported Tasks and Leaderboards Languages The majority of the data is in Vietnamese, but there is still some English from some bilingual workbooks. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hllj/vi_math_problem_crawl.texttext-generation10K<n<100K1 likes52 downloads3y agoHugging Face16vlinhd11 /medical_sft_crawl_vi_10k_v1 ViMed-SFT: Vietnamese Medical Conversational Dataset Dataset Description ViMed-SFT is a high-quality Vietnamese medical conversational dataset designed for Supervised Fine-Tuning (SFT) of Large Language Models. The dataset contains medical Q&A conversations between users and a virtual medical assistant. Dataset Summary Attribute Value Language Vietnamese Domain Healthcare / Medical Task Conversational AI, Instruction Tuning Samples… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/medical_sft_crawl_vi_10k_v1.texttext-generation10K<n<100K0 likes51 downloads19d agoHugging Face17mlx-community /dolma3_mix-common_crawl-art_and_design-160kThe 160K subset of AllenAI's common_crawl-art_and_design Pretraining dataset split into train a valid saamples. Train set size: 159436 Valid set size: 160 Direct usage in MLX-LM-LoRA: python -m mlx_lm.lora \ --train \ --model Qwen/Qwen3-0.6B-Base \ --data mlx-community/dolma3_mix-common_crawl-art_and_design-160k \ --num-layers 4 \ --iters 1000 \ --batch-size 1 \ --steps-per-report 50 \ --max-seq-length 1028 \ --adapter-path path/to/adapter Direct usage in MLX-LM: python -m mlx_lm.lora \… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/dolma3_mix-common_crawl-art_and_design-160k.texttext-generation100K<n<1M2 likes47 downloads9mo agoHugging Face18Tinuade /common-crawl-docx-sample Common Crawl DOCX Sample A sample of normalized text extracted from DOCX records in Common Crawl. Source Common Crawl release: CC-MAIN-YYYY-NN Source index: Common Crawl URL Index Pipeline: marin-community/marin Pipeline revision: REPLACE_WITH_GIT_SHA Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible. Processing The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.tabulartext-generation1K<n<10K0 likes47 downloads7d agoHugging Face19crawlfeeds /HomeDepot-Smart-Home-Dataset Home Depot Smart Home Product Dataset A structured dataset of Smart Home products from Home Depot, featuring detailed product specifications, pricing, category taxonomy, highlights, color variants, dimensions, and brand data. Ideal for training product recommendation models, smart home AI assistants, price intelligence systems, and e-commerce search engines. Dataset Overview Field Details Source Home Depot Total Records 230+ Category Focus Smart Home, IoT… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/HomeDepot-Smart-Home-Dataset.tabulartext-classificationn<1K0 likes43 downloads6mo agoHugging Face20crawlfeeds /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes37 downloads5mo agoHugging Face21quannguyen204 /medical_sft_crawl_vi_10k_v1 ViMed-SFT: Vietnamese Medical Conversational Dataset Dataset Description ViMed-SFT is a high-quality Vietnamese medical conversational dataset designed for Supervised Fine-Tuning (SFT) of Large Language Models. The dataset contains medical Q&A conversations between users and a virtual medical assistant. Dataset Summary Attribute Value Language Vietnamese Domain Healthcare / Medical Task Conversational AI, Instruction Tuning Samples 10,085… See the full description on the dataset page: https://huggingface.co/datasets/quannguyen204/medical_sft_crawl_vi_10k_v1.texttext-generation10K<n<100K0 likes36 downloads10mo agoHugging Face22espsluar /crawlerlm-html-to-json CrawlerLM: HTML to JSON Extraction A synthetic instruction-tuning dataset for training language models to extract structured JSON from HTML. Dataset Description This dataset contains HTML paired with structured JSON extraction tasks in chat format. It's designed for fine-tuning small language models to perform structured data extraction from messy, real-world HTML across multiple domains. Key Features 447 examples in instruction-tuning chat format Real HTML… See the full description on the dataset page: https://huggingface.co/datasets/espsluar/crawlerlm-html-to-json.texttext-generationn<1K2 likes36 downloads9mo agoHugging Face23BRODILOFT /wikifacts_crawled Wikipedia Did You Know Facts A plain-text collection of ~41,755 short, self-contained facts taken from Wikipedia's "Did you know..." (DYK) hooks, with April Fools' Day joke hooks removed. Published by BRODILOFT. Dataset Summary Each line of the file is one fact written as a complete sentence, for example: Sow thistles are named because they were fed to lactating sows. Mary Hallaren was the first woman to join the United States Army. The facts cover a wide range… See the full description on the dataset page: https://huggingface.co/datasets/BRODILOFT/wikifacts_crawled.texttext-generation10K<n<100K1 likes35 downloads3d agoHugging Face24yasalma /tt-crawl Dataset Summary In an effort to democratize research on low-resource languages, we release TatarCrawl dataset, a web news corpus consisting of materials from nearly 15 unique sources in the Tatar Language. To load and use dataset, run this script: from datasets import load_dataset tt_crawl=load_dataset("neurotatarlar/tt-crawl") texttext-generation1M<n<10M0 likes32 downloads2y agoHugging Face25crawlfeeds /Medium-Articles-Corpus Medium Articles Corpus (10K Sample) The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers. This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles Dataset Features This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.imagetext-classification10K<n<100K2 likes32 downloads1y agoHugging Face26crawlfeeds /walmart-reviews-dataset 🛒 Walmart Product Reviews Dataset (6.7K Records) This dataset contains 6,700+ structured customer reviews from Walmart.com. Each entry includes product-level metadata along with review details, making it ideal for small-scale machine learning models, sentiment analysis, and ecommerce insights. 📑 Dataset Fields Column Description url Direct product page URL name Product name/title sku Product SKU (Stock Keeping Unit) price Product price (numeric, USD)… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/walmart-reviews-dataset.tabulartext-classification1K<n<10K0 likes27 downloads1y agoHugging Face27quannguyen204 /medical_sft_crawl_vi_7.1k_v1gated ViMed-SFT: Vietnamese Medical Conversational Dataset Dataset Description ViMed-SFT is a high-quality Vietnamese medical conversational dataset designed for Supervised Fine-Tuning (SFT) of Large Language Models. The dataset contains medical Q&A conversations between users and a virtual medical assistant. Dataset Summary Attribute Value Language Vietnamese Domain Healthcare / Medical Task Conversational AI, Instruction Tuning Samples 7,124… See the full description on the dataset page: https://huggingface.co/datasets/quannguyen204/medical_sft_crawl_vi_7.1k_v1.texttext-generation1K<n<10K0 likes21 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.