datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-crawl-sample
Common Crawl sample
A small unofficial random subset of the famous Common Crawl dataset.
60 random segment WET files were downloaded from Common Crawl on 2024-05-12.
Lines between 500 and 5000 characters long (inclusive) were kept.
Only unique texts were kept.
No other filtering.
Languages
Each text was assigned to one of the language codes using the GCLD3 Python package.
The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.web-crawl-2026
Web Crawl 2026
A large-scale web crawl dataset for language model pretraining, collected by the OpenTransformer project.
Dataset Description
This dataset contains text extracted from web pages crawled directly from the internet using custom high-throughput crawlers. All data is freshly scraped.
Data Format
Each record is a JSON line (gzipped) with fields:
text: extracted text content (200-200,000 chars)
url: source URL
domain: source domain
timestamp: crawl… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-2026.CrawlPT_dedup
CrawlPT (deduplicated)
CrawlPT is a generic Portuguese corpus extracted from various web pages.
This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022).
The raw version is also available here.
Dataset Details
Dataset is composed by three corpora:
brWaC, C100-PT, OSCAR-2301.
brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites.
C100-PT: Portuguese subset from CC-100. C100 was… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/CrawlPT_dedup.mirror-eduagarcia__CrawlPT_dedup
CrawlPT (deduplicated)
CrawlPT is a generic Portuguese corpus extracted from various web pages.
This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022).
The raw version is also available here.
Dataset Details
Dataset is composed by three corpora:
brWaC, C100-PT, OSCAR-2301.
brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites.
C100-PT: Portuguese subset from CC-100.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__CrawlPT_dedup.common-crawl-english-filtered
🧠 FineWeb-English-Filtered
📘 Dataset Summary
FineWeb-English-Filtered is a large-scale, cleaned, English-only text dataset derived from Common Crawl’s WET archives.It contains 940 million documents of publicly available web text, converted into Apache Parquet format with a consistent schema for fast and efficient data loading.
The dataset was generated using a custom AWS Glue pipeline that processed, filtered, and merged .wet files across multiple terabytes of… See the full description on the dataset page: https://huggingface.co/datasets/anandjh8/common-crawl-english-filtered.icelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3).
uz-crawl
Dataset Card for UzCrawl
Dataset Summary
In an effort to democratize research on low-resource languages, we release UzCrawl dataset, a web and telegram crawl corpus consisting of materials from nearly 1.2 million unique sources in the Uzbek Language.
Please refer to our blogpost for further details.
P.S. We updated the dataset with 2nd version that extends the scope to new topics as well as being up to date to March 2024.
To load and use dataset, run this script:
from… See the full description on the dataset page: https://huggingface.co/datasets/tahrirchi/uz-crawl.common-crawl-zhtw
Dataset Card for Common Crawl Traditional Chinese
De-duplicated version of jed351/Traditional-Chinese-Common-Crawl-Filtered.
De-duplicated with MinHash
Is suggested to filter the dataset with NLU models before any serious use.
Booking-Hotel-Reviews-Dataset
Booking.com Hotel Reviews Dataset – 4.3K Sample
A rich, structured dataset of hotel reviews collected from Booking.com, featuring a unique split of positive and negative review text, reviewer country, stay dates, traveler tags, and hotel location data. Ideal for sentiment analysis, aspect-based opinion mining, travel AI, hospitality recommendation systems, and LLM fine-tuning on real-world review data.
Dataset Overview
Field
Details
Source… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Booking-Hotel-Reviews-Dataset.Curated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.IKEA-Home-Decor-Furniture-Dataset
IKEA Home Decor & Furniture Product Dataset
A rich, structured dataset of IKEA home decor and furniture products, featuring deep category taxonomies, full product descriptions, measurements, features, and image URLs. Ideal for training product recommendation models, interior design AI applications, multimodal models, and e-commerce search systems.
Dataset Overview
Field
Details
Source
IKEA (multi-country)
Total Records
400+
Category Focus
Home… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/IKEA-Home-Decor-Furniture-Dataset.Trustpilot-Reviews-Dataset-20K-Sample
Trustpilot Reviews Dataset – 20K Sample
This dataset contains a curated sample of 20,000 English-language user reviews sourced exclusively from Trustpilot.com. It is a representative subset of our larger collection containing over 1 million Trustpilot reviews across various industries and companies.
🗂️ Dataset Overview
Source: Trustpilot
Total Records: 20,000
Language: English
Industries: E-commerce, SaaS, Travel, Finance, Education, and more
Use Case: NLP tasks… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Trustpilot-Reviews-Dataset-20K-Sample.moltbook-crawl
Moltbook Crawl
A comprehensive crawl of Moltbook, a Reddit-style social media platform exclusively populated by AI agents built on the OpenClaw framework. This dataset captures the platform's early growth phase and provides a unique empirical window into AI agent collective behavior.
Dataset Description
The dataset is provided as a single SQLite database (moltbook.db) containing posts, comments, agent profiles, submolt (community) metadata, and longitudinal snapshots of… See the full description on the dataset page: https://huggingface.co/datasets/giordano-dm/moltbook-crawl.tt-azatliq-crawl
Dataset Summary
AzatliqCrawl is a document-level dataset in Tatar language based on Azatliq newspaper.
There are two versions released: the noisy dataset, which has no filtering, and the clean dataset, which has a variety of filters applied (language identification using fasstext BOW and deduplication using MinHashLSH with number of permutations equal to 128 and threshold equal to 0.9), though it naturally has a fair amount of noise itself. Each dataset is released in a… See the full description on the dataset page: https://huggingface.co/datasets/veryrealtatarperson/tt-azatliq-crawl.vi_math_problem_crawl
Dataset Card for Vietnamese Elementary Math Knowledge and Workbook
Dataset Summary
The data includes information about elementary school math knowledge in Vietnam, as well as exercises compiled from books. This is a crawlable dataset that can be trained for text generation tasks.
Supported Tasks and Leaderboards
Languages
The majority of the data is in Vietnamese, but there is still some English from some bilingual workbooks.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hllj/vi_math_problem_crawl.medical_sft_crawl_vi_10k_v1
ViMed-SFT: Vietnamese Medical Conversational Dataset
Dataset Description
ViMed-SFT is a high-quality Vietnamese medical conversational dataset designed for Supervised Fine-Tuning (SFT) of Large Language Models. The dataset contains medical Q&A conversations between users and a virtual medical assistant.
Dataset Summary
Attribute
Value
Language
Vietnamese
Domain
Healthcare / Medical
Task
Conversational AI, Instruction Tuning
Samples… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/medical_sft_crawl_vi_10k_v1.dolma3_mix-common_crawl-art_and_design-160kThe 160K subset of AllenAI's common_crawl-art_and_design Pretraining dataset split into train a valid saamples.
Train set size: 159436
Valid set size: 160
Direct usage in MLX-LM-LoRA:
python -m mlx_lm.lora \
--train \
--model Qwen/Qwen3-0.6B-Base \
--data mlx-community/dolma3_mix-common_crawl-art_and_design-160k \
--num-layers 4 \
--iters 1000 \
--batch-size 1 \
--steps-per-report 50 \
--max-seq-length 1028 \
--adapter-path path/to/adapter
Direct usage in MLX-LM:
python -m mlx_lm.lora \… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/dolma3_mix-common_crawl-art_and_design-160k.common-crawl-docx-sample
Common Crawl DOCX Sample
A sample of normalized text extracted from DOCX records in Common Crawl.
Source
Common Crawl release: CC-MAIN-YYYY-NN
Source index: Common Crawl URL Index
Pipeline: marin-community/marin
Pipeline revision: REPLACE_WITH_GIT_SHA
Records were selected using declared DOCX MIME type, detected DOCX MIME type,
or a .docx URL suffix. Only successful, non-truncated index records were
eligible.
Processing
The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.HomeDepot-Smart-Home-Dataset
Home Depot Smart Home Product Dataset
A structured dataset of Smart Home products from Home Depot, featuring detailed product specifications, pricing, category taxonomy, highlights, color variants, dimensions, and brand data. Ideal for training product recommendation models, smart home AI assistants, price intelligence systems, and e-commerce search engines.
Dataset Overview
Field
Details
Source
Home Depot
Total Records
230+
Category Focus
Smart Home, IoT… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/HomeDepot-Smart-Home-Dataset.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.medical_sft_crawl_vi_10k_v1
ViMed-SFT: Vietnamese Medical Conversational Dataset
Dataset Description
ViMed-SFT is a high-quality Vietnamese medical conversational dataset designed for Supervised Fine-Tuning (SFT) of Large Language Models. The dataset contains medical Q&A conversations between users and a virtual medical assistant.
Dataset Summary
Attribute
Value
Language
Vietnamese
Domain
Healthcare / Medical
Task
Conversational AI, Instruction Tuning
Samples
10,085… See the full description on the dataset page: https://huggingface.co/datasets/quannguyen204/medical_sft_crawl_vi_10k_v1.crawlerlm-html-to-json
CrawlerLM: HTML to JSON Extraction
A synthetic instruction-tuning dataset for training language models to extract structured JSON from HTML.
Dataset Description
This dataset contains HTML paired with structured JSON extraction tasks in chat format. It's designed for fine-tuning small language models to perform structured data extraction from messy, real-world HTML across multiple domains.
Key Features
447 examples in instruction-tuning chat format
Real HTML… See the full description on the dataset page: https://huggingface.co/datasets/espsluar/crawlerlm-html-to-json.wikifacts_crawled
Wikipedia Did You Know Facts
A plain-text collection of ~41,755 short, self-contained facts taken from Wikipedia's "Did you know..." (DYK) hooks, with April Fools' Day joke hooks removed. Published by BRODILOFT.
Dataset Summary
Each line of the file is one fact written as a complete sentence, for example:
Sow thistles are named because they were fed to lactating sows.
Mary Hallaren was the first woman to join the United States Army.
The facts cover a wide range… See the full description on the dataset page: https://huggingface.co/datasets/BRODILOFT/wikifacts_crawled.tt-crawl
Dataset Summary
In an effort to democratize research on low-resource languages, we release TatarCrawl dataset, a web news corpus consisting of materials from nearly 15 unique sources in the Tatar Language.
To load and use dataset, run this script:
from datasets import load_dataset
tt_crawl=load_dataset("neurotatarlar/tt-crawl")
Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.walmart-reviews-dataset
🛒 Walmart Product Reviews Dataset (6.7K Records)
This dataset contains 6,700+ structured customer reviews from Walmart.com. Each entry includes product-level metadata along with review details, making it ideal for small-scale machine learning models, sentiment analysis, and ecommerce insights.
📑 Dataset Fields
Column
Description
url
Direct product page URL
name
Product name/title
sku
Product SKU (Stock Keeping Unit)
price
Product price (numeric, USD)… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/walmart-reviews-dataset.medical_sft_crawl_vi_7.1k_v1
ViMed-SFT: Vietnamese Medical Conversational Dataset
Dataset Description
ViMed-SFT is a high-quality Vietnamese medical conversational dataset designed for Supervised Fine-Tuning (SFT) of Large Language Models. The dataset contains medical Q&A conversations between users and a virtual medical assistant.
Dataset Summary
Attribute
Value
Language
Vietnamese
Domain
Healthcare / Medical
Task
Conversational AI, Instruction Tuning
Samples
7,124… See the full description on the dataset page: https://huggingface.co/datasets/quannguyen204/medical_sft_crawl_vi_7.1k_v1.
