datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-2
🥂 FineWeb2
A sparkling update with 1000s of languages
What is it?
This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages.
The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments.
In particular, on the set of 9 diverse languages we used to guide our processing decisions, 🥂… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-2.FineWeb2-HQ
FineWeb2-HQ
Dataset summary
FineWeb2-HQ is a high-quality, model-filtered pretraining dataset derived as a subset of FineWeb2, spanning 20 languages. It enables around 6x faster pretraining compared to the base dataset. FineWeb2-HQ was created by selecting the top 10% quality documents of FineWeb2 in each language, based on scores assigned by a deep learning classifier trained to identify structured and knowledge-rich samples using XLM-RoBERTa embeddings.
Validation… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-HQ.FineWeb2-embedded
FineWeb2-embedded
Dataset summary
FineWeb2-embedded is an extension of the FineWeb2 dataset, annotated with document-level XLM-RoBERTa embeddings for 20 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research.
Since XLM-RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean-pooling 512 token chunks of the XLM-RoBERTa output. Therefore, longer texts… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-embedded.fineweb-2-dutchFineWeb2024
FineWeb-Edu 2024 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2024.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2024
Rows
162,500,784… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2024.fineweb2-bagaco
Bagaço 🍷🇵🇹
Bagaço is a pretraining dataset for European Portuguese. It filters the Fineweb2 dataset to URLs from Portuguese domains (e.g., .pt/). Each document is classified into one of 9 categories and scored for educational quality.
Filtering
Source: HuggingFaceFW/fineweb-2, subset por_Latn, split train
Filter: URLs containing .pt/ (Portuguese top-level domain)
Document classification
Each document is classified into one of 9 categories: Society, Arts… See the full description on the dataset page: https://huggingface.co/datasets/duarteocarmo/fineweb2-bagaco.FineWeb2025
FineWeb-Edu 2025 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2025.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2025
Rows
99,022,205… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2025.FineWeb-2023
FineWeb-Edu 2023 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2023.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2023
Rows
104,280,950… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb-2023.FineWeb-2020
FineWeb-Edu 2020 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2020.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2020
Rows
123,382,457… See the full description on the dataset page: https://huggingface.co/datasets/DataHound26/FineWeb-2020.FineWeb-2021
FineWeb-Edu 2021 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2021.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2021
Rows
139,636,993… See the full description on the dataset page: https://huggingface.co/datasets/DataHound26/FineWeb-2021.FineWeb-2022
FineWeb-Edu 2022 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2022.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2022
Rows
106,753,442… See the full description on the dataset page: https://huggingface.co/datasets/DataHound26/FineWeb-2022.fineweb_2_samples_hq
fineweb_2_samples_hq
FINEWEB2-HQ dataset
Dataset Structure
This dataset contains 5 JSONL files with a total size of 26415.31 MB.
Files:
ukr_Cyrl_sample_001.jsonl: 6163.40 MB
ron_Latn_sample_001.jsonl: 3739.69 MB
kor_Hang_sample_001.jsonl: 4120.89 MB
hin_Deva_sample_001.jsonl: 6681.96 MB
heb_Hebr_sample_001.jsonl: 5709.37 MB
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/this/dataset")
Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.fineweb-2-et
FineWeb2-et
The Estonian subset of HuggingFaceFW/fineweb-2 reuploaded for ease of access.
Licensing Information
The dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. The use of this dataset is also subject to CommonCrawl's Terms of Use.
Citation Information
@misc{penedo2025fineweb2pipelinescale,
title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language}… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/fineweb-2-et.fineweb-2-autocurate
fineweb-2-autocurate
Autonomously curated subsets of HuggingFaceFW/fineweb-2.
An LLM agent iteratively samples documents, identifies quality problems, and proposes heuristic fixes. Each fix is validated by training a small language model for 5 minutes and measuring BPB improvement on a Wikipedia eval set. Only fixes that improve BPB are kept.
Built with autocurate.
Subsets
Language
Subset
Original Docs
Kept Docs
Kept %
BPB Before → After
Improvement
Danish… See the full description on the dataset page: https://huggingface.co/datasets/bowang0911/fineweb-2-autocurate.fineweb-2-turkish-categorized-long
altaidevorg/fineweb-2-turkish-categorized long filtered Turkish texts
Source: altaidevorg/fineweb-2-turkish-categorized (config: default).
The script streamed 10,000,000 raw source rows before stopping. Categories
ads, adult content, sports, tabloid were rejected before length and quality filtering.
Retained rows contain 3,000–16,500 characters
and passed the iteration-5 Turkish language,
repetition, glue-word, punctuation, SEO, and soft information-density filters.
Selected… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/fineweb-2-turkish-categorized-long.fineweb-2-pol-latest
ReactiveAI - FineWeb2 PL subset
This dataset is derived from polish subset of FineWeb2 by HuggingFace. Includes latest ~8.5M examples.
Original dataset description below
🥂 FineWeb2
A sparkling update with 1000s of languages
What is it?
This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages.
The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/fineweb-2-pol-latest.fineweb-2-zhtw
Dataset Card for fineweb-2-zhtw
fineweb-2-zhtw 是以 Hugging Face FineWeb-2 的 cmn_Hani(官話/漢字)子集為來源,經多層產地與字形/用語過濾後,取出以台灣繁體中文為主的網頁語料,保留 FineWeb-2 的完整 metadata,可作為繁中持續預訓練語料。
與單純「挑出繁體字」的做法不同:本資料集把「繁體字」和「台灣繁體內容」當成兩件事處理。中國網站的 BIG5 轉碼閘道、香港媒體、以及訂房網站的機器翻譯頁面都會產出繁體字,但用語與語感並非台灣中文,這些都在過濾流程中被排除。
Dataset Details
來源 cmn_Hani/train 共 370 個 parquet 分片、約 636,058,984 列(約 1.6 TB)。經過濾後保留 11,416,204 列,留存率 **1.81%**。
Dataset Sources
Repository: lianghsun/fineweb-2-zhtw… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/fineweb-2-zhtw.fineweb-2-edu-zhtw
Dataset Card for fineweb-2-edu-zhtw
UltraX 清洗欄位
原有欄位(text 與 zhtw_*、edu_* 等共 22 欄)維持不變,另新增以下欄位,記錄以
openbmb/UltraX-0.6B-Preview 清洗的結果。
UltraX 不做端到端改寫,而是預測結構化編輯操作(keep_all / remove_all / remove_lines /
replace_str / add_line),再由程式確定性套用,因此每個決策都可被檢驗與否決。
欄位
說明
cleaned_text
清洗後文字,緊接 text 之後;needs_review 為 true 時等同原文
word_count / token_count
原文的詞數與 token 數(本資料集原本沒有,一併補上)
cleaned_word_count / cleaned_token_count
清洗後對應數值,同一套算法
word_reduction_pct /… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/fineweb-2-edu-zhtw.fineweb2-id-filtered-10k
FineWeb2-ID Filtered (min 10.000 karakter)
Dataset ini adalah subset turunan dari HuggingFaceFW/fineweb-2,
subset bahasa Indonesia (ind_Latn), yang disaring dengan syarat:
Panjang teks (setelah strip whitespace) minimal 10.000 karakter per dokumen
Diambil secara streaming dari split train, urutan asli (tanpa shuffle)
Jumlah sampel: 2,097,152 dokumen, tersebar di 80 file JSONL.
Sumber & Lisensi
Seluruh isi teks berasal dari FineWeb-2 (lisensi ODC-By), yang pada… See the full description on the dataset page: https://huggingface.co/datasets/tampakwilll/fineweb2-id-filtered-10k.
