datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb_2_samples_hq
fineweb_2_samples_hq
FINEWEB2-HQ dataset
Dataset Structure
This dataset contains 5 JSONL files with a total size of 26415.31 MB.
Files:
ukr_Cyrl_sample_001.jsonl: 6163.40 MB
ron_Latn_sample_001.jsonl: 3739.69 MB
kor_Hang_sample_001.jsonl: 4120.89 MB
hin_Deva_sample_001.jsonl: 6681.96 MB
heb_Hebr_sample_001.jsonl: 5709.37 MB
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/this/dataset")
Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.Fineweb_2_500k_removedFineWeb2-HQ-ja-20B元のデータセットFineWeb2-HQ
元のデータセットは多言語で巨大なため、扱いやすい用に日本語データを約200GBだけ抽出したデータセットです
wc 結果
1763269 38541549 5370473709 fineweb_jpn_Jpan_chunk_0.jsonl
1784158 37430170 5370514369 fineweb_jpn_Jpan_chunk_1.jsonl
1639554 40065129 5370372344 fineweb_jpn_Jpan_chunk_10.jsonl
1575127 42167166 5370298354 fineweb_jpn_Jpan_chunk_11.jsonl
1686375 39225898 5370402506 fineweb_jpn_Jpan_chunk_12.jsonl
1786948 36456352 5370498572… See the full description on the dataset page: https://huggingface.co/datasets/dahara1/FineWeb2-HQ-ja-20B.fineweb_2_500k_both_deduplicatedFineweb_2_500k_filteredFineweb_2_500k_bothfineweb2-chinese
FineWeb2 - Chinese
From cmn_Hani subset
Region
Rows
MAINLAND_CHINA
834356
TAIWAN
83875
HONG_KONG
10411
AMBIGUOUS_TRADITIONAL_TW_OR_HK
67700
OTHER
3658
fineweb-200-weightedepfml-FineWeb2-HQ-sample
epfml/FineWeb2-HQ
A curated subset of the epfml/FineWeb2-HQ dataset featuring high-quality multilingual text.
Details
First 25 000 rows per config (language and script pair)
Duplicates removed
Texts truncated to 512 LLaMA 3.1 tokens
Scores transformed with log10
Rows shuffled and 20% of the rows split into the test set (stratified by config)
Example
{
"text": "爵士大师Tim Garland 深圳专场 - [jazz]\nTim Garland Lighthouse… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/epfml-FineWeb2-HQ-sample.AraToken-FineWeb2-HQ-ar
AraToken FineWeb2-HQ Arabic splits
These are the exact document splits used in
AraToken (code). They were drawn from the arb_Arab part of
epfml/FineWeb2-HQ and are redistributed
under its ODC-By 1.0 license.
A document goes to a split by blake2b(id, digest_size=8) mod 1000, so the splits are disjoint by
construction.
split
buckets
role
documents
characters
lm-train
0–599
LEP and CPT adaptation
662,685
2.00 B
tok-train
600–899
tokenizer training and pruning
159,826… See the full description on the dataset page: https://huggingface.co/datasets/mariklolik/AraToken-FineWeb2-HQ-ar.fineweb2-ro-human
FineWeb2-Ro-Human
FineWeb2-Ro-Human is a human annotated sample that is extended into FineWeb2-Ro-LLM.
More details can be found here.
Usage
You can load this dataset using the Hugging Face datasets library:
from datasets import load_dataset
dataset = load_dataset("OpenLLM-Ro/fineweb2-ro-human", split="train")
FineWeb2-Edu-JA-EN
FineWeb2-Edu-JA-EN: Japanese-English Parallel Dataset
FineWeb2-Edu-JA-EN is a high-quality, information-rich Japanese-English parallel corpus curated specifically for educational and academic domains. It is derived from a subset of hotchpotch/fineweb-2-edu-japanese, which was filtered, clustered, and translated into English using Google Translate.
Original Dataset: hotchpotch/fineweb-2-edu-japanese
Size: 2,168 rows
License: ODC-BY
Dataset Creation & Methodology… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/FineWeb2-Edu-JA-EN.fineweb2hq-vs-c4This dataset includes 5000 rows per language from each of two sources: the higher-quality epfml/FineWeb2-HQ
and the lower-quality allenai/c4. The data is split 80/20 into training and test sets.
Languages were carefully chosen to ensure balanced representation across both splits:
Arabic, Chinese, Czech, Danish, Dutch, French, German, Greek, Hungarian, Indonesian, Italian, Japanese, Persian, Polish, Portuguese, Russian, Spanish, Swedish, Turkish, and Vietnamese.
aleph-alpha-germanweb-fineweb2-filtered-fair
Aleph Alpha GermanWeb — Filtered Fair FineWeb2 corpus
This gated repository preserves the large filtered FineWeb2 corpus consumed by
the final Aleph Alpha GermanWeb fair pipeline.
Contents
data/: Aleph-Alpha-GermanWeb-fineweb2-filtered-fair
Access and provenance
This repository is publicly visible but requires manual access approval.
Recipients must comply with upstream FineWeb2 and Aleph Alpha GermanWeb terms.
anima-corpus-ko-fineweb2-broad
anima-corpus-ko-fineweb2-broad
🇰🇷 한국어 broad (일반) 코퍼스 for anima conv 303M byte-level pretraining.
anima chat register 표준(a_chat_registers)의 4칸 {ko·en} × {일반·SNS} 중 ko-일반 칸을 메우기 위한 데이터셋이다. 기존 ko-일반 source 가 ~1.7MB 로 빈약(en ~202MB 대비)했던 갭을 FineWeb-2 한국어로 보강한다.
Source
Upstream: HuggingFaceFW/fineweb-2, config kor_Hang (한국어 한글 스크립트).
Extracted from: train parquet data/kor_Hang/train/000_00000.parquet (1 of 25 shards).
Field: text 컬럼만 추출 (raw UTF-8, byte-vocab256… See the full description on the dataset page: https://huggingface.co/datasets/dancinlab/anima-corpus-ko-fineweb2-broad.fineweb-2-ml4mfineweb-2-ml4mhindi-fineweb2-40bfineweb2-id-filtered-10k
FineWeb2-ID Filtered (min 10.000 karakter)
Dataset ini adalah subset turunan dari HuggingFaceFW/fineweb-2,
subset bahasa Indonesia (ind_Latn), yang disaring dengan syarat:
Panjang teks (setelah strip whitespace) minimal 10.000 karakter per dokumen
Diambil secara streaming dari split train, urutan asli (tanpa shuffle)
Jumlah sampel: 2,097,152 dokumen, tersebar di 80 file JSONL.
Sumber & Lisensi
Seluruh isi teks berasal dari FineWeb-2 (lisensi ODC-By), yang pada… See the full description on the dataset page: https://huggingface.co/datasets/tampakwilll/fineweb2-id-filtered-10k.
