datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb_2_samples_hq
fineweb_2_samples_hq
FINEWEB2-HQ dataset
Dataset Structure
This dataset contains 5 JSONL files with a total size of 26415.31 MB.
Files:
ukr_Cyrl_sample_001.jsonl: 6163.40 MB
ron_Latn_sample_001.jsonl: 3739.69 MB
kor_Hang_sample_001.jsonl: 4120.89 MB
hin_Deva_sample_001.jsonl: 6681.96 MB
heb_Hebr_sample_001.jsonl: 5709.37 MB
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/this/dataset")
Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.epfml-FineWeb2-HQ-sample
epfml/FineWeb2-HQ
A curated subset of the epfml/FineWeb2-HQ dataset featuring high-quality multilingual text.
Details
First 25 000 rows per config (language and script pair)
Duplicates removed
Texts truncated to 512 LLaMA 3.1 tokens
Scores transformed with log10
Rows shuffled and 20% of the rows split into the test set (stratified by config)
Example
{
"text": "爵士大师Tim Garland 深圳专场 - [jazz]\nTim Garland Lighthouse… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/epfml-FineWeb2-HQ-sample.AraToken-FineWeb2-HQ-ar
AraToken FineWeb2-HQ Arabic splits
These are the exact document splits used in
AraToken (code). They were drawn from the arb_Arab part of
epfml/FineWeb2-HQ and are redistributed
under its ODC-By 1.0 license.
A document goes to a split by blake2b(id, digest_size=8) mod 1000, so the splits are disjoint by
construction.
split
buckets
role
documents
characters
lm-train
0–599
LEP and CPT adaptation
662,685
2.00 B
tok-train
600–899
tokenizer training and pruning
159,826… See the full description on the dataset page: https://huggingface.co/datasets/mariklolik/AraToken-FineWeb2-HQ-ar.anima-corpus-ko-fineweb2-broad
anima-corpus-ko-fineweb2-broad
🇰🇷 한국어 broad (일반) 코퍼스 for anima conv 303M byte-level pretraining.
anima chat register 표준(a_chat_registers)의 4칸 {ko·en} × {일반·SNS} 중 ko-일반 칸을 메우기 위한 데이터셋이다. 기존 ko-일반 source 가 ~1.7MB 로 빈약(en ~202MB 대비)했던 갭을 FineWeb-2 한국어로 보강한다.
Source
Upstream: HuggingFaceFW/fineweb-2, config kor_Hang (한국어 한글 스크립트).
Extracted from: train parquet data/kor_Hang/train/000_00000.parquet (1 of 25 shards).
Field: text 컬럼만 추출 (raw UTF-8, byte-vocab256… See the full description on the dataset page: https://huggingface.co/datasets/dancinlab/anima-corpus-ko-fineweb2-broad.fineweb2-id-filtered-10k
FineWeb2-ID Filtered (min 10.000 karakter)
Dataset ini adalah subset turunan dari HuggingFaceFW/fineweb-2,
subset bahasa Indonesia (ind_Latn), yang disaring dengan syarat:
Panjang teks (setelah strip whitespace) minimal 10.000 karakter per dokumen
Diambil secara streaming dari split train, urutan asli (tanpa shuffle)
Jumlah sampel: 2,097,152 dokumen, tersebar di 80 file JSONL.
Sumber & Lisensi
Seluruh isi teks berasal dari FineWeb-2 (lisensi ODC-By), yang pada… See the full description on the dataset page: https://huggingface.co/datasets/tampakwilll/fineweb2-id-filtered-10k.
