datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B.
fineweb-nanochatbpe-100M
fineweb-nanochatbpe-100M
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
This is a 100-million-token slice for data-constrained experiments. The train.bin
is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent
alexkstern/fineweb-nanochatbpe-20B
train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.fineweb-edu-zh-chengyu-cpt
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus
A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on
cultural knowledge in figurative language, built from the highest-quality
tier of opencsg/Fineweb-Edu-Chinese-V2.1.
Each document is educational Chinese text containing at least one culturally
vetted chengyu, with an appended 【成语注释】 knowledge block listing every
matched idiom's figurative meaning(s) and classical source citation.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.fineweb-nanochatbpe-20B
fineweb-nanochatbpe-20B
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
20,000,000,000
val.bin
val
52,336,096
train and val are disjoint held-out partitions. Each .bin is a raw
little-endian uint16 stream (no header); token count = filesize / 2, and
train.meta.json / val.meta.jsoncarry the full… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-20B.fineweb-edufineweb_2_samples_hq
fineweb_2_samples_hq
FINEWEB2-HQ dataset
Dataset Structure
This dataset contains 5 JSONL files with a total size of 26415.31 MB.
Files:
ukr_Cyrl_sample_001.jsonl: 6163.40 MB
ron_Latn_sample_001.jsonl: 3739.69 MB
kor_Hang_sample_001.jsonl: 4120.89 MB
hin_Deva_sample_001.jsonl: 6681.96 MB
heb_Hebr_sample_001.jsonl: 5709.37 MB
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/this/dataset")
Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.fineweb_2_500k_both_deduplicatedFineweb_2_500k_removedfineweb-edu-pretokenized-llama3-100b
FineWeb-Edu Pretokenized with Llama 3.1 (100B)
This repository contains the sample/100BT subset of FineWeb-Edu, pretokenized for Megatron-LM-style training with meta-llama/Meta-Llama-3.1-8B.
It is a derived, pretokenized version of the upstream data, not an official Hugging Face FineWeb release.
Dataset summary
140 indexed shards
97,270,686 non-empty documents
97,458,793,013 tokens
English web text from FineWeb-Edu sample/100BT
Source dataset revision:… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-llama3-100b.fineweb-c-progressfineweb_score4_sortedfineweb-eduFineweb_2_500k_bothfineweb_class5-0Fineweb_2_500k_filteredfinewebs-scandeval-results
ScandEval Results on English NLU
We use ScandEval in revision 8766d2a to conduct experiments with our pretrained FineWeb LMs.
Additionally, results for BERT, RoBERTa and ELECTRA were also performed to have a nice comparison.
Model ID
Avg. Score
CoNLL-En
SST5
ScaLA-En
SQuAD
model-garden-lms/bert-base-finewebs-1m
69.03
88.98 ± 0.43 / 88.67 ± 0.36
58.11 ± 1.2 / 59.77 ± 1.49
57.29 ± 3.57 / 77.15 ± 2.17
55.82 ± 1.35 / 66.46 ± 1.51
model-garden-lms/bert-base-finewebs-951k… See the full description on the dataset page: https://huggingface.co/datasets/model-garden-lms/finewebs-scandeval-results.fineweb2-chinese
FineWeb2 - Chinese
From cmn_Hani subset
Region
Rows
MAINLAND_CHINA
834356
TAIWAN
83875
HONG_KONG
10411
AMBIGUOUS_TRADITIONAL_TW_OR_HK
67700
OTHER
3658
fineweb-200-weighted1M-finewebedu-samples2048tTotal tokens in matching entries: 1_392_312_785
Average tokens per entry: 1392.31
FlofloB__smollm2-135M_pretrained_1000k_fineweb-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1000k_fineweb
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1000k_fineweb
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1000k_fineweb-details.FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-details.FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.1k-creative-writing-8kt-fineweb-edu-sampleTotal tokens in matching entries: 5_575_157
Average tokens per entry: 5575.16
20k-finewebedu-samples-8ktcreative-writing-2048-fineweb-edu-sampleCreative Writing:
keywords:
- "creative writing"
- "storytelling"
- "roleplaying"
- "narrative structure"
- "character development"
- "worldbuilding"
- "plot devices"
- "genre fiction"
- "writing techniques"
- "literary elements"
- "RPG storytelling"
- "interactive narrative"
max_entries: 2048
min_tokens: 512
max_tokens: 2048
min_int_score: 4
Total tokens in matching entries: 2218544
20k-finewebedu-samples-512tTotal tokens in matching entries: 7595043
Average tokens per entry: 379.75
100k-finewebedu-samples-8kt1M-finewebedu-samples256tTotal tokens in matching entries: 196_670_428
Average tokens per entry: 196.67
100k-finewebedu-samples4096tTotal tokens in matching entries: 275_639_417
Average tokens per entry: 2756.39
quantum-computing-2048-fineweb-sampleTest dataset for testing a data sampling and filtering script.
Configuration for filtering:
dataset: "HuggingFaceFW/fineweb-edu"
output_file: "quantum_computing_entries.jsonl"
state_file: "quantum_computing_dataset_state.json"
# Processing parameters
keywords:
- "quantum computing"
- "qubit"
- "quantum entanglement"
- "quantum supremacy"
- "quantum algorithm"
- "quantum error correction"
max_entries: 1024
min_tokens: 1024
max_tokens: 2048
min_int_score: 4
# Shuffling… See the full description on the dataset page: https://huggingface.co/datasets/Lambent/quantum-computing-2048-fineweb-sample.
