datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-pretraining-corpusnemotron_qa_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
nemotron_actual_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
nemotron_fineinstructions_1T_exp_chat
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
nemotron_synthetic_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
Chinese-pretraining-datasetData source: https://github.com/CVI-SZU/Linly/wiki/Linly-OpenLLaMA
tweaktron-pretraining-data-2agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.gliner-biomed-pre-training
GLiNER-BioMed pre-training dataset
This dataset, used for the pre-training stage of the GLiNER-BioMed models, was introduced in the paper GLiNER-BioMed: a suite of efficient models for open biomedical named entity recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@article{yazdani2026gliner,
author = {Yazdani, Anthony and Stepanov, Ihor and Teodoro, Douglas},
title = {{GLiNER-BioMed}: a suite of efficient… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-pre-training.nemotron_wrap_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
HuatuoGPT2-Pretraining-Instruction
HuatuoGPT2-Pretraining-Instruction-5200K
Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT.
This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible.
Data Volume
The following table details the volume and distribution of pre-training data for HuatuoGPT2:
Data Source
Data Volume
Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.wikipedia_en_512_for_pretraining
Cleaned Wikipedia 512 Pretraining Dataset
Dataset Description
This dataset is a cleaned version of the Hugging Face dataset lucadiliello/wikipedia_512_pretraining.
The original dataset contains English Wikipedia text prepared for language-model pretraining.
This derivative version applies additional filtering intended to remove malformed, duplicated, or incomplete training samples while otherwise preserving the original text.
Source
Original… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/wikipedia_en_512_for_pretraining.steuerllm_pretraining_dataset
SteuerLLM Pretraining Dataset
Project page | Paper | GitHub
Pretraining Dataset for German Tax Law filtered from FineWeb. This dataset was used for the continual pretraining stage of SteuerLLM, a specialized large language model for German tax law analysis.
Dataset Description
The SteuerLLM pretraining dataset is a domain-specific subset filtered from large-scale web corpora. It focuses on identifying and extracting tax-related content from German web data to adapt… See the full description on the dataset page: https://huggingface.co/datasets/windprak/steuerllm_pretraining_dataset.ipt_actual_all_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
ipt_fineinstructions_all_exp_chat
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
ipt_synthetic_all_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
ipt_fineinstructions_all_judged_exp_chat
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
ipt_fineinstructions_all_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
DCLM-pretraining-datasetjapanese-deduptest-codeact-pretrainingai-theories-corpus-en-pretraining
ai-theories en コーパス
ai-theories(https://github.com/kojikojiprg/ai-theories)プロジェクトの成果物。
006(小型 GPT の事前学習)・008 の英語条件用のコーパス。
ライセンスについての注記
このデータセット自体のライセンスは クリエイティブ・コモンズ 表示-継承 4.0 国際
(CC BY-SA 4.0) である。ai-theoriesの他の成果物(トークナイザなど)は
cc-by-nc-4.0を採用しているが、本データセットの内容はフリー百科事典
『ウィキペディア(Wikipedia)』の本文そのもの(wikitext を平文に変換したのみで、
内容は改変していない)であり、Wikipedia 本文自体のライセンス(CC BY-SA 4.0、
表示・継承の条件)を継承する必要があるため、別のライセンスとしている。
由来… See the full description on the dataset page: https://huggingface.co/datasets/kojikojiprg/ai-theories-corpus-en-pretraining.FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.pretraining_subsets_corpusEnglish-Pretraining-Datasetrepro-gram-modular-pretraining-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Arabic_Pretraining_10KBotanic1-pretraining
Botanic1 pretraining data
This dataset accompanies the Botanic1 technical report.
It preserves the genomic sequence shards used to pretrain the Botanic1 models,
including the original sequence case, augmentation margins, train/test separation,
source manifests, and assembly metadata.
Release contents
The initial release contains the main 8 kbp corpus used by Botanic1-S, M, L, and XL.
The context extension corpora at 16, 32, 64, and 128 kbp are planned for this… See the full description on the dataset page: https://huggingface.co/datasets/living-models/Botanic1-pretraining.150M-0.025x-DCLM-pretraining-dataset
