datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretraining_v1-omega_bookscarbon-pretraining-corpus
🧬 Carbon Pretraining Corpus
Description
173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model.
This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species.
Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.filtering-pretraining-mix-arrow-formatts-icl-pretraining-corpus
TS-ICL Pretraining Corpus (community reconstruction)
A unified, cleaned reconstruction of the univariate pretraining corpus described in
Table 5 of TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via
In-Context Learning (Le Naour, Nabil & Petralia, EDF R&D; arXiv:2606.05878). The TS-ICL
authors did not release their pretraining data pipeline, so this corpus is rebuilt from the
named upstream sources (LOTSA, Chronos, and the TempoPFN synthetic generators) and… See the full description on the dataset page: https://huggingface.co/datasets/JuaAI/ts-icl-pretraining-corpus.tabula-pretraining-corpus-v2
Tabula Pretraining Corpus v2
A large-scale synthetic tabular dataset for pretraining transformer-based in-context learning models for tabular data (similar to TabPFN).
Overview
Metric
Value
Total rows
272,271,776
Total datasets
10,867
Shards
135
Mean utility AUC
0.851
Format
Parquet (float32)
Schema
Each shard is a Parquet file with a fixed-width schema:
feat_0 through feat_63: Float32 feature columns. Unused slots are NaN.
target:… See the full description on the dataset page: https://huggingface.co/datasets/avewright/tabula-pretraining-corpus-v2.seq2seq-mixed-pretraining-SmolLM2deep-ignorance-pretraining-mix
Deep Ignorance Model Suite
We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-pretraining-mix.Sombench-pretraining-data
SomBench Pre-training Corpus: Multimodal Lunar Tiles
Dataset Summary
This includes a small sample from SomBench: a corpus of co-registered, multimodal lunar image tiles built for
large-scale self-supervised (foundation-model) pre-training. It contains a subset of modalities from the
low-resolution (WAC-anchored) and high-resolution (NAC-anchored) tracks specifically used in pretraining.
Tiles are anchored to individual LROC Experiment Data Record (EDR) image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-pretraining-data.contrastive-pretraining
Contrastive Pretraining
Per-language query/document pairs produced by the retrieval-common-crawl pipeline.
Each config corresponds to a single language or source with identical LightOn-style schema.
Config overview
Configs available are: fw-edu, fw2-arb_Arab, fw2-ces_Latn, fw2-cmn_Hani, fw2-dan_Latn, fw2-deu_Latn, fw2-ell_Grek, fw2-fas_Arab, fw2-fra_Latn, fw2-hun_Latn, fw2-ind_Latn, fw2-ita_Latn, fw2-jpn_Jpan, fw2-nld_Latn, fw2-pol_Latn, fw2-por_Latn, fw2-rus_Cyrl… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/contrastive-pretraining.mixed-pretraining-tokenizedncbi-dataset-for-genome-network-pretrainingcontrol-pretraining-filter-annotated
control-pretraining-filter-annotated
climbmix_full — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content)
climbmix_ai_docs — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content)
climbmix_long — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by… See the full description on the dataset page: https://huggingface.co/datasets/sudoers/control-pretraining-filter-annotated.pretraining-corpus
🧠 Kiy-K Synthetic Pretraining Corpus
Author: Khoi K. (@Kiy-K)License: Apache 2.0Last Updated: 2025-10-30
📘 Overview
The Kiy-K Synthetic Pretraining Corpus is a large-scale collection of synthetically generated English text designed for language model pretraining and instruction-tuning research.
All data is synthetic, created using open-source large language models such as GPT-OSS, NVIDIA Nemotron, and DeepSeek, under full control of the author.No real user… See the full description on the dataset page: https://huggingface.co/datasets/Kiy-K/pretraining-corpus.glossapi-greek-nanochat-pretraining-dataset-v2
GlossAPI Greek pretraining corpus v2
HPLT filtering method
The HPLT component is HPLT/ell_Grek_ge8_no_mt_clean60. It retains HPLT quality bins 8, 9, 10 (GE8), uses the pre-applied no-MT/register filter, requires greek_badness_score <= 60, and applies Wave4 Greek re-cleaning and normalization. The standalone filtered slice contained 48,728,774 documents; 48,629,460 remain after corpus-wide deduplication.
GlossAPI datasets and token counts
GlossAPI… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset-v2.ipt_fineinstructions_all
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
tiny-slm-pretraining-corpus
🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB)
A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures.
100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders.
📊 Dataset Statistics
Total Documents: 20,066,075
Train: 19,663,898
Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.pretraining_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 969,
"total_frames": 259815,
"total_tasks": 6,
"total_videos": 1938,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:969"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/pretraining_v1.tabula-pretraining-corpus
Tabula Pretraining Corpus
A continuously growing tabular pretraining corpus for the Tabula foundation model
(tabPFN-style in-context learning). Built by an autonomous agent that alternates
between harvesting permissively-licensed real datasets and generating high-quality
synthetic ones.
Usage
from datasets import load_dataset
# Load a specific batch config
ds = load_dataset("avewright/tabula-pretraining-corpus", name="datagen_001")
# Load all configs
from… See the full description on the dataset page: https://huggingface.co/datasets/avewright/tabula-pretraining-corpus.peptideclm-2-pretraining-data
PeptideMTR Training Data
This repository contains the dataset for the PeptideMTR paper. It is designed for SMILES encoder models trained by masked-language modeling (MLM) and/or multi-target regression (MTR) tasks, focusing on mapping peptide sequences to biochemical properties.
Link to the manuscript will be added here when available.
Dataset Summary
The dataset includes peptide sequences paired with 99 RDKit-derived descriptors representing various physicochemical… See the full description on the dataset page: https://huggingface.co/datasets/aaronfeller/peptideclm-2-pretraining-data.pretraining-high-quality
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness of the text
lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.wikipedia_en_512_for_pretraining
Cleaned Wikipedia 512 Pretraining Dataset
Dataset Description
This dataset is a cleaned version of the Hugging Face dataset lucadiliello/wikipedia_512_pretraining.
The original dataset contains English Wikipedia text prepared for language-model pretraining.
This derivative version applies additional filtering intended to remove malformed, duplicated, or incomplete training samples while otherwise preserving the original text.
Source
Original… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/wikipedia_en_512_for_pretraining.bookcorpusopen_with_ids_chunked
Dataset Card for "bookcorpusopen_with_ids_chunked"
More Information needed
pretraining-lower-quality
Dataset Card for Lapa Pretraining Lower Quality Dataset
Dataset Description
Dataset Summary
This dataset is a high quality (but lower quality than https://huggingface.co/datasets/lapa-llm/pretraining-high-qualit) subset of pretraining corpus for Ukrainian language.
It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality.pretraining-pretokenized-smollm3
SmolLM3 Pretokenized Pretraining Sources
Datatrove/Nanotron tokenized-byte versions of three public pretraining sources:
fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT
finemath-4plus: HuggingFaceTB/finemath, finemath-4plus
stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure
All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.DCLM-pretraining-datasetgotriple-pretraining-dataset
GoTriple Pretraining dataset
Summary
The GoTriple Pre-training Dataset is a multilingual corpus built from open-access research artefacts harvested via the GoTriple platform. It focuses on Social Sciences and Humanities (SSH) content, addressing their limited presence in standard LLM pre-training corpora.
Current release includes History, Sociology, Environmental Sciences, Psychology and Geography texts (~23.14B tokens).
Intended Use
Continuous… See the full description on the dataset page: https://huggingface.co/datasets/odoma/gotriple-pretraining-dataset.pretraining-priors-political-partisan
pretraining-priors political partisan answers
For every one of the 6,820 IdeoINST questions (Chen et al., EMNLP 2024,
arXiv:2402.11725; the prompt set Eugleo/pretraining-priors-political-eval), one claude-sonnet-5 call wrote a very
strongly and clearly left-leaning answer and a very strongly and clearly right-leaning
answer (about three sentences each, first person, no partisan self-label), and tagged the
question with one 0/1 flag per domain. A flag cat_<domain> is 1 when a… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-political-partisan.ORD-Pretraining-v5GenoJEPA-Pretraining
GenoJEPA-Pretraining
This dataset provides the pre-training resources used for GenoJEPA, a genomic representation learning framework based on joint-embedding predictive architecture.
GenoJEPA learns semantic representations of DNA sequences by shifting the optimization target from nucleotide-level reconstruction to latent-space semantic alignment. The pre-training data is used to construct global and local sequence views for self-supervised genomic representation learning.… See the full description on the dataset page: https://huggingface.co/datasets/ChengsenWang/GenoJEPA-Pretraining.ai-theories-corpus-en-pretraining
ai-theories en コーパス
ai-theories(https://github.com/kojikojiprg/ai-theories)プロジェクトの成果物。
006(小型 GPT の事前学習)・008 の英語条件用のコーパス。
ライセンスについての注記
このデータセット自体のライセンスは クリエイティブ・コモンズ 表示-継承 4.0 国際
(CC BY-SA 4.0) である。ai-theoriesの他の成果物(トークナイザなど)は
cc-by-nc-4.0を採用しているが、本データセットの内容はフリー百科事典
『ウィキペディア(Wikipedia)』の本文そのもの(wikitext を平文に変換したのみで、
内容は改変していない)であり、Wikipedia 本文自体のライセンス(CC BY-SA 4.0、
表示・継承の条件)を継承する必要があるため、別のライセンスとしている。
由来… See the full description on the dataset page: https://huggingface.co/datasets/kojikojiprg/ai-theories-corpus-en-pretraining.
