datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KapInstruct-100M
KapInstruct-100M: Curated 100-Million Token Instruction Tuning Dataset
KapInstruct-100M is a high-fidelity, 100-million-token instruction-tuning dataset engineered for Supervised Fine-Tuning (SFT) and alignment of compact language models (under 1 billion parameters). Formatted with the Qwen ChatML chat template and tokenized using Qwen/Qwen3.5-0.8B-Base, the dataset enforces strict assistant-only loss masking (masking user prompts and structural delimiters to -100)… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapInstruct-100M.lucky-initialization-atlas-100m-v2
Lucky initialization atlas v2 evidence
Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs,
provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses
after those artifacts complete. It excludes
credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B
binary logs.
WestGenesis-Coder-SFT-100M
Dataset Overview
WestGenesis-Coder-Dataset is a meticulously curated coding dataset designed specifically for instruction-based model tuning and fine-tuning of existing models with enhanced code generation capabilities. This represents one of the largest and most comprehensively filtered corpora of publicly available coding data on the Hugging Face platform, with a non-thinking approach that emphasizes direct, concise code outputs for rapid model training.
Key… See the full description on the dataset page: https://huggingface.co/datasets/isthatshan/WestGenesis-Coder-SFT-100M.Rainbow-Pony-100m-Flutter-steps-eval
Rainbow-Pony-100M Flutter — Steps Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-steps,
a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files.
In steps mode, the model is given an existing file and an edit instruction and
generates a sequence of localized search/replace edit actions, each mechanically
applied to the current file state before the next action is generated… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-steps-eval.Rainbow-Pony-100m-Flutter-direct-eval
Rainbow-Pony-100M Flutter — Direct Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-direct,
a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files.
In direct mode, the model is given an existing file and an edit instruction and
generates the complete modified file in a single forward pass (as opposed to the
steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-direct-eval.bulgarian-medical-cpt-100m
Bulgarian text for MOSS continued pretraining
Exactly 100 million training tokens: 10M medical and 90M general Bulgarian.
An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers.
No model training has been performed as part of this dataset build.
Medical data
Exactly 10,000,000 training tokens and 50,000 additional validation tokens,
including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.simple-100m-pretrain-1b
Simple-100M Pretraining Dataset (1B Tokens)
A training-optimized, packed pretraining dataset for ~100M parameter language models. Built for reproducibility, minimal runtime overhead, and exact mixing ratios.
🎯 Purpose
This dataset was created to train Simple-100M, a decoder-only Transformer targeting:
✅ Beat GPT-2-70M perplexity with minimal complexity
✅ Reproducible artifacts with exact token accounting
✅ Zero runtime preprocessing (ready-to-train)
Target… See the full description on the dataset page: https://huggingface.co/datasets/Dodosoomro/simple-100m-pretrain-1b.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.sutra-100M
Sutra 100M Pretraining Dataset
A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens.
Dataset Description
This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through:
Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.RedPajama-Data-V2-100M
RedPajama-Data-V2-100M
Dataset Description
This is a 100.0 Million token subset of krisbailey/RedPajama-Data-V2-1B, which is a subset of togethercomputer/RedPajama-Data-V2.
Motivation
100M tokens is a standard size for:
CI/CD Pipelines: Fast enough to download and train for unit tests.
Debugging: Verifying training loops without waiting for hours.
Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B).
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-100M.swerebench-openhands-100m-max64k
SWE-rebench OpenHands 100M SFT Subset (max 64k)
This is a deterministic, representative subset of
nebius/SWE-rebench-openhands-trajectories,
augmented with exact sequence and supervised-loss token counts. The source trajectories were
collected with Qwen3-Coder-480B-A35B-Instruct and OpenHands v0.54.0. This derivative preserves
the source dataset's CC BY 4.0 license and attribution.
Filters and size
7,867 trajectories
100,095,655 assistant loss tokens
352,709,237… See the full description on the dataset page: https://huggingface.co/datasets/hanspeterlyngsoeraaschoujensen/swerebench-openhands-100m-max64k.MS-GPT-Pretraining-100M
MS-GPT 100M pretraining corpus
This repository contains the molecule-only pretraining corpus released with
MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced
Posterior Querying of a Molecule-Language Model.
The released metadata reports 100,007,359 molecule records, 4,096 fingerprint
bits, and 9,413 excluded InChIKeys. The corpus is accompanied by formula vectors
and formula-group metadata.
File
Description
data.arrow
Molecule pretraining… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/MS-GPT-Pretraining-100M.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think
4,000 evaluation attempts: 250 tasks × 16 samples. Original samples 0–3 and
twelve additional seeded samples 4–15. The train split contains evaluation
records, not training data. The evaluated model is violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m at
revision 0c295885100d6eba4f514752aa081c5b0c73fdec.
Cohort
Attempts
All-criteria-pass rate ± task-level SEM
Original four
1,000
8.000% ±… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think.sutra-improved-100M
Sutra Improved 100M
A self-improved pedagogical dataset for LLM pretraining, containing 413,899 entries totaling 110,038,011 tokens (~110 million). This dataset was created by applying an iterative self-improvement process to the Sutra-10B dataset, where each sample was rewritten using Gemma-3-4B-IT and only the better version (original or rewritten) was kept, followed by comprehensive deduplication and quality filtering.
Dataset Description
This dataset explores… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-improved-100M.cosmopedia-100M
cosmopedia-100M
Dataset Description
This is a 100.0 Million token subset of krisbailey/cosmopedia-1B, which is a subset of HuggingFaceTB/cosmopedia.
Motivation
100M tokens is a standard size for:
CI/CD Pipelines: Fast enough to download and train for unit tests.
Debugging: Verifying training loops without waiting for hours.
Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B).
Dataset Details
Total Tokens: 100,000,060… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-100M.clt-tokenized-control-ar-zh-ko-ja-100m
Balanced Arabic-Chinese-Korean-Japanese 4B-token training data
Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 100,000,000 tokens in complete documents. Total target: 400,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard.
Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>.
Vertex-0.6-100M-self-identification
Vertex 0.6 100M Self Identification
A self-identification SFT dataset for Vertex-0.6-100M-8192-Instruct:
459 ChatML-style conversations that teach the model who it is: its name, creator, family,
architecture, parameter count and knowledge cutoff.
Made from SupraLabs/LLM-self-identification
(Apache-2.0), with every {{SELF_ID.*}} marker replaced with the Vertex 0.6 100M identity:
Marker
Value
MODEL_ID
VertexResearch/Vertex-0.6-100M-8192-Instruct
MODEL_NAME
Vertex 0.6… See the full description on the dataset page: https://huggingface.co/datasets/VertexResearch/Vertex-0.6-100M-self-identification.dclm_data_100m
Data-Constrained Language Model Pretraining (100M)
This dataset contains pre-tokenized .pt files containing packed GPT-2-tokenized sequences. It was used in the research presented in the paper Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws.
The repository provides the 100M token training split along with a validation split, specifically prepared for experiments in data-constrained, compute-rich regimes.
Links
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/zhiwei555/dclm_data_100m.KYS-DCLM-Refinedweb-100M-Scored
KYS-DCLM-Refinedweb-100M-Scored
The candidate document pool behind Know Your Sources: Data Selection Matters when Rewriting for
Data-Constrained Pretraining — 99,949,162 web documents sampled from
DCLM-RefinedWeb, each annotated
with three independent quality scores, their tie-aware global percentiles, a 24-way
WebOrganizer topic label, and a Llama-2 token count.
Every source-selection strategy in the paper is a different way of ranking this table.
Contents
200… See the full description on the dataset page: https://huggingface.co/datasets/blab-jhu/KYS-DCLM-Refinedweb-100M-Scored.falcon-refinedweb-100M
falcon-refinedweb-100M
Dataset Description
This is a 100.0 Million token subset of krisbailey/falcon-refinedweb-1B, which is a subset of tiiuae/falcon-refinedweb.
Motivation
100M tokens is a standard size for:
CI/CD Pipelines: Fast enough to download and train for unit tests.
Debugging: Verifying training loops without waiting for hours.
Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B).
Dataset Details
Total Tokens:… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-100M.
