datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
carbon-pretraining-corpus
🧬 Carbon Pretraining Corpus
Description
173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model.
This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species.
Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.pretraining-corpus
🧠 Kiy-K Synthetic Pretraining Corpus
Author: Khoi K. (@Kiy-K)License: Apache 2.0Last Updated: 2025-10-30
📘 Overview
The Kiy-K Synthetic Pretraining Corpus is a large-scale collection of synthetically generated English text designed for language model pretraining and instruction-tuning research.
All data is synthetic, created using open-source large language models such as GPT-OSS, NVIDIA Nemotron, and DeepSeek, under full control of the author.No real user… See the full description on the dataset page: https://huggingface.co/datasets/Kiy-K/pretraining-corpus.tiny-slm-pretraining-corpus
🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB)
A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures.
100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders.
📊 Dataset Statistics
Total Documents: 20,066,075
Train: 19,663,898
Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.pretraining-high-quality
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness of the text
lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.pretraining-lower-quality
Dataset Card for Lapa Pretraining Lower Quality Dataset
Dataset Description
Dataset Summary
This dataset is a high quality (but lower quality than https://huggingface.co/datasets/lapa-llm/pretraining-high-qualit) subset of pretraining corpus for Ukrainian language.
It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality.synthetic-pretraining-transformers-v1
Synthetic Pre-training Transformers v1.0.0
Dataset Description
This is a synthetic pre-training dataset generated from transformer architecture patterns. It contains paraphrased, augmented, and interpolated content derived from validated seed data about neural sequence modeling and attention mechanisms.
Dataset Summary
Total Samples: 100
Total Tokens: 6,084
Average Tokens per Sample: 60.84
Format: Parquet
Version: 1.0.0
License: CC-BY-4.0
Supported… See the full description on the dataset page: https://huggingface.co/datasets/gugarosa/synthetic-pretraining-transformers-v1.KYS-1.5B-Pretraining-Corpora
KYS-1.5B-Pretraining-Corpora
The six 10B-token pretraining mixtures from Know Your Sources: Data Selection Matters when
Rewriting for Data-Constrained Pretraining, stored without duplication.
The key idea: one anchor, six remainders
Every setting trains on the same 10B-token recipe:
10B mixture = 5B shared anchor + 5B strategy-specific tokens
(identical in all (this is the ONLY thing
six settings, that differs… See the full description on the dataset page: https://huggingface.co/datasets/blab-jhu/KYS-1.5B-Pretraining-Corpora.pretraining-high-quality-10k-workshop
Lapa HQ 10k Workshop Corpus
A small deterministic subset of lapa-llm/pretraining-high-quality for tokenizer-transfer workshop runs.
Provenance
Source dataset: lapa-llm/pretraining-high-quality
Source config: default
Source split: train
Rows: 10000
Selection: first 10000 rows by dataset-server row order
Download window size: 100
Parallel workers: 20
Created at UTC: 2026-06-20T09:22:21.460168+00:00
Added columns:
source_row_idx
mini_corpus_index
glossapi-greek-nanochat-pretraining-dataset
Glossapi Greek Nanochat Pretraining Dataset
This repository contains the source-separated Greek corpus used to build nanochat Greek pretraining mixtures. It is intentionally not a pre-split train/validation/test export: builders load data/*.parquet, preserve source_dataset, and create deterministic experiment-specific mixes and splits downstream.
Current Snapshot
Total rows: 49474947
Total characters: 248276390721
Included source datasets: 19
Data files: 273… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset.
