datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm1bA benchmark corpus to be used for measuring progress in statistical language modeling. This has almost one billion words in the training data.Lumina-Math-Foundations-1B
Lumina-Math-Foundations-1B
Lumina-Math-Foundations-1B is an industrial-scale foundational mathematical reasoning dataset in Indonesian,
Dataset Summary
Language: Indonesian (id) with LaTeX mathematical formulas.
Scale: 1 Billion Synthetic High-Fidelity Mathematical Reasoning instances.---
Data Schema & Field Breakdown
Field Name
Type
Description
problem
string
100% pure human natural language problem statement with standard LaTeX math… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/Lumina-Math-Foundations-1B.sutra-1B
Sutra 1B Pretraining Dataset
A high-quality pedagogical dataset designed for LLM pretraining, containing 948,709 educational entries totaling over 1 billion tokens.
Dataset Description
This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through:
Clear pedagogical structure: Content follows proven educational patterns
Cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-1B.spai-ss6-llm-1b-thai-corpus
Thai Medical And Health Corpus
Thai public medical and health web corpus collected for research and LLM dataset
experimentation, with optional imported Thai medical/health datasets from
Hugging Face stored as separate configs.
Public Web Corpus
Config: default
Split: train
Records: 3660 deduplicated articles
Columns: 16
Format: Parquet
Latest collection profile: free_1000
Latest generated at: 2026-06-06T17:41:38.787978+00:00
Source And Method
The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.shreyansh-1B-SLM-pretrain-stem-english
📚 Vigyan Pretrain Corpus: 13GB Web-Scale Scientific & Technical Text
The Vigyan Pretrain Corpus is a web-scale, curated raw text pre-training dataset comprising 13.14 GB of high-density STEM literature, textbooks, open-access research papers, and technical documentations.
🔬 Dataset Overview
Designed specifically for pre-training and continuous pre-training (CPT) of Small Language Models (SLMs) in the 1B–3B parameter regime:
High Information Density: Filtered to… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-1B-SLM-pretrain-stem-english.human-templated-captions-1bcsv delimiter is = ".,|,."
apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading.
This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon.
Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.RedPajama-Data-V2-1B
RedPajama-Data-V2 1B
Dataset Description
This is a 1.01 Billion token subset of the togethercomputer/RedPajama-Data-V2 dataset (specifically derived from the sample-10B config). It was created by randomly sampling the source data.
Motivation
RedPajama V2 is a state-of-the-art web dataset with rich quality signals. This 1B token subset allows for rapid testing of these quality signals or other filtering experiments without needing to process the full… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-1B.cosmopedia-1b
Cosmopedia 1B
Dataset Description
This is a 1 Billion token subset of the krisbailey/cosmopedia-10B dataset, which itself is a 10B subset of HuggingFaceTB/cosmopedia.
It was created by uniformly sampling approximately 9.5% of the 10B dataset, ensuring the data distribution remains consistent with the source.
Motivation
While the 10B dataset is a "Goldilocks" size for many experiments, 1B tokens is the standard size for rapid prototyping, scaling law… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-1b.SYN-1B
SYN-1B
Dataset Summary
SYN-1B is a 1.04B-token synthetic language-modeling corpus of rule-governed
text streams. Each row is a decoded instance in which the text establishes
facts, mappings, bindings, or simple generative rules; later spans may revise
those rules, swap bindings, delay a query over long filler, or present a null
control with event-like surface text that should not change the answer.
The dataset is intended as structured pretraining data, not as a… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/SYN-1B.LFM2.5-8B-A1B-KO-CPT-DATA
LFM2.5-8B-A1B Korean CPT Data
Prepared Korean continued-pretraining data for LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL.
Files
data/ko_cpt_mix_full_lfmstyle_20260627.jsonl: prepared full CPT corpus with one JSON object per line and a text field
metadata/ko_cpt_mix_full_lfmstyle_20260627.stats.json: corpus statistics
metadata/ko_cpt_mix_full_lfmstyle_20260627.stats.json.full_report.json: per-source preprocessing report
metadata/ko_cpt_sources_full_20260627.json:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-DATA.QuixiMath-1B
QuixiMath-1B
QuixiMath is brought to you by Eric Hartford and QuixiAI
https://github.com/QuixiAI/QuixiMath
Dataset Summary
QuixiMath-1B is a synthetic math reasoning corpus generated from the QuixiMath
procedural problem generators. Each record contains a natural-language problem,
explicit step-by-step scratchpad opcodes, a canonical final answer, and metadata
for filtering or reweighting by skill, operation, grade band, and relative
difficulty.
The canonical… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/QuixiMath-1B.sphragis-olmo1b-adaptation-corpus
Sphragis OLMo-1B adaptation corpus
Version-controlled input for adapting allenai/OLMo-1B-hf to Ancient Greek
before authorship-language-model training. It contains only OGA whole works
whose TLG author occurs in neither Sphragis benchmark.
Text has the exact model-facing benchmark surface form: polytonic-aware
lowercasing with grc_utils.lower_grc, removal of all editorial punctuation,
normalization of whitespace, and removal of consonant-final elision marks.
Splits are made over… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-olmo1b-adaptation-corpus.lm1b
LM1B - One Billion Word Benchmark
Dataset Description
The One Billion Word Benchmark is a large language modeling dataset.
It contains approximately one billion words of training data derived from news articles.
How was this dataset built?
We download the full LM1B dataset from TensorFlow Datasets (TFDS) and convert it to HuggingFace format automatically. The full script is in lm1b.py. The required environment is:
tensorflow==2.20.0
tensorflow-datasets==4.9.9… See the full description on the dataset page: https://huggingface.co/datasets/FrankCCCCC/lm1b.falcon-refinedweb-1B
Falcon RefinedWeb 1B
Dataset Description
This is a 1.01 Billion token subset of the tiiuae/falcon-refinedweb dataset. It was created by streaming the dataset with a large shuffle buffer to ensure a random, representative sample of the web data.
Motivation
RefinedWeb is a high-quality filtered web dataset, but the full version is massive. This 1B token slice provides a perfect testbed for evaluating model architecture changes or for use in curriculum learning… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-1B.simple-100m-pretrain-1b
Simple-100M Pretraining Dataset (1B Tokens)
A training-optimized, packed pretraining dataset for ~100M parameter language models. Built for reproducibility, minimal runtime overhead, and exact mixing ratios.
🎯 Purpose
This dataset was created to train Simple-100M, a decoder-only Transformer targeting:
✅ Beat GPT-2-70M perplexity with minimal complexity
✅ Reproducible artifacts with exact token accounting
✅ Zero runtime preprocessing (ready-to-train)
Target… See the full description on the dataset page: https://huggingface.co/datasets/Dodosoomro/simple-100m-pretrain-1b.1b-smollm-corpus
SmolLM-Corpus — 1B Token Subset
A curated 1-billion-token English pretraining corpus sampled from
HuggingFaceTB/smollm-corpus,
designed for training small language models (~20M parameters).
Dataset Composition
Source
Ratio
Tokens
Documents
FineWeb-Edu (dedup)
87%
~870M
849,577
Cosmopedia v2
13%
~130M
161,889
Total
100%
~1B
1,011,466
Rationale for the Split
The 87/13 ratio mirrors the natural token distribution of the full… See the full description on the dataset page: https://huggingface.co/datasets/ecreeth/1b-smollm-corpus.fineweb-edu-1B
FineWeb-Edu 1B
Dataset Description
FineWeb-Edu 1B is a high-quality, stratified subset of the HuggingFaceFW/fineweb-edu dataset. It contains approximately 1 billion tokens of educational web text, carefully sampled to preserve the original distribution of source data (CommonCrawl dumps).
This dataset provides an accessible, lightweight alternative to the larger FineWeb-Edu subsets (like sample-10BT or sample-100BT) while maintaining the same data diversity and quality… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/fineweb-edu-1B.fineweb-edu-1b
MLBricks FineWeb-Edu 1B
MLBricks curated datasetUpstream: HuggingFaceFW/fineweb-edu @ 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9Source config: sample-10BTSource split: trainLicense: ODC-By 1.0
MLBricks maintains this dataset as a stable Studio preset and does not claim
ownership of the underlying source material.
Edition
Destination: MLBricks/fineweb-edu-1b
Rows: 969,429
GPT-2 tokens: 1,000,000,000
Main Studio column: text
Source-rights notice… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/fineweb-edu-1b.Chinese_qa
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/sdbhud1b/Chinese_qa.GSM8K-Aug-Llama-3.2-1B-Instruct-Correct-CoT
Verified self-generated GSM8K reasoning
64 independently sampled completions are generated per prepared question.
Final answers are checked against the source answer. Among complete, correctly
formatted correct completions whose CoT passes the final-result-statement and
combined length checks, one sample is selected uniformly at random using a
reproducible per-question seed. CoT length does not rank eligible samples.
The final result belongs
in the separate final-answer line of… See the full description on the dataset page: https://huggingface.co/datasets/hanseungwook/GSM8K-Aug-Llama-3.2-1B-Instruct-Correct-CoT.openwebmath-1b
MLBricks OpenWebMath 1B
MLBricks curated datasetUpstream: open-web-math/open-web-math @ fde8ef8de2300f5e778f56261843dab89f230815Source config: defaultSource split: trainLicense: ODC-By 1.0
MLBricks maintains this dataset as a stable Studio preset and does not claim
ownership of the underlying source material.
Edition
Destination: MLBricks/openwebmath-1b
Rows: 469,065
GPT-2 tokens: 1,000,000,000
Main Studio column: text
Source-rights notice… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/openwebmath-1b.ultra-fineweb-tokenized-1B-short
Ultra-FineWeb Tokenized 2B
Status: complete
A tokenized subset streamed from [openbmb/Ultra-FineWeb] using its en split.
The Parquet data has exactly one column: input_ids (list<int32>).
Processing
Source revision: 7ddd4170ce03e0afbd7d9b80d4bc0b8eebf877e4
Tokenizer: Ba2han/TR_CPT1
Tokenizer revision: d32fb40763740cca0a7c04b4549d2129edeafa9a
Source field: content
Filter: score > 0.75
Stored sequence length: 20 to 888 tokens, inclusive
Target: approximately 1,000… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/ultra-fineweb-tokenized-1B-short.fineweb-1BT
FineWeb-1BT: 1 Billion Token Subset
Dataset Description
FineWeb-1BT is a carefully curated 1 billion token subset of the HuggingFaceFW/fineweb dataset, sampled exclusively from the official 10BT FineWeb subset. It was created using true random sampling to ensure unbiased representation across the entire 10BT corpus.
Provenance: This 1BT subset is a uniform sample drawn from the official FineWeb 10BT split (not from the full raw stream), preserving its language/quality… See the full description on the dataset page: https://huggingface.co/datasets/javiersgjavi/fineweb-1BT.sf9212
传奇私服分布式路由与自动化接口索引库 - Batch 005
本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。
📂 区域节点集群子目录 (Spider Pool Indexes)
👉 新开传奇私服迷失版-2026传奇SF 补丁包-2026传奇私服统治 —— 承载源站 🌐 min-sf.0002sf.com
👉 传奇私服发布网推荐排行-传奇私服平台高爆服-2026传奇私服汉化 —— 承载源站 🌐 min-sf.0211sf.cn
👉 传奇私服大全-传奇私服发布网免费-1.76复古传奇私服 —— 承载源站 🌐 min-sf.06sf.cn
👉 传奇SF 之选-经典传奇私服-找私服-手机版传奇世界sf直装包 —— 承载源站 🌐 min-sf.34sf.net
👉 本月传奇私服新服-2026传奇SF 侠客版-2026传奇SF 传奇发布 —— 承载源站 🌐 min-sf.5566sf.com
👉 本月传奇私服任务-元宝好打的传奇私服-本月传奇私服亚服 —— 承载源站 🌐… See the full description on the dataset page: https://huggingface.co/datasets/Neil1b/sf9212.sf9213
传奇私服分布式路由与自动化接口索引库 - Batch 004
本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。
📂 区域节点集群子目录 (Spider Pool Indexes)
👉 耐玩传奇私服新开发布-传奇SF 神器版-1.76金币小极品私服 —— 承载源站 🌐 new-sf.0002sf.com
👉 全新传奇私服星辰版-火爆传奇私服网站-本月传奇私服微变版 —— 承载源站 🌐 new-sf.34sf.net
👉 传奇sf官网-传奇私服发布网变态版-火爆传奇私服-找私服 —— 承载源站 🌐 new-sf.6sifu.com
👉 原汁原味热血传奇私服-传奇sf开服预告今天-复古传奇私服网 —— 承载源站 🌐 new-sf.7xingame.com
👉 传奇SF 门户-传奇私服发布网任务-传奇新开私服 —— 承载源站 🌐 new-sf.8thgames.com
👉 传奇SF 副本-本月传奇私服火爆版-传奇sf封号怎么解 —— 承载源站 🌐… See the full description on the dataset page: https://huggingface.co/datasets/Neil1b/sf9213.minicpm5-1b-quantization-benchmark
openbmb/MiniCPM5-1B 次世代量子化(Quanto FP8 / INT4 vs BNB 4bit)実測ベンチマークレポート
対象モデル: openbmb/MiniCPM5-1B (1.16B parameters, 128k context, LlamaForCausalLM)
検証ハードウェア: NVIDIA GeForce RTX 4070 Ti (12GB GDDR6X, Ada Lovelace, Compute Capability 8.9, 第4世代Tensor Core)
実行環境: Windows / Python 3.13 / PyTorch 2.6.0+cu124 / transformers 4.57.6 / optimum-quanto 0.2.7 / bitsandbytes 0.50.0
検証日: 2026-09-19 12:12:34
1. エグゼクティブサマリー(全体比較)
NVIDIA GeForce RTX 4070 Ti 実機環境において、標準ネイティブ… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/minicpm5-1b-quantization-benchmark.sf9215
传奇私服分布式路由与自动化接口索引库 - Batch 001
本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。
📂 区域节点集群子目录 (Spider Pool Indexes)
👉 2026全新传奇私服开服-精品新开[传奇私服]-[传奇私服]网今日发布 —— 承载源站 🌐 api.718games.com.cn
👉 传奇私服发布网-变态[传奇私服]发布网-[传奇私服]网欧服-2026全新传奇私服开服 —— 承载源站 🌐 api.999sif.com.cn
👉 传奇私服发布网-[传奇私服]稳定发布网-网通低延迟热血[传奇sf]-2026全新传奇私服开服 —— 承载源站 🌐 api.chuangqi-sifu.com.cn
👉 传奇私服发布网-今日[传奇私服]新开服-[传奇私服]榜单-2026全新传奇私服开服 —— 承载源站 🌐 api.chuanqisfu.com.cn
👉 2026全新传奇私服开服-今日[传奇私服]雷霆版-今日新开传奇高爆服 —— 承载源站 🌐… See the full description on the dataset page: https://huggingface.co/datasets/Neil1b/sf9215.cq92106
传奇私服分布式路由与自动化接口索引库 - Batch 003
本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。
📂 区域节点集群子目录 (Spider Pool Indexes)
👉 2026全新传奇私服开服-[传奇私服]找不到黑屏-1.76合击传奇发布网新开服 —— 承载源站 🌐 g1.718games.com.cn
👉 2026全新传奇私服开服-今日[传奇私服]新服榜-本月[传奇私服]下载站 —— 承载源站 🌐 g1.999sif.com.cn
👉 2026全新传奇私服开服-公益[传奇sf]-微变[传奇私服]发布网 —— 承载源站 🌐 g1.chuangqi-sifu.com.cn
👉 2026全新传奇私服开服-今日[传奇私服]新服推荐-绿色最新[传奇私服] —— 承载源站 🌐 g1.chuanqisfu.com.cn
👉 2026全新传奇私服开服-修仙单职业热血[传奇sf]-[传奇私服]开服公告 —— 承载源站 🌐 g1.game-12hf.cn
👉… See the full description on the dataset page: https://huggingface.co/datasets/Neil1b/cq92106.sf9210
传奇私服分布式路由与自动化接口索引库 - Batch 006
本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。
📂 区域节点集群子目录 (Spider Pool Indexes)
👉 传奇私服排行榜-传奇私服发布网火爆发布-找sf哪个好 —— 承载源站 🌐 min-sf.1440game.com
👉 热血传奇sf经典怀旧-传奇私服发布网-嘟嘟传奇热血传奇私服 —— 承载源站 🌐 min-sf.16hsf.com
👉 2026传奇SF 推荐网-刚开私服晚上8点沙捐首区-传奇私服开服教程 —— 承载源站 🌐 min-sf.18183-game.com.cn
👉 2026传奇SF 将军-传奇SF 心得-传奇私服汉化版 —— 承载源站 🌐 min-sf.18183gamesf.com.cn
👉 传奇sf发布网-sf999传奇发布网入口-耐玩传奇私服发布网 —— 承载源站 🌐 min-sf.1satori.com
👉 本月传奇私服私服榜-新开传奇私服主站-传奇SF… See the full description on the dataset page: https://huggingface.co/datasets/Neil1b/sf9210.experimental-pretrain-1b
Dataset Card for Experimental Pretraining Dataset 1B
Dataset Details
Dataset Description
A meticulously curated 1 billion token dataset optimized for experimental pretraining of small language models. This dataset represents a balanced mixture of the highest quality educational content (60%), mathematical reasoning (30%), and Python code (10%), specifically designed for rapid experimentation and research in language model training.
Curated by: Yxanul… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/experimental-pretrain-1b.
