ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2
ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.
ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2)
Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI).
v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its non-standard license — download it directly from openbmb/factnet_factsense if you want to replicate exactly.
Source mix (final weights used for base pretrain)
* = external source (openbmb/factnet_factsense) referenced in training but not mirrored here due to license.
Category breakdown
Reasoning-level tags
The OpenThoughts subset is preprocessed to include [REASONING: LEVEL] tags before each assistant response, where LEVEL is one of LOW, MEDIUM, HIGH, XHIGH, MAX based on response length. This teaches the model to associate the tag with reasoning depth for controllable thinking at inference.
Mapping (assistant response chars):
- LOW: < 2000
- MEDIUM: 2000 - 5000
- HIGH: 5000 - 10000
- XHIGH: 10000 - 20000
- MAX: > 20000
Usage
from datasets import load_dataset
ds = load_dataset("ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2", split="train", streaming=True)
for row in ds:
text = row.get("text") or row.get("content") or ""
...Different sources have different columns:
Training recipe
- Model: 1B total / 47M active MoE (top-1 over 35 experts), hidden 384, 16 layers, GQA 6:2, 64k vocab, 5120 seq
- Base LR: 2e-5 → 1e-6 (WSD linear decay over 7.5 B token budget)
- Rolling saves: every 100M tokens
- Warm-start init: Pluto Nano 0.5 (32k) → Pluto Nano 1.0 (64k) via Qwen2.5-1.5B sub-token decomposition (100% non-random)
See training script: `scripts/113_pretrain_streaming.py`
Attribution
All source datasets remain under their original licenses. This bundle is redistributed unchanged with proper attribution:
HuggingFaceFW/fineweb-edu&fineweb-2— © HuggingFace, ODC-BY 1.0openbmb/UltraData-Math— © OpenBMB, Apache 2.0mlabonne/open-perfectblend— © Maxime Labonneopen-thoughts/OpenThoughts3-1.2M— © open-thoughts team, Apache 2.0 (annotated with QwQ-32B)crumb/openstax-text— © crumb (OpenStax textbooks), CC-BY-4.0allenai/peS2o— © Allen AI, ODC-BYopenbmb/factnet_factsense— © OpenBMB (external, not mirrored)
Related
- Pluto Nano 1.0 model: `ASTRAI-labs/pluto-nano-1.0` (coming soon)
- Predecessor v1: `ASTRAI-labs/Pluto-Nano-1.0-Pretrain`
- Predecessor model: `ASTRAI-labs/pluto-nano-0.5`
