CoolFace
Datasetpublic

ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2

ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
2likes549downloads
Dataset Card

ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2)

Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI).

v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its non-standard license — download it directly from openbmb/factnet_factsense if you want to replicate exactly.

Source mix (final weights used for base pretrain)

WeightCategorySourceLicenseSubdir
10 %EN WebHuggingFaceFW/fineweb-eduODC-BY 1.0en_fineweb_edu/
15 %Mathopenbmb/UltraData-Math (L3-Conv)Apache 2.0math/
8 %Chat+Codemlabonne/open-perfectblendApache 2.0perfectblend/
12 %CoT Reasonopen-thoughts/OpenThoughts3-1.2MApache 2.0openthoughts/
4 %Textbookscrumb/openstax-textCC-BY-4.0science_openstax/
8 %Papersallenai/peS2o (subset ~5 GB)ODC-BYscience_pes2o/
14 %*Factsopenbmb/factnet_factsense (external — see below)othernot in repo
12 %ZH WebHuggingFaceFW/fineweb-2 (cmn_Hani)ODC-BY 1.0zh_fineweb2/
5 %PT WebHuggingFaceFW/fineweb-2 (por_Latn)ODC-BY 1.0fineweb2_pt/
5 %ES WebHuggingFaceFW/fineweb-2 (spa_Latn)ODC-BY 1.0fineweb2_es/
7 %HI WebHuggingFaceFW/fineweb-2 (hin_Deva)ODC-BY 1.0fineweb2_hi/

* = external source (openbmb/factnet_factsense) referenced in training but not mirrored here due to license.

Category breakdown

CategoryShare
Language (web)39 %
Math15 %
CoT reasoning12 %
Science12 %
Facts (external)14 %*
Chat + Code8 %

Reasoning-level tags

The OpenThoughts subset is preprocessed to include [REASONING: LEVEL] tags before each assistant response, where LEVEL is one of LOW, MEDIUM, HIGH, XHIGH, MAX based on response length. This teaches the model to associate the tag with reasoning depth for controllable thinking at inference.

Mapping (assistant response chars):

  • —LOW: < 2000
  • —MEDIUM: 2000 - 5000
  • —HIGH: 5000 - 10000
  • —XHIGH: 10000 - 20000
  • —MAX: > 20000

Usage

python
from datasets import load_dataset
ds = load_dataset("ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2", split="train", streaming=True)
for row in ds:
    text = row.get("text") or row.get("content") or ""
    ...

Different sources have different columns:

SubdirText field
en_fineweb_edu/text
zh_fineweb2/text
fineweb2_*/text
math/content
perfectblend/conversations
openthoughts/conversations
science_openstax/text
science_pes2o/text

Training recipe

  • —Model: 1B total / 47M active MoE (top-1 over 35 experts), hidden 384, 16 layers, GQA 6:2, 64k vocab, 5120 seq
  • —Base LR: 2e-5 → 1e-6 (WSD linear decay over 7.5 B token budget)
  • —Rolling saves: every 100M tokens
  • —Warm-start init: Pluto Nano 0.5 (32k) → Pluto Nano 1.0 (64k) via Qwen2.5-1.5B sub-token decomposition (100% non-random)

See training script: `scripts/113_pretrain_streaming.py`

Attribution

All source datasets remain under their original licenses. This bundle is redistributed unchanged with proper attribution:

  • —HuggingFaceFW/fineweb-edu & fineweb-2 — © HuggingFace, ODC-BY 1.0
  • —openbmb/UltraData-Math — © OpenBMB, Apache 2.0
  • —mlabonne/open-perfectblend — © Maxime Labonne
  • —open-thoughts/OpenThoughts3-1.2M — © open-thoughts team, Apache 2.0 (annotated with QwQ-32B)
  • —crumb/openstax-text — © crumb (OpenStax textbooks), CC-BY-4.0
  • —allenai/peS2o — © Allen AI, ODC-BY
  • —openbmb/factnet_factsense — © OpenBMB (external, not mirrored)

Related