CoolFace
Datasetpublic

ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2

ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
2likes549downloads
13 commits on main
34ba8573mo ago

Upload folder using huggingface_hub

ShinMK3
81edb923mo ago

Upload folder using huggingface_hub

ShinMK3
a14e3423mo ago

Upload folder using huggingface_hub

ShinMK3
733df053mo ago

Upload folder using huggingface_hub

ShinMK3
175e4893mo ago

Upload folder using huggingface_hub

ShinMK3
5f84a7e3mo ago

Upload folder using huggingface_hub

ShinMK3
fcbb64a3mo ago

Upload folder using huggingface_hub

ShinMK3
641dba83mo ago

Upload folder using huggingface_hub

ShinMK3
c93ec763mo ago

Upload folder using huggingface_hub

ShinMK3
7b2272e3mo ago

Upload folder using huggingface_hub

ShinMK3
2ad138f3mo ago

Upload scripts/113_pretrain_streaming.py with huggingface_hub

ShinMK3
4d33cef3mo ago

Upload README.md with huggingface_hub

ShinMK3
aeff4cf3mo ago

initial commit

ShinMK3