CoolFace
Datasetpublic

nirmalendu01/abir177m-pretrain-balanced20-ezhijaru

abir177m pretrain mix — balanced20 en/zh/hi/ja/ru Frozen packed-token shards for reproducible abir177m GPT-2–style pretraining. Languages: 20% each en, zh, hi, ja, ru Source streams: FineWeb (en) + FineWeb-2 (zh/hi/ja/ru) Tokenizer: mistralai/Mistral-Nemo-Base-2407 Packing: 2048-token causal LM blocks (input_ids, labels identical) Target budget: 3.55B tokens (1,733k sequences) See meta.json for exact mixture + dataset map + seed.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes128downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face