CoolFace
Datasetpublic

semran1/Nemotron-Pretraining-Specialized-v1

Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web… See the full description on the dataset page: https://huggingface.co/datasets/semran1/Nemotron-Pretraining-Specialized-v1.

sourceHugging Faceotherupdated 10mo agoView on Hugging Face
2likes63downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
semran1/Nemotron-Pretraining-Specialized-v1 · CoolFace