CoolFace
Datasetpublicgated

nvidia/Nemotron-CC-v2.1

Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1.

sourceHugging Faceotherupdated 9mo agoView on Hugging Face
139likes23kdownloads

No commit history came back for main. The revision may not exist, or the source declined the request.