CoolFace
Datasetpublic

lillian039/nemotron_cc_v2_hq_packed4096_200shard

Nemotron-CC-v2 High-Quality, packed to 4096 tokens (train) Documents from nvidia/Nemotron-CC-v2 High-Quality subset, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS appended per document) and greedily packed into sequences of at most 4096 tokens. A document is never split across a pack boundary; documents longer than 4096 are truncated to their own pack. Every pack ends on an EOS/document boundary. Schema index (int64): running pack id… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096_200shard.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes212downloads
settings

This repository belongs to lillian039 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenemotron_cc_v2_hq_packed4096_200shard
visibilitypublic
licenceother
gatedno
ownerlillian039
Account settings
lillian039/nemotron_cc_v2_hq_packed4096_200shard · CoolFace