lillian039/nemotron_cc_v2_hq_packed4096_200shard
Nemotron-CC-v2 High-Quality, packed to 4096 tokens (train) Documents from nvidia/Nemotron-CC-v2 High-Quality subset, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS appended per document) and greedily packed into sequences of at most 4096 tokens. A document is never split across a pack boundary; documents longer than 4096 are truncated to their own pack. Every pack ends on an EOS/document boundary. Schema index (int64): running pack id… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096_200shard.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face