CoolFace
Datasetpublic

lillian039/nemotron_cc_v2_hq_packed4096_200shard

Nemotron-CC-v2 High-Quality, packed to 4096 tokens (train) Documents from nvidia/Nemotron-CC-v2 High-Quality subset, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS appended per document) and greedily packed into sequences of at most 4096 tokens. A document is never split across a pack boundary; documents longer than 4096 are truncated to their own pack. Every pack ends on an EOS/document boundary. Schema index (int64): running pack id… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096_200shard.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes212downloads
7 commits on main
21631e82mo ago

Add files using upload-large-folder tool

lillian039
8b67a522mo ago

Add files using upload-large-folder tool

lillian039
e20190f2mo ago

Add files using upload-large-folder tool

lillian039
dcb813c2mo ago

Add files using upload-large-folder tool

lillian039
55d2df62mo ago

Add files using upload-large-folder tool

lillian039
201a3462mo ago

Upload README.md with huggingface_hub

lillian039
4987f602mo ago

initial commit

lillian039