CoolFace
Datasetpublic

Fredithefish/Nemotron-CC-HQ-20B

Nemotron-CC-HQ-20B This Dataset consists of approximately 20B tokens of Nemotron-CC-HQ, consisting of randomly sampled slices from crawls in the range CC-MAIN-2013-20-part-00012 to CC-MAIN-2019-04-part-00007. For more information about Nemotron-CC check the Paper by Nvidia Disclaimer: Derived from Nemotron-CC (Common Crawl). No ownership of underlying content is claimed. Data may be subject to third-party rights. Use at your own risk and in compliance with… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Nemotron-CC-HQ-20B.

sourceHugging Faceupdated 6mo agoView on Hugging Face
1likes2kdownloads
4 commits on main
1bcdb7f6mo ago

Update README.md

Fredithefish
e5160c36mo ago

Create README.md

Fredithefish
39f57eb6mo ago

Upload folder using huggingface_hub

Fredithefish
65c6c596mo ago

initial commit

Fredithefish