CoolFace
Datasetpublic

toksuitebackup/meta-llama-Llama-3.2-1B-toksuite-detokenized

Training data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes988downloads
7 commits on main
8d3bf4410mo ago

Upload folder using huggingface_hub

gsaltintas
066d4cf11mo ago

Upload folder using huggingface_hub

gsaltintas
d03b77011mo ago

Upload README.md with huggingface_hub

gsaltintas
c4cab4111mo ago

Upload README.md with huggingface_hub

gsaltintas
70dd0c611mo ago

Upload README.md with huggingface_hub

gsaltintas
f36f72c11mo ago

Upload README.md with huggingface_hub

gsaltintas
cbe966b11mo ago

initial commit

gsaltintas