CoolFace
Datasetpublic

toksuitebackup/aya-expanse-8b-toksuite-detokenized

Training data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes697downloads
7 commits on main
5a3b5d610mo ago

Upload folder using huggingface_hub

gsaltintas
4228d3111mo ago

Upload README.md with huggingface_hub

gsaltintas
ab17d3b11mo ago

Upload README.md with huggingface_hub

gsaltintas
526d71711mo ago

Upload README.md with huggingface_hub

gsaltintas
d6af31911mo ago

Upload README.md with huggingface_hub

gsaltintas
d33f98111mo ago

Upload folder using huggingface_hub

gsaltintas
7ea754f11mo ago

initial commit

gsaltintas