CoolFace
Datasetpublic

toksuitebackup/Qwen-Qwen3-8B-toksuite-detokenized

Training data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes203downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
toksuitebackup/Qwen-Qwen3-8B-toksuite-detokenized · CoolFace