CoolFace
Datasetpublic

toksuitebackup/byt5-small-toksuite-detokenized

Training data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes158downloads

toksuitebackup/byt5-small-toksuite-detokenized · main · files are served by the source, never re-hosted here