toksuitebackup/byt5-small-toksuite-detokenized
Training data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
0158
Upload folder using huggingface_hub
Upload README.md with huggingface_hub
initial commit
