datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
byt5-small-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
byT5_base_zero_shot_eval_result_1000_summary_promptbyT5_base_zero_shot_eval_result_200_no_promptbyT5_base_zero_shot_eval_result_200_summary_promptbyT5_base_zero_shot_eval_result_500_no_promptbyT5_base_zero_shot_eval_result_1000_no_promptbyT5_base_zero_shot_eval_result_500_summary_promptdhivehi-byt5-data
