t5
Datasets
All datasets matching “t5”c4_t5_corrupted_seqlen256
Dataset Card for "c4_t5_corrupted_seqlen256"
More Information needed
openwebtext-t5Mintaka_Graph_Features_T5-xl-ssm
Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm"
More Information needed
c4-t5-ragged
C4, T5 tokenized, in ragged array format
Processed distribution of Google's C4 dataset: a colossal, cleaned version of Common Crawl's web crawl corpus.
Uses the text data from allenai/c4.
Includes en subset only.
T5 tokenizer was applied to the text.Distributed as a ragged array.
Converted via json_to_ragged.py.
Download size of all shards:
Split
Data+Lengths Size
Divided across n Shards
Typical shard size: data.npy
Typical shard size: len.npy
Train
293G
1024
344M
1.4M… See the full description on the dataset page: https://huggingface.co/datasets/Birchlabs/c4-t5-ragged.xsum_validation_t5capstone_sakuga_iblip_t5_embeddings
