CoolFace
Datasetpublic

TemryL/tokenized_wikipedia_20220301.en_train_512

Tokenized English Wikipedia Dataset Dataset Description This dataset contains tokenized chunks of text from the English Wikipedia dump of March 1, 2022. Each entry in the dataset represents a chunk of text from Wikipedia, with information about which document and position within the document it comes from. Dataset Creation Source Dataset: Wikipedia (20220301.en) Tokenizer: BERT base uncased Chunk Size: 512 tokens (including special tokens)… See the full description on the dataset page: https://huggingface.co/datasets/TemryL/tokenized_wikipedia_20220301.en_train_512.

sourceHugging Facecc-by-sa-3.0updated 2y agoView on Hugging Face
0likes343downloads

TemryL/tokenized_wikipedia_20220301.en_train_512 · main · files are served by the source, never re-hosted here