TemryL/tokenized_wikipedia_20220301.en_train_512
Tokenized English Wikipedia Dataset Dataset Description This dataset contains tokenized chunks of text from the English Wikipedia dump of March 1, 2022. Each entry in the dataset represents a chunk of text from Wikipedia, with information about which document and position within the document it comes from. Dataset Creation Source Dataset: Wikipedia (20220301.en) Tokenizer: BERT base uncased Chunk Size: 512 tokens (including special tokens)… See the full description on the dataset page: https://huggingface.co/datasets/TemryL/tokenized_wikipedia_20220301.en_train_512.
This repository belongs to TemryL on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
