CoolFace
Datasetpublic

TemryL/tokenized_wikipedia_20220301.en_train_512

Tokenized English Wikipedia Dataset Dataset Description This dataset contains tokenized chunks of text from the English Wikipedia dump of March 1, 2022. Each entry in the dataset represents a chunk of text from Wikipedia, with information about which document and position within the document it comes from. Dataset Creation Source Dataset: Wikipedia (20220301.en) Tokenizer: BERT base uncased Chunk Size: 512 tokens (including special tokens)… See the full description on the dataset page: https://huggingface.co/datasets/TemryL/tokenized_wikipedia_20220301.en_train_512.

sourceHugging Facecc-by-sa-3.0updated 2y agoView on Hugging Face
0likes343downloads
settings

This repository belongs to TemryL on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nametokenized_wikipedia_20220301.en_train_512
visibilitypublic
licencecc-by-sa-3.0
gatedno
ownerTemryL
Account settings
TemryL/tokenized_wikipedia_20220301.en_train_512 · CoolFace