eoinf/wikitext_gptneox
Dataset Card for eoinf/wikitext_gptneox Original dataset Original dataset: Salesforce/wikitext Dataset Details Total Tokens: 122,236,928 Total Sequences: 119,372 Context Length: 1024 tokens Tokenizer: EleutherAI/gpt-neox-20b Format: Each example contains a single field tokens with a list of 1024 token IDs Preprocessing Each document was: Tokenized using the EleutherAI/gpt-neox-20b tokenizer Prefixed with a BOS (beginning of… See the full description on the dataset page: https://huggingface.co/datasets/eoinf/wikitext_gptneox.
05
