CoolFace
Datasetpublic

eoinf/wikitext_gptneox

Dataset Card for eoinf/wikitext_gptneox Original dataset Original dataset: Salesforce/wikitext Dataset Details Total Tokens: 122,236,928 Total Sequences: 119,372 Context Length: 1024 tokens Tokenizer: EleutherAI/gpt-neox-20b Format: Each example contains a single field tokens with a list of 1024 token IDs Preprocessing Each document was: Tokenized using the EleutherAI/gpt-neox-20b tokenizer Prefixed with a BOS (beginning of… See the full description on the dataset page: https://huggingface.co/datasets/eoinf/wikitext_gptneox.

sourceHugging Facemitupdated 11mo agoView on Hugging Face
0likes5downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
eoinf/wikitext_gptneox · CoolFace