CoolFace
Datasetpublic

alexkstern/owt-gpt2bpe-9B

owt-gpt2bpe-9B OpenWebText (from apollo-research/Skylion007-openwebtext-tokenizer-gpt2), pre-tokenized with the gpt2bpe tokenizer (vocab 50,257) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 9,015,870,208 val.bin val 20,000,000 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/owt-gpt2bpe-9B.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes24downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
alexkstern/owt-gpt2bpe-9B · CoolFace