CoolFace
Datasetpublic

AverageMetaheuristicsEnjoyer/fineweb-edu-100BT-gpt2-bin

fineweb-edu 100BT — GPT-2 pre-tokenized (.bin) Pre-tokenized HuggingFaceFW/fineweb-edu :: sample/100BT for LLM pretraining without on-the-fly tokenization or HF streaming (flat uint16 token ids, nanoGPT layout). Tokenizer: gpt2 (tiktoken == HF AutoTokenizer('gpt2'), ids identical) Format: uint16 little-endian, headerless Layout: documents concatenated, eos=50256 appended after each doc eos token id: 50256 train tokens: 100,146,465,071 val tokens: 20,000,000 Usage… See the full description on the dataset page: https://huggingface.co/datasets/AverageMetaheuristicsEnjoyer/fineweb-edu-100BT-gpt2-bin.

sourceHugging Faceodc-byupdated 2mo agoView on Hugging Face
0likes11downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
AverageMetaheuristicsEnjoyer/fineweb-edu-100BT-gpt2-bin · CoolFace