minhnguyent546/Fineweb-Edu-10BT
Dataset Summary This is the tokenized Fineweb-Edu (10BT subset) using SmolLM2-135M tokenzier. Data is divided into shards (.npy files) for easier to load with PyTorch IterableDataset. Each .npy file can be loaded with numpy.load('file_name.npy'). Split # Documents # Shards # Tokens train 9,575,380 101 10,004,991,326 (10.0B) val 96,721 1 101,807,253 (0.1B) Total 9,672,101 102 10,106,798,579 (10.1B) Example of usage uvx hf download… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/Fineweb-Edu-10BT.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face