CoolFace
Datasetpublicgated

BabyLM-community/babylm-eng

BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: eng Script: Latn Tier: 100M Byte Premium Factor: 1.000000 Size (MB): 539.18 Expected Size (MB): 543.00 Number of Documents: 137,710 Total Tokens: 98,878,321 Tokenizer: separate by whitespace Tokens Per Category child-available-speech: 9,102,166 tokens child-books:… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-eng.

sourceHugging Faceunknownupdated 1y agoView on Hugging Face
3likes37downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.