CoolFace
Datasetpublicgated

BabyLM-community/babylm-ces

BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: ces Script: Latn Tier: 1M Byte Premium Factor: 1.035849 Size (MB): 5.64 Expected Size (MB): 5.62 Number of Documents: 540 Total Tokens: 762,576 Tokenizer: separate by whitespace Tokens Per Category child-directed-speech: 377,313 tokens padding-fineweb-c: 78,540… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ces.

sourceHugging Faceunknownupdated 1y agoView on Hugging Face
0likes5downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.