CoolFace
Datasetpublicgated

BabyLM-community/babylm-zho

BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: zho Script: Hani, Hans, Latn Tier: > 100M Byte Premium Factor: 0.935966 Size (MB): 518.85 Expected Size (MB): 508.23 Number of Documents: 203,891 Total Tokens: 137,835,046 Tokenizer: Qwen/Qwen3-0.6B Tokens Per Category child-available-speech: 7,403,441 tokens… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-zho.

sourceHugging Faceunknownupdated 8mo agoView on Hugging Face
1likes35downloads

BabyLM-community/babylm-zho · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.