BabyLM-community/babylm-zho
BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: zho Script: Hani, Hans, Latn Tier: > 100M Byte Premium Factor: 0.935966 Size (MB): 518.85 Expected Size (MB): 508.23 Number of Documents: 203,891 Total Tokens: 137,835,046 Tokenizer: Qwen/Qwen3-0.6B Tokens Per Category child-available-speech: 7,403,441 tokens… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-zho.
145
No card is published for this repository, or it could not be fetched from Hugging Face right now.
