BabyLM-community/babylm-fas
BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: fas Script: Arab Tier: 100M Byte Premium Factor: 1.597326 Size (MB): 867.30 Expected Size (MB): 867.35 Number of Documents: 217,776 Total Tokens: 98,506,081 Tokenizer: separate by whitespace Tokens Per Category child-books: 67,165 tokens educational: 94,320,928… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-fas.
021
No card is published for this repository, or it could not be fetched from Hugging Face right now.
