BabyLM-community/babylm-fra
BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: fra Script: Latn Tier: 100M Byte Premium Factor: 1.173979 Size (MB): 634.88 Expected Size (MB): 637.47 Number of Documents: 81,950 Total Tokens: 126,580,785 Tokenizer: separate by whitespace Tokens Per Category child-available-speech: 1,989,852 tokens child-books:… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-fra.
017
No card is published for this repository, or it could not be fetched from Hugging Face right now.
