BabyLM-community/babylm-bug
BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: bug Script: Latn Tier: 1M Byte Premium Factor: 1.227860 Size (MB): 6.67 Expected Size (MB): 6.67 Number of Documents: 7,777 Total Tokens: 1,002,579 Tokenizer: separate by whitespace Tokens Per Category child-books: 41,174 tokens padding: 961,405 tokens… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-bug.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face