CoolFace
Datasetpublicgated

BabyLM-community/babylm-heb

BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: heb Script: Hebr Tier: 1M Byte Premium Factor: 1.355477 Size (MB): 7.37 Expected Size (MB): 7.36 Number of Documents: 210 Total Tokens: 818,910 Tokenizer: separate by whitespace Tokens Per Category child-directed-speech: 309,854 tokens padding-wikipedia: 509,056… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-heb.

sourceHugging Faceunknownupdated 1y agoView on Hugging Face
0likes20downloads

BabyLM-community/babylm-heb · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.