BabyLM-community/babylm-ar-subtitles
babylm-ara Dataset Description This dataset is part of the BabyLM multilingual collection. Dataset Summary Language: ara Script: Unknown Number of Documents: 65951 Total Tokens: 399142332 Tokens Per Category subtitles: 399142332 tokens Data Fields text: The document text doc_id: Unique identifier for the document category: Type of content (e.g., child-directed-speech, educational, etc.) data-source: Original source… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ar-subtitles.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face