4factors/arabic-msa-sample
Arabic — Modern Standard Arabic (MSA) Sample Native-written, human-verified Modern Standard Arabic. No scraping. No machine translation. No synthetic generation. Every sentence written from scratch by a first-language speaker in formal news / official-statement register, then reviewed line by line against a written checklist and measured for structural diversity across the whole set. A public demonstration sample (50 items). Larger MSA datasets and other varieties (Levantine… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-msa-sample.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face