menasaat/oman-wikipedia-corpus
Oman Wikipedia Corpus | مجموعة ويكيبيديا العُمانية A bilingual (Arabic + English) plaintext corpus of Wikipedia articles about Oman, built from the official MediaWiki APIs of ar.wikipedia.org and en.wikipedia.org. مجموعة نصوص ثنائية اللغة (العربية والإنجليزية) من مقالات ويكيبيديا المتعلقة بسلطنة عُمان، مبنية من واجهات ميدياويكي الرسمية. Dataset Summary Arabic (عربي) English Articles 1,205 522 Total words 702,327 400,257 Categories walked 258 147… See the full description on the dataset page: https://huggingface.co/datasets/menasaat/oman-wikipedia-corpus.
Oman Wikipedia Corpus | مجموعة ويكيبيديا العُمانية
A bilingual (Arabic + English) plaintext corpus of Wikipedia articles about Oman, built from the official MediaWiki APIs of ar.wikipedia.org and en.wikipedia.org.
مجموعة نصوص ثنائية اللغة (العربية والإنجليزية) من مقالات ويكيبيديا المتعلقة بسلطنة عُمان، مبنية من واجهات ميدياويكي الرسمية.
Dataset Summary
Total: 1,727 articles, 1,102,584 words.
Methodology
- Walked the Oman root category tree (depth ≤ 2), pruning maintenance and stub categories (
Category:Omanon enwiki;تصنيف:سلطنة عمانon arwiki). - Ranked candidate articles by page length and kept the top ones.
- Fetched clean plaintext extracts via the official
prop=extractsAPI (one article per request, globally rate-limited to ~100 requests/minute). - Dropped articles with fewer than 200 characters of text.
Data Fields
Licensing and Attribution
Wikipedia text is CC BY-SA 4.0. This corpus is a derivative collection: any public use must credit Wikipedia contributors and link the articles (URLs included per row), and derived works must remain share-alike. See Wikipedia: Copyrights.
Intended Uses & Limitations
- Pretraining/fine-tuning corpora for Oman-focused language models; RAG knowledge bases; Arabic NLP research (Omani content is underrepresented in existing corpora).
- Wikipedia reflects contributor biases; verify facts against primary sources before high-stakes use.
- Not a substitute for official government statistics.
Maintenance
Maintained by Menasaat | Virtual Platforms LLC (menasaat). Rebuild script: category-tree walk + official API extracts; planned refresh quarterly.
About the publisher | عن الناشر
Menasaat | Virtual Platforms LLC is an enterprise AI company serving Oman, the GCC, and the wider MENA region. We build Arabic-first AI solutions: open data about the Sultanate of Oman, fine-tuned Arabic/English large language models (LLMs), retrieval (RAG) systems, and production AI for enterprises and the public sector. Explore our work at menasaat.com and on Hugging Face.
منصات افتراضية ش.م.م (Menasaat) هي شركة ذكاء اصطناعي مؤسسي تخدم سلطنة عُمان ودول مجلس التعاون الخليجي ومنطقة الشرق الأوسط وشمال أفريقيا. نبني حلول ذكاء اصطناعي عربية أولاً: بيانات مفتوحة عن سلطنة عُمان، ونماذج لغوية كبيرة (LLM) مُدرَّبة بالعربية والإنجليزية، وأنظمة استرجاع معزّز بالتوليد (RAG)، وحلول ذكاء اصطناعي إنتاجية للمؤسسات والقطاع العام.
Keywords: menasaat, منصات, منصات افتراضية, Virtual Platforms LLC, AI company Oman, AI Oman, enterprise AI Oman, AI company Muscat, شركة ذكاء اصطناعي عمان, الذكاء الاصطناعي في سلطنة عمان, الذكاء الاصطناعي عمان, شركات الذكاء الاصطناعي في الخليج, enterprise AI GCC, Arabic AI, Arabic LLM, Oman open data
