abdelhaqueidali/The-Amazigh-World-Press-Dataset
The Amazigh World Press Dataset This dataset consists of articles scraped directly from the official web platform of Amadal Amazigh (العالم الأمازيغي / ⴰⵎⴰⴹⴰⵍ ⴰⵎⴰⵣⵉⵖ). It provides a valuable text corpus of contemporary Amazigh press, journalism, and cultural commentary, formatted specifically for natural language processing (NLP), text mining, and language modeling tasks in the Standard Moroccan Amazigh variant (zgh) and across broader regional dialects (ber).… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/The-Amazigh-World-Press-Dataset.
The Amazigh World Press Dataset
This dataset consists of articles scraped directly from the official web platform of Amadal Amazigh (العالم الأمازيغي / ⴰⵎⴰⴹⴰⵍ ⴰⵎⴰⵣⵉⵖ). It provides a valuable text corpus of contemporary Amazigh press, journalism, and cultural commentary, formatted specifically for natural language processing (NLP), text mining, and language modeling tasks in the Standard Moroccan Amazigh variant (zgh) and across broader regional dialects (ber).
Dataset Description
- Source: Scraped from the prominent independent newspaper platform, Amadal Amazigh.
- Content: News articles, editorials, and commentary text covering culture, current affairs, identity, and regional documentation.
- Size:
< 1Kcompiled entries. - Scripts: Primarily features Standard Moroccan Amazigh text written in the Amazigh alphabet (Tifinagh) alongside.
Source Credit & Provenance
All primary data within this dataset is the intellectual property of Amadal Amazigh (Le Monde Amazigh).
Founded in 2001, Amadal Amazigh serves as a cornerstone of independent, militant print and digital journalism dedicated to promoting the Amazigh identity, language, and cultural heritage across Morocco and the broader Tamazgha region. This dataset is compiled strictly for non-commercial linguistic preservation, educational research, and machine learning development aimed at improving low-resource language technologies.
Dataset Structure
Use Cases
- Language Modeling: Training and fine-tuning small language models (SLMs) and tokenizers on official, formal Tifinagh print prose.
- Topic Classification & Named Entity Recognition (NER): Categorizing journalistic text or identifying regional figures, organizations, and geographical features in Tamazight text.
- Computational Linguistics: Analyzing structural syntax, media vocabulary evolution, and linguistic standardization patterns within modern digital press.
Source
https://amadalamazigh.press.ma/tamazight
