CoolFace
Datasetpublic

abdelhaqueidali/The-Amazigh-World-Press-Dataset

The Amazigh World Press Dataset This dataset consists of articles scraped directly from the official web platform of Amadal Amazigh (العالم الأمازيغي / ⴰⵎⴰⴹⴰⵍ ⴰⵎⴰⵣⵉⵖ). It provides a valuable text corpus of contemporary Amazigh press, journalism, and cultural commentary, formatted specifically for natural language processing (NLP), text mining, and language modeling tasks in the Standard Moroccan Amazigh variant (zgh) and across broader regional dialects (ber).… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/The-Amazigh-World-Press-Dataset.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes9downloads
Dataset Card

The Amazigh World Press Dataset

This dataset consists of articles scraped directly from the official web platform of Amadal Amazigh (العالم الأمازيغي / ⴰⵎⴰⴹⴰⵍ ⴰⵎⴰⵣⵉⵖ). It provides a valuable text corpus of contemporary Amazigh press, journalism, and cultural commentary, formatted specifically for natural language processing (NLP), text mining, and language modeling tasks in the Standard Moroccan Amazigh variant (zgh) and across broader regional dialects (ber).

Dataset Description

  • —Source: Scraped from the prominent independent newspaper platform, Amadal Amazigh.
  • —Content: News articles, editorials, and commentary text covering culture, current affairs, identity, and regional documentation.
  • —Size: < 1K compiled entries.
  • —Scripts: Primarily features Standard Moroccan Amazigh text written in the Amazigh alphabet (Tifinagh) alongside.

Source Credit & Provenance

All primary data within this dataset is the intellectual property of Amadal Amazigh (Le Monde Amazigh).

Founded in 2001, Amadal Amazigh serves as a cornerstone of independent, militant print and digital journalism dedicated to promoting the Amazigh identity, language, and cultural heritage across Morocco and the broader Tamazgha region. This dataset is compiled strictly for non-commercial linguistic preservation, educational research, and machine learning development aimed at improving low-resource language technologies.


Dataset Structure


Use Cases

  • —Language Modeling: Training and fine-tuning small language models (SLMs) and tokenizers on official, formal Tifinagh print prose.
  • —Topic Classification & Named Entity Recognition (NER): Categorizing journalistic text or identifying regional figures, organizations, and geographical features in Tamazight text.
  • —Computational Linguistics: Analyzing structural syntax, media vocabulary evolution, and linguistic standardization patterns within modern digital press.

Source

https://amadalamazigh.press.ma/tamazight