abdeljalilELmajjodi/facebook_darija_dataset
Darija Facebook Posts Dataset Dataset Details Dataset Description This dataset consists of more than 5k public posts from Facebook. Each post contains text content, metadata. This dataset containt more than 400K darija tokens. Curated by: @abdeljalilELmajjodi Language(s) (NLP): Multiple (primarily Moroccon Arabic) Uses This dataset could be used for: Training and testing language models on social media content Analyzing social… See the full description on the dataset page: https://huggingface.co/datasets/abdeljalilELmajjodi/facebook_darija_dataset.
Darija Facebook Posts Dataset
Dataset Details
Dataset Description
<!-- Provide a longer summary of what this dataset is. --> This dataset consists of more than 5k public posts from Facebook. Each post contains text content, metadata. This dataset containt more than 400K darija tokens.
- Curated by: @abdeljalilELmajjodi
- Language(s) (NLP): Multiple (primarily Moroccon Arabic)
Uses
This dataset could be used for:
- Training and testing language models on social media content
- Analyzing social media posting patterns
- Studying conversation structures and reply networks
- Research on social media content moderation
- Natural language processing tasks using social media datas
Dataset Structure
Contains the following fields for each post:
- text: The main content of the post
- PageName: Timestamp of post creation
- langidentity: Language code (aryArab,ara_Arab,...)
Other
To Identify posts language we used Gherbal classifier by @Sawalni, and it gives the next results:


Bias, Risks, and Limitations
The goal of this dataset is for you to have fun :)
