CoolFace
Datasetpublic

abdeljalilELmajjodi/facebook_darija_dataset

Darija Facebook Posts Dataset Dataset Details Dataset Description This dataset consists of more than 5k public posts from Facebook. Each post contains text content, metadata. This dataset containt more than 400K darija tokens. Curated by: @abdeljalilELmajjodi Language(s) (NLP): Multiple (primarily Moroccon Arabic) Uses This dataset could be used for: Training and testing language models on social media content Analyzing social… See the full description on the dataset page: https://huggingface.co/datasets/abdeljalilELmajjodi/facebook_darija_dataset.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
6likes64downloads
Dataset Card

Darija Facebook Posts Dataset

Dataset Details

Dataset Description

<!-- Provide a longer summary of what this dataset is. --> This dataset consists of more than 5k public posts from Facebook. Each post contains text content, metadata. This dataset containt more than 400K darija tokens.

  • —Curated by: @abdeljalilELmajjodi
  • —Language(s) (NLP): Multiple (primarily Moroccon Arabic)

Uses

This dataset could be used for:

  • —Training and testing language models on social media content
  • —Analyzing social media posting patterns
  • —Studying conversation structures and reply networks
  • —Research on social media content moderation
  • —Natural language processing tasks using social media datas

Dataset Structure

Contains the following fields for each post:

  • —text: The main content of the post
  • —PageName: Timestamp of post creation
  • —langidentity: Language code (aryArab,ara_Arab,...)

Other

To Identify posts language we used Gherbal classifier by @Sawalni, and it gives the next results:

image/png

image/png

Bias, Risks, and Limitations

The goal of this dataset is for you to have fun :)