datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
asjp-cerist
ASJP / CERIST Academic PDF Corpus
Overview
DarjaCore ASJP/CERIST is a large-scale collection of academic PDF documents collected from the Algerian ASJP (Algerian Scientific Journal Platform) ecosystem associated with CERIST (Centre de Recherche sur l'Information Scientifique et Technique).
The dataset preserves the original PDF documents rather than converting them into extracted text. This allows researchers to perform their own OCR, document parsing… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/asjp-cerist.algerian-darja-corpus
Algerian Darja Corpus
A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts.
Dataset Summary
The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.algerian-darja-forum-posts
Algerian Darja Dataset
A large-scale Algerian Darja conversational dataset prepared for NLP, language-model training, instruction tuning, and conversational AI research.
3,209,157 samples · 1.193B tokens · 371.83 tokens/sample on average
Dataset at a Glance
Property
Value
Samples
3,209,157
Total tokens
1,193,257,847
Approx. tokens
1.193B
Average tokens / sample
371.83
Language
Algerian Darja
Format
Conversational JSON
Storage format… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-forum-posts.algerian-darja-sample
Algerian Darja Sample
A growing Algerian Darja text corpus collected for NLP and language-modeling research.
This dataset is updated incrementally as new sources are collected and processed. Exact sample counts, file sizes, and statistics change between releases — refer to the Dataset Viewer on the repository page for current figures rather than any numbers in this card.
Dataset at a Glance
Sample count, file size, character/word counts, and other metrics are… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-sample.algerian-darja-corpus
Algerian Darja Corpus
11,151 long-form conversational transcripts in Algerian Darja for language modeling of real spoken Algerian, by Kamel Touati (Independent AI Researcher, Algiers, ORCID 0009-0000-5330-2123, Hugging Face touati-kamel), mirrored on the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-darja-corpus: 11,151 train rows) and re-counted row-by-row with datasets streaming… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-darja-corpus.darja-tounsi1darja-tunisie-chat
Dataset Card for "darja-tunisie-chat"
More Information needed
darja-blindspot-eval
Blind Spot Evaluation: Algerian Darja, Arabizi, and French Code-Switching
Author: Khadija Abderrahmane
Model evaluated: Qwen/Qwen2.5-1.5B-Instruct (1.5B parameters)
1. The blind spot
Nearly every widely used Arabic NLP benchmark — ArabicMMLU, ARLUE, AraSentiment, and most Arabic instruction-tuning datasets — is built almost entirely on Modern Standard Arabic (MSA), with limited coverage of major spoken dialects (Egyptian, Gulf, Levantine). Algerian Darja, the… See the full description on the dataset page: https://huggingface.co/datasets/khadidjaabderrahmane/darja-blindspot-eval.koch_small_block_2cam_120ep_datadarja_sentiment.csvdarjaAI_datasetdarja-14msa-darja-pairs-completekoch_small_block_datakoch_small_block_120ep_datakoch_small_block__2cam_140ep_dataeval_koch_testkoch_test_datamsa-darja-pairs
