datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
algerian-darja-corpus
Algerian Darja Corpus
A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts.
Dataset Summary
The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.Algerian-Darija
Overview
This dataset contains text in Algerian Darija, collected from a variety of sources including existing datasets on Hugging Face, web scraping, and YouTube transcript APIs.
The train split consists more then 2k rows of uncleaned text data.
The v1 split consists more than 170k rows of split and partially cleaned text.
Sources
The text data was gathered from:
Hugging Face Datasets: Pre-existing datasets relevant to Algerian Darija.
Web Scraping: Content… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/Algerian-Darija.algerian-darja-forum-posts
Algerian Darja Dataset
A large-scale Algerian Darja conversational dataset prepared for NLP, language-model training, instruction tuning, and conversational AI research.
3,209,157 samples · 1.193B tokens · 371.83 tokens/sample on average
Dataset at a Glance
Property
Value
Samples
3,209,157
Total tokens
1,193,257,847
Approx. tokens
1.193B
Average tokens / sample
371.83
Language
Algerian Darja
Format
Conversational JSON
Storage format… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-forum-posts.Algerian-Youtube-Comments
Algerian Youtube Comments
55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows).
The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.TinyStories-Algerian-Darija
TinyStories Algerian Darija
Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous).
The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.algerian-darja-corpus
Algerian Darja Corpus
11,151 long-form conversational transcripts in Algerian Darja for language modeling of real spoken Algerian, by Kamel Touati (Independent AI Researcher, Algiers, ORCID 0009-0000-5330-2123, Hugging Face touati-kamel), mirrored on the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-darja-corpus: 11,151 train rows) and re-counted row-by-row with datasets streaming… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-darja-corpus.algerian-darja-sample
Algerian Darja Sample
A growing Algerian Darja text corpus collected for NLP and language-modeling research.
This dataset is updated incrementally as new sources are collected and processed. Exact sample counts, file sizes, and statistics change between releases — refer to the Dataset Viewer on the repository page for current figures rather than any numbers in this card.
Dataset at a Glance
Sample count, file size, character/word counts, and other metrics are… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-sample.DziriAlign
DziriAlign
1,000 preference pairs (prompt, chosen, rejected) for aligning language models with Algerian Darja and its sociocultural norms, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/DziriAlign: 1,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/DziriAlign", split="train", streaming=True): 1,000 rows).
The default config answers: when two replies compete, which one… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/DziriAlign.algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.algerian-darija-dictionary-v1
Description
This dataset is a comprehensive dictionary of Algerian Darija words and popular proverbs. It aims to document the rich linguistic heritage and daily expressions used in Algeria.
⚠️ Quality Warning
Please note that the current version of the dataset contains noise and spam. It is a work in progress and is not yet 100% accurate.
Dataset Statistics
Number of entries: 4,636
Columns: word, french_writing, definition, example… See the full description on the dataset page: https://huggingface.co/datasets/awras/algerian-darija-dictionary-v1.Algerian-Darija
Overview
This dataset contains text in Algerian Darija, collected from a variety of sources including existing datasets on Hugging Face, web scraping, and YouTube transcript APIs.
The train split consists more then 2k rows of uncleaned text data.
The v1 split consists more than 170k rows of split and partially cleaned text.
Sources
The text data was gathered from:
Hugging Face Datasets: Pre-existing datasets relevant to Algerian Darija.
Web Scraping: Content… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/Algerian-Darija.
