CoolFace
Datasetpublic

touati-kamel/algerian-darja-corpus

Algerian Darja Corpus A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts. Dataset Summary The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.

sourceHugging Facecc-by-4.0updated 9d agoView on Hugging Face
6likes303downloads
Dataset Card

Algerian Darja Corpus

![DOI](https://doi.org/10.5281/zenodo.22723182) ![License: CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)

A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts.

Dataset Summary

The Algerian Darja Corpus consists of transcribed dialogue segments (specifically from podcasts and talk shows, featuring various hosts and guests). Due to the nature of modern Algerian speech, the text exhibits strong code-switching characteristics.

  • —Total Documents (Transcripts): 11,151
  • —Total Word Count: 34,336,366 words
  • —Total Character Count: 186,106,376 characters
  • —Data Format: JSON Lines (.jsonl)

Advanced Preprocessing & Data Cleaning

The corpus has undergone rigorous NLP preprocessing tailored specifically for Algerian Darja and code-switched dialogues:

  1. 1.Timestamp & Subtitle Cleanup:
  2. 2.Stripped all SRT timestamp lines (00:01:43,300 --> 00:01:49,400), standalone time markers (00:00:37, [01:23], 0:18-0:20), and standalone subtitle index numbers.
  3. 3.Removed sound, music, ambient action tags, and ASR paralinguistic speech event tags (e.g., [Music], [موسيقى], (ضحك), (laughs), [تصفيق], *Music playing*, <laugh>, <cry>, <weep>, <sob>, <scream>, <shout>, <whisper>, <sigh>, <gasp>, <groan>, <moan>, <pause>, <hes>, <stutter>, <breath>, <sniff>, <cough>, <throat_clear>).
  4. 4.Stripped sporadic speaker turn labels and prefixes (e.g., Guest:, Host:, Speaker 1:, المتحدث:, المذيع:) to ensure seamless continuous prose flow for causal language modeling.
  1. 1.Deduplication & Redundancy Reduction:
  2. 2.Deduplicated identical and consecutive repeated subtitle lines.
  3. 3.Removed immediate word/phrase stutters and overlapping ASR loops (n-grams repeating immediately).
  4. 4.Collapsed character elongations to a maximum of 2 occurrences (e.g., هههههههههه -> هه, بزاااااف -> بزااف).
  5. 5.Deduplicated identical documents across the corpus.
  1. 1.Noise & Foreign Character Filtering:
  2. 2.Filtered leaked foreign scripts (Indic/Devanagari, Tamil, Chinese, Cyrillic, Hebrew) from stray subtitle translations.
  3. 3.Stripped Unicode control characters (\x00-\x1F, \x7F-\x9F), zero-width formatting spaces (\u200B, \u200C, \u200D, \uFEFF), and replacement characters (\uFFFD).
  4. 4.Cleaned HTML tags (<...>) and unescaped HTML entities (&amp; -> &, &quot; -> ").
  5. 5.Removed raw URLs and emojis.
  1. 1.Normalization & Dialect Preservation:
  2. 2.Applied Unicode NFKC normalization.
  3. 3.Stripped Arabic Tatweel / Kashida (ـ).
  4. 4.Removed short vowel diacritics (tashkeel) while preserving Shaddah (ّ) for dialectal and phonetic integrity.
  5. 5.Preserved original Arabic letters (avoiding aggressive MSA Alef/Hamza flattening) to maintain French loanword distinctions.
  6. 6.Filtered out uninformative empty or tiny documents (< 5 words).
  1. 1.Strict Dialect Filtering & Language Purification:
  2. 2.Audited the entire corpus against comprehensive Algerian dialectal markers (particles, verbs, pronouns, interrogatives, negations, vocatives, and Arabizi tokens).
  3. 3.Purged non-dialect documents (pure English translated subtitles, pure French business podcasts, and pure Modern Standard Arabic / non-Algerian texts).
  4. 4.Guaranteed that 100% of the retained 11,151 documents contain genuine Algerian Darja dialect content and authentic code-switched dialogues.

Dataset Structure & Configurations

The corpus is provided in two official configurations:

  1. 1.`default` (`algerian_darja_corpus.jsonl`): The clean, filtered long-form conversational transcripts in Algerian Darja (11,151 documents, 34,336,366 words).
  2. 2.`conditioned` (`algerian_darja_corpus_conditioned.jsonl`): The domain-conditioned version where each document is categorized and prefixed with a dedicated domain control token for steerable continued pre-training (CPT) and domain-specific generation.

Domain Distribution (conditioned configuration)

Domain TagThematic ScopeDocuments% DocsVolume (Words)% Volume
<general>Open discussions, social debates, culture, multi-thematic7,13464.0 %20,906,87560.9 %
<cooking>Traditional dishes, pastry, kitchen recipes, culinary1,85316.6 %3,414,7589.9 %
<tech>Smartphones, PC hardware, unboxing, software, specs1,28311.5 %4,602,46413.4 %
<sports>Algerian football, national team (Fennecs), leagues3813.4 %2,080,3016.1 %
<podcast>Entrepreneurship, startups, business, career, self-growth2432.2 %1,547,9024.5 %
<story>True crime, real-life narratives, judicial inquiries, folktales2161.9 %1,529,8334.5 %
<lifestyle_vlog>Daily routines, travel vlogs, shopping hauls, home care300.3 %142,7330.4 %
<comedy>Pranks, hidden cameras, stand-up sketches, satire110.1 %111,5000.3 %
Total11,151100.0 %34,336,366100.0 %

Data Fields

Standard Configuration (default)
  • —text (string): The full transcript text of the conversation segment.
  • —word_count (int): Total word count of the transcript.
  • —character_count (int): Total character count of the transcript.
Conditioned Configuration (conditioned)
  • —text (string): The intact original transcript.
  • —domain (string): The classified domain tag (e.g., <cooking>, <tech>, <podcast>).
  • —conditioned_text (string): The transcript prefixed with the domain control tag (<domain> text...) ready for causal LM pre-training.
  • —word_count (int): Total word count.
  • —character_count (int): Total character count.

Data Instance Examples

Default
json
{
  "text": "سلام عليكم ومرحباً بكم في حلقة جديدة من podcast FluentlyTalk... لاباس، والله غير الحمد لله...",
  "word_count": 4173,
  "character_count": 23513
}
Conditioned
json
{
  "text": "سلام عليكم ومرحباً بكم في حلقة جديدة من podcast FluentlyTalk... لاباس، والله غير الحمد لله...",
  "domain": "<podcast>",
  "conditioned_text": "<podcast> سلام عليكم ومرحباً بكم في حلقة جديدة من podcast FluentlyTalk...",
  "word_count": 4173,
  "character_count": 23513
}

Language & Writing Systems

Algerian Darja is a spoken dialect that doesn't have a single standardized writing system. This dataset reflects real-world orthographic choices:

  • —Arabic Script: Traditional Arabic letters used to write Darja words phonetically.
  • —Latin Script (Arabizi / Franco-Algerian): Text written using the Latin alphabet, often incorporating French spelling conventions or English loanwords.
  • —Code-Switching: Frequent alternation between Algerian Darja, French, and English within the same sentence or turn.

How to Use

You can load this dataset directly using the Hugging Face datasets library:

python
from datasets import load_dataset

# 1. Load the default cleaned dialectal corpus
dataset = load_dataset("touati-kamel/algerian-darja-corpus", "default")
print("First default document:", dataset['train'][0]['text'][:120])

# 2. Load the domain-conditioned corpus for steerable pre-training
dataset_cond = load_dataset("touati-kamel/algerian-darja-corpus", "conditioned")
print("Domain tag:", dataset_cond['train'][0]['domain'])
print("First conditioned document:", dataset_cond['train'][0]['conditioned_text'][:120])

Data Sources

The transcripts in this corpus were collected from the following YouTube channels:

  • —بودكاست المفيد El Moufid Podcast
  • —Keepodcast
  • —BelkadiManel
  • —takicharni
  • —ilyesderradji
  • —MohamedDjamelTaleb
  • —ramziZRT
  • —Khoubai
  • —Intaj
  • —Omar Rahmoun
  • —Mouslem khirouni
  • —Fluently
  • —Zaki Agha CasTea
  • —Anis Hamidi
  • —BrainerX
  • —Raouf Talks
  • —Alias Djamel
  • —Entrepreneur podcast - بودكاست المقاول
  • —NOEST Express
  • —Dar El Mic – دار الميك
  • —Flown marketing
  • —King's Podcast
  • —السبق قبل حدوثه للمحلل خزناجي نواري
  • —Dr Courage
  • —بودكاست أبيض (Abyadh Podcast)
  • —Itz Bcf
  • —Achraf Bcf
  • —Wassim Guessoum
  • —Sohaib & Abdou | صهيب و عبدو
  • —ZEERO
  • —Zaki Rb
  • —Aljadidtv
  • —حدود اللَّه
  • —سياسة الجزائر
  • —LAMIN Tube
  • —الشبل Chible Dz
  • —Nora Trabelsi
  • —YOUCEF ZAROUTA
  • —Mimho
  • —Viper Beyaz
  • —Raouf Belkacemi
  • —Sifeddin Oukhlif
  • —La3ziz Dz
  • —Abdou boutaleb
  • —Ayman Cha7el
  • —Ayman Cha7el 2
  • —DZ.Horror
  • —Raklita
  • —Raklita Extra
  • —Oum Walid
  • —مطبخ أم جواد جزائرية 🇩🇿
  • —RIFKA
  • —Sultan Achour
  • —TH8 professional
  • —Afinou
  • —Mamin Zeroual React
  • —Maamar Killer Bee
  • —sidahmed kick
  • —Raiid Vlogs
  • —حلويات إقتصادية و راقية (Oum Yara)
  • —مطبخ أم لينا الشاوية
  • —وصفتي - Wasfati
  • —Sabah kitchen
  • —cuisine samira dz مطبخ سميرة
  • —Zinou 2.0
  • —Dz21 (@skikda-k8p)
  • —Le Bon Coin Algerien
  • —ابن الدولة الجزائرية الحرة
  • —Khaled Ben arts (@khaledbenarts)
  • —MaghrebTech (@TECHMAGHREB)
  • —AMINE TAYEBI (@logiofficiel)
  • —yakoub tech (@yakoubtech)
  • —Just AINAR (@justAINAR)
  • —Barid Dz (@bariddz)
  • —ألجيريا ويب - Algeria Web (@AlgeriaWeb)
  • —Firmi (@Firmi09)
  • —Omar Tech -عمر تك (@omartech02)
  • —YACINE BEN (@yacine_benamara)
  • —OSSRA TECH (@OSSRATECH)
  • —Hocine Sellaoui (@HocineSellaoui)
  • —Bilal DN (@bilaldn)

and from huggingface datasets:

  • —oddadmix/arabic-audio-collection-algerian-kahwa-postcast
  • —oddadmix/arabic-audio-collection-algerian-loubna-stories
  • —oddadmix/arabic-audio-collection-algerian-rawi

License

This dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.

Citation

If you use this corpus in your research or project, please cite it using the following BibTeX entry:

bibtex
@dataset{touati_2026_22723182,
  author       = {Touati, Kamel},
  title        = {{Algerian Darja Corpus: A 34-Million-Word Long-Form Conversational Dataset for Dialectal Language Modeling}},
  month        = sep,
  year         = 2026,
  publisher    = {Zenodo},
  version      = {2.1.0},
  doi          = {10.5281/zenodo.22723182},
  url          = {https://doi.org/10.5281/zenodo.22723182}
}

@article{touati2026algerian,
  author       = {Touati, Kamel},
  title        = {{Algerian Darja Corpus: A 34-Million-Word Long-Form Conversational Dataset for Dialectal Language Modeling}},
  journal      = {arXiv preprint},
  year         = {2026},
  doi          = {10.5281/zenodo.22723182},
  url          = {https://doi.org/10.5281/zenodo.22723182}
}