datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStories-Algerian-DarijaAlgerian-Youtube-Commentsarabic-audio-collection-algerian-loubna-stories
Loubna Stories Arabic Speech Dataset
Dataset Summary
The Loubna Stories Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 237 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-loubna-stories.arabic-audio-collection-algerian-kahwa-postcast
Kahwa Postcast Arabic Speech Dataset
Dataset Summary
The Kahwa Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 110 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-kahwa-postcast.algerian-darja-corpus
Algerian Darja Corpus
A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts.
Dataset Summary
The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.Algerian-Darija
Overview
This dataset contains text in Algerian Darija, collected from a variety of sources including existing datasets on Hugging Face, web scraping, and YouTube transcript APIs.
The train split consists more then 2k rows of uncleaned text data.
The v1 split consists more than 170k rows of split and partially cleaned text.
Sources
The text data was gathered from:
Hugging Face Datasets: Pre-existing datasets relevant to Algerian Darija.
Web Scraping: Content… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/Algerian-Darija.algerian_ttsalgerian-darja-forum-posts
Algerian Darja Dataset
A large-scale Algerian Darja conversational dataset prepared for NLP, language-model training, instruction tuning, and conversational AI research.
3,209,157 samples · 1.193B tokens · 371.83 tokens/sample on average
Dataset at a Glance
Property
Value
Samples
3,209,157
Total tokens
1,193,257,847
Approx. tokens
1.193B
Average tokens / sample
371.83
Language
Algerian Darja
Format
Conversational JSON
Storage format… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-forum-posts.Algerian-Youtube-Comments
Algerian Youtube Comments
55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows).
The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.TinyStories-Algerian-Darija
TinyStories Algerian Darija
Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous).
The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.arabic-audio-collection-algerian-rawi
Rawi Postcast Arabic Speech Dataset
Dataset Summary
The Rawi Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 51 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-rawi.Algerian-STT-Super-Dataset-V2algerian-license-plates
Algerian License Plates Detection Dataset
First open-source annotated Algerian license plate detection dataset.
1950 images with YOLO-format bounding box annotations, collected from
publicly available Algerian media sources.
Annotation Format
Each .txt label file: one row per plate
0 cx cy w h
where cx, cy, w, h are normalized to [0, 1]. Class 0 = plate.
Model trained on this dataset
YOLOv8s achieves mAP@50 = 0.993.
Live demo: https://tobni-algerplate.hf.space… See the full description on the dataset page: https://huggingface.co/datasets/tobni/algerian-license-plates.cafe-algerian-codeswitch-speech
CAFE Algerian Codeswitch Speech
This dataset contains Algerian Arabic and French code-switched speech.
Repository Path: FatimahEmadEldin/cafe-algerian-codeswitch-speech
algerian-darja-sample
Algerian Darja Sample
A growing Algerian Darja text corpus collected for NLP and language-modeling research.
This dataset is updated incrementally as new sources are collected and processed. Exact sample counts, file sizes, and statistics change between releases — refer to the Dataset Viewer on the repository page for current figures rather than any numbers in this card.
Dataset at a Glance
Sample count, file size, character/word counts, and other metrics are… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-sample.algerian-darja-corpus
Algerian Darja Corpus
11,151 long-form conversational transcripts in Algerian Darja for language modeling of real spoken Algerian, by Kamel Touati (Independent AI Researcher, Algiers, ORCID 0009-0000-5330-2123, Hugging Face touati-kamel), mirrored on the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-darja-corpus: 11,151 train rows) and re-counted row-by-row with datasets streaming… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-darja-corpus.Algerian-STT-Cleaned-V5algerian-arabic-english-translation-50k
Algerian Arabic English Translation 50K
50,000 aligned Algerian Darja to English sentence pairs for machine translation, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-arabic-english-translation-50k: 50,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/algerian-arabic-english-translation-50k", split="train", streaming=True): 50,000 rows).
The default config answers:… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-arabic-english-translation-50k.DziriEval
DziriEval
1,000 native multiple-choice questions in Algerian Darja for LLM evaluation, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/DziriEval: 1,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/DziriEval", split="train", streaming=True): 1,000 rows).
The default config answers: does the model understand Algerian culture, geography, history, everyday life, and the Darja… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/DziriEval.algerian-arabic-english-translation-50k
Algerian Arabic (Darija) ↔ English Translation (50K) 🇩🇿
An open-source parallel corpus of 50,000 Algerian Arabic (Darija) sentences paired with their English translations. Released to help the research community and developers build and evaluate NLP models, translation systems, and LLMs for Algerian Arabic / Maghrebi Dialect — an under-resourced variety of Arabic.
Use it freely for machine translation, LLM fine-tuning, evaluation, dialectal Arabic NLP, and data augmentation.… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-arabic-english-translation-50k.DziriAlign
DziriAlign
1,000 preference pairs (prompt, chosen, rejected) for aligning language models with Algerian Darja and its sociocultural norms, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/DziriAlign: 1,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/DziriAlign", split="train", streaming=True): 1,000 rows).
The default config answers: when two replies compete, which one… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/DziriAlign.Algerian-STT-Cleaned-V3algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.Algerian-Arabic***This dataset contains 1.5k Algerian Arabic sentiment comments classified into two classes
subjective positive, subjective negative.
***This dataset is collected and annotated by RANIM for Arabic NLP Solutions, feel free to use it.
***We appreciate citing our company name "RANIM for Arabic NLP Solutions" when using this dataset.
***For more data/information visit our website : https://ranim-for-nlp.web.app
or contact us : ranim.for.nlp@gmail.com
"RANIM for… See the full description on the dataset page: https://huggingface.co/datasets/ranim/Algerian-Arabic.algerian-darija-dictionary-v1
Description
This dataset is a comprehensive dictionary of Algerian Darija words and popular proverbs. It aims to document the rich linguistic heritage and daily expressions used in Algeria.
⚠️ Quality Warning
Please note that the current version of the dataset contains noise and spam. It is a work in progress and is not yet 100% accurate.
Dataset Statistics
Number of entries: 4,636
Columns: word, french_writing, definition, example… See the full description on the dataset page: https://huggingface.co/datasets/awras/algerian-darija-dictionary-v1.algerian-realestate-ner-dataset
Algerian-realestate-NER-dataset
Dataset Description
This is a specialized Named Entity Recognition dataset extracted from the complex reality of the Algerian digital real-estate market in Facebook groups, it contains 13 labeled entity and 7138 training example
Real estate advertisements in Algeria (found on Facebook groups) are unstructured , noisy and bloated with code-switching between Algerian Darja (dialect), Arabizi, Standard Arabic, and French, This dataset… See the full description on the dataset page: https://huggingface.co/datasets/81melody/algerian-realestate-ner-dataset.algerian-family-law-qa
algerian-family-law-qa
A retrieval and reranking dataset for Algerian family-law question answering — real user questions scraped from Algerian Facebook legal-advice groups, paired with relevant articles from the Algerian Family Code (Law No. 84-11).
Questions are written in Algerian Darja (dialect), Arabizi (Arabic in Latin script), Modern Standard Arabic, and French code-switched text. Documents are articles from the Algerian Family Code covering divorce, custody, alimony, and… See the full description on the dataset page: https://huggingface.co/datasets/81melody/algerian-family-law-qa.algerian-code-switched-asr-eval
Dataset Description
This dataset contains 335 manually transcribed audio segments (approximately 30 minutes total) extracted from Algerian political and misinformation-related social-media videos, collected from public interviews on TikTok and YouTube in September 2025. It was built as the evaluation set for What WER Hides: A Closer Look at Algerian Code-Switched ASR, and is intended for benchmarking automatic speech recognition (ASR) systems on Algerian dialectal and… See the full description on the dataset page: https://huggingface.co/datasets/oist/algerian-code-switched-asr-eval.algerian-cars-realestate-search-dataset
algerian-cars-realestate-search-dataset
A multilingual query/listing relevance dataset for Algerian marketplace search, built from algerian search listings (vehicles and real estate) with user queries in Darja, Arabizi, French, and Arabic.
Each example pairs a search query with a listing document and a relevance score in [0, 1]. Positive pairs come from true query-listing matches; negatives were mined using the 81melody/algerianME5 dense retriever (hard negatives: top-retrieved… See the full description on the dataset page: https://huggingface.co/datasets/81melody/algerian-cars-realestate-search-dataset.algerian-drug-nomenclature
🇩🇿 Algerian Drug Nomenclature — Nomenclature Nationale des Médicaments
A clean, structured, PII-free reference dataset of pharmaceutical products authorised for the Algerian market, derived from the official Nomenclature Nationale des Médicaments (Ministère de la Santé, Algeria).
This is reference product data — registration numbers, INNs, brand names, dosages, manufacturers, and reimbursement status. It contains no personal, patient, or transactional data of any kind.… See the full description on the dataset page: https://huggingface.co/datasets/tkawen/algerian-drug-nomenclature.
