datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStories-Algerian-Darijaalgerian-darja-corpus
Algerian Darja Corpus
A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts.
Dataset Summary
The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.TinyStories-Algerian-Darija
TinyStories Algerian Darija
Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous).
The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.algerian-darja-corpus
Algerian Darja Corpus
11,151 long-form conversational transcripts in Algerian Darja for language modeling of real spoken Algerian, by Kamel Touati (Independent AI Researcher, Algiers, ORCID 0009-0000-5330-2123, Hugging Face touati-kamel), mirrored on the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-darja-corpus: 11,151 train rows) and re-counted row-by-row with datasets streaming… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-darja-corpus.algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.algerian-family-law-qa
algerian-family-law-qa
A retrieval and reranking dataset for Algerian family-law question answering — real user questions scraped from Algerian Facebook legal-advice groups, paired with relevant articles from the Algerian Family Code (Law No. 84-11).
Questions are written in Algerian Darja (dialect), Arabizi (Arabic in Latin script), Modern Standard Arabic, and French code-switched text. Documents are articles from the Algerian Family Code covering divorce, custody, alimony, and… See the full description on the dataset page: https://huggingface.co/datasets/81melody/algerian-family-law-qa.algerian-cars-realestate-search-dataset
algerian-cars-realestate-search-dataset
A multilingual query/listing relevance dataset for Algerian marketplace search, built from algerian search listings (vehicles and real estate) with user queries in Darja, Arabizi, French, and Arabic.
Each example pairs a search query with a listing document and a relevance score in [0, 1]. Positive pairs come from true query-listing matches; negatives were mined using the 81melody/algerianME5 dense retriever (hard negatives: top-retrieved… See the full description on the dataset page: https://huggingface.co/datasets/81melody/algerian-cars-realestate-search-dataset.algerian-real-estatealgerian-wav2vec2-resultsalgerian-mms-results
