algerian-nlp/algerian-arabic-english-translation-50k
Algerian Arabic English Translation 50K 50,000 aligned Algerian Darja to English sentence pairs for machine translation, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-arabic-english-translation-50k: 50,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/algerian-arabic-english-translation-50k", split="train", streaming=True): 50,000 rows). The default config answers:… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-arabic-english-translation-50k.
Algerian Arabic English Translation 50K
50,000 aligned Algerian Darja to English sentence pairs for machine translation, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-arabic-english-translation-50k: 50,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/algerian-arabic-english-translation-50k", split="train", streaming=True): 50,000 rows).
The default config answers: how do you translate everyday Algerian Darja into English and back.
from datasets import load_dataset
ds = load_dataset("algerian-nlp/algerian-arabic-english-translation-50k", "default", split="train")default — aligned Darija-English pairs
50,000 rows, all in train. No dev or test split ships with this repo — split your own and publish the seed, or scores will not compare.
The gap this config fixes, with the measurement: Algerian Darja has almost no parallel data at this scale. 8,699 of 50,000 Arabic sides (17.4%, counted by streaming filter len(Arabic) > 200) exceed 200 characters, so the set skews toward short conversational sentences rather than documents.
Key on (id): ids are dense integers, safe to join against. Four file formats ship the same pairs (.parquet 13 MB recommended, .csv 22.8 MB, .jsonl 24.5 MB, .json 25.6 MB — author sizes, not reweighed by us); the default config reads the parquet only. What breaks if ignored: mixing formats across runs can reorder rows — the parquet id order is the canonical one.
Scoring protocol
SacreBLEU and chrF on a held-out split you define, reported with the signature (BLEU|nrefs:1|...) and the split seed. Scores without a signature and seed are NOT comparable. Unmeasured: no baseline scores have been produced by us on this protocol.
Results
Unmeasured. No baseline results scored by us in the protocol above.
Usage
default: Darija to English and English to Darija translation.
from datasets import load_dataset
ds = load_dataset("algerian-nlp/algerian-arabic-english-translation-50k", "default", split="train")
ds[0]
# id: 0
# Arabic: علاش ما راكش حاب تآمن بلي نحبك صح؟ ...
# English: Why don't you want to believe that I truly love you? ...
# Longer Arabic sides (measured 2026-09-17, streaming, 50,000 rows scanned):
subset = ds.filter(lambda row: len(row["Arabic"]) > 200)
len(subset) # 8699What this dataset does NOT contain — and what breaks if you assume it does. No dev or test split, no dialect-region tags, no translator ids, and no per-row source id, licence, or retrieval date. Do not assume translation direction: nothing records whether a pair was authored Darija-first or English-first.
Script & code-switching distribution
Unmeasured. No per-script counts were published and none were counted by us. The Arabic side is Arabic-script by inspection of row 0; Arabizi coverage, if any, is undocumented.
Provenance and limits
- The
defaultconfig is authored unser translation pairs and inherits Apache-2.0. The sentences are the creators', not ours. - Per-row provenance is missing: rows carry no source id, no licence, and no retrieval date beyond the integer
id. - Filtration audit. Unmeasured by us. No translator count, review pass, or quality-sampling rate is published.
- Bounded, not cleared. Translationese and MSA-flavoured Darija cannot be excluded: no dialect-purity measurement ships with the set.
- Out of scope. Arabizi input, code-switched input, and document-level translation were not targeted; scoring them here would corrupt the sentence-translation measurement.
- Decontamination of any training corpus against this dataset is unmeasured.
Citation
@misc{algerian_nlp_algerian_arabic_english_translation_50k,
title = {Algerian Arabic English Translation 50K: 50,000 Darija-English pairs},
author = {Algerian NLP Collective},
year = {2026},
url = {https://huggingface.co/datasets/algerian-nlp/algerian-arabic-english-translation-50k}
}No DOI, arXiv ID, or named author is published on the repo; the author field above is the collective fallback. No upstream sources to cite.
Licence
Apache-2.0 for this data. Patent grant and attribution terms apply as written in the licence text.
