CoolFace
Datasetpublic

algerian-nlp/algerian-arabic-english-translation-50k

Algerian Arabic English Translation 50K 50,000 aligned Algerian Darja to English sentence pairs for machine translation, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-arabic-english-translation-50k: 50,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/algerian-arabic-english-translation-50k", split="train", streaming=True): 50,000 rows). The default config answers:… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-arabic-english-translation-50k.

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
0likes92downloads
Dataset Card

Algerian Arabic English Translation 50K

50,000 aligned Algerian Darja to English sentence pairs for machine translation, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-arabic-english-translation-50k: 50,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/algerian-arabic-english-translation-50k", split="train", streaming=True): 50,000 rows).

The default config answers: how do you translate everyday Algerian Darja into English and back.

python
from datasets import load_dataset

ds = load_dataset("algerian-nlp/algerian-arabic-english-translation-50k", "default", split="train")

default — aligned Darija-English pairs

50,000 rows, all in train. No dev or test split ships with this repo — split your own and publish the seed, or scores will not compare.

The gap this config fixes, with the measurement: Algerian Darja has almost no parallel data at this scale. 8,699 of 50,000 Arabic sides (17.4%, counted by streaming filter len(Arabic) > 200) exceed 200 characters, so the set skews toward short conversational sentences rather than documents.

field
idinteger row id — row 0 carries id 0; dense 0-to-49,999 ordering is assumed by the filename convention, not verified row by row — key on (id) only after asserting uniqueness yourself
Arabicthe Algerian Darja sentence
Englishthe English translation, aligned one-to-one

Key on (id): ids are dense integers, safe to join against. Four file formats ship the same pairs (.parquet 13 MB recommended, .csv 22.8 MB, .jsonl 24.5 MB, .json 25.6 MB — author sizes, not reweighed by us); the default config reads the parquet only. What breaks if ignored: mixing formats across runs can reorder rows — the parquet id order is the canonical one.

Scoring protocol

SacreBLEU and chrF on a held-out split you define, reported with the signature (BLEU|nrefs:1|...) and the split seed. Scores without a signature and seed are NOT comparable. Unmeasured: no baseline scores have been produced by us on this protocol.

Results

Unmeasured. No baseline results scored by us in the protocol above.

Usage

default: Darija to English and English to Darija translation.

python
from datasets import load_dataset

ds = load_dataset("algerian-nlp/algerian-arabic-english-translation-50k", "default", split="train")
ds[0]
# id: 0
# Arabic: علاش ما راكش حاب تآمن بلي نحبك صح؟ ...
# English: Why don't you want to believe that I truly love you? ...

# Longer Arabic sides (measured 2026-09-17, streaming, 50,000 rows scanned):
subset = ds.filter(lambda row: len(row["Arabic"]) > 200)
len(subset)  # 8699

What this dataset does NOT contain — and what breaks if you assume it does. No dev or test split, no dialect-region tags, no translator ids, and no per-row source id, licence, or retrieval date. Do not assume translation direction: nothing records whether a pair was authored Darija-first or English-first.

Script & code-switching distribution

Unmeasured. No per-script counts were published and none were counted by us. The Arabic side is Arabic-script by inspection of row 0; Arabizi coverage, if any, is undocumented.

Provenance and limits

  • —The default config is authored unser translation pairs and inherits Apache-2.0. The sentences are the creators', not ours.
  • —Per-row provenance is missing: rows carry no source id, no licence, and no retrieval date beyond the integer id.
  • —Filtration audit. Unmeasured by us. No translator count, review pass, or quality-sampling rate is published.
  • —Bounded, not cleared. Translationese and MSA-flavoured Darija cannot be excluded: no dialect-purity measurement ships with the set.
  • —Out of scope. Arabizi input, code-switched input, and document-level translation were not targeted; scoring them here would corrupt the sentence-translation measurement.
  • —Decontamination of any training corpus against this dataset is unmeasured.

Citation

bibtex
@misc{algerian_nlp_algerian_arabic_english_translation_50k,
  title  = {Algerian Arabic English Translation 50K: 50,000 Darija-English pairs},
  author = {Algerian NLP Collective},
  year   = {2026},
  url    = {https://huggingface.co/datasets/algerian-nlp/algerian-arabic-english-translation-50k}
}

No DOI, arXiv ID, or named author is published on the repo; the author field above is the collective fallback. No upstream sources to cite.

Licence

Apache-2.0 for this data. Patent grant and attribution terms apply as written in the licence text.