CoolFace
Datasetpublic

nawaralseelawi/mizan-iraqi-arabic-benchmark

Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set Mizan is the first comprehensive, originally-authored evaluation benchmark for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2 public development set: 340 originally-authored, dually-reviewed items across two tracks (MSA baseline / Iraqi) and six axes. 📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865 🏆 Live leaderboard: https://mizan-bench.onrender.com 💻 Code… See the full description on the dataset page: https://huggingface.co/datasets/nawaralseelawi/mizan-iraqi-arabic-benchmark.

sourceHugging Facecc-by-4.0updated 15d agoView on Hugging Face
1likes77downloads
Dataset Card

Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set

Mizan is the first comprehensive, originally-authored evaluation benchmark for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2 public development set: 340 originally-authored, dually-reviewed items across two tracks (MSA baseline / Iraqi) and six axes.

  • —📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865
  • —🏆 Live leaderboard: https://mizan-bench.onrender.com
  • —💻 Code (Apache-2.0): https://github.com/nawaralseelawi/mizan-platform

Composition

340 items = 265 Iraqi-track + 75 MSA-track:

AxisFormatItemsScoring
Dialect comprehension (Iraqi 50 + MSA 20)multiple choice70auto
Iraq/general knowledge (Iraqi 50 + MSA 20)multiple choice70auto
Official-document extractionstructured extraction50auto (strict field match)
Dialect generationopen generation50human-judged (pending)
MSA↔Iraqi translation (both directions)open generation50human-judged (pending)
Safety in Iraqi contextopen generation50human-judged (pending)

Answer positions in multiple-choice items are balanced and audited (chi-square ≤ 1.20, df = 3 for every track–axis group). All official documents are fully simulated; no real records or personal data. Safety items describe harm categories without instantiating harmful content.

Fields

Each line of pilot-0.2-all.jsonl is one item: item_id, track (arabic/iraqi), axis, question_format (multiplechoice / extraction / opengeneration), tier (all public_dev in this release), a payload with the question, choices/gold or document text and expected fields, and a dialect-region tag on Iraqi items.

Intended use & contamination notice

This is the development tier, released for research transparency and reproduction. A sealed private test tier ships with v1.0. If you train on this data, please say so when reporting Mizan scores.

Citation

bibtex
@misc{alseelawi2026mizan,
  author = {Alseelawi, Nawar S. and Aljumaily, Mustafa S.},
  title  = {Mizan: A National Benchmark for Evaluating Large Language Models
            on Iraqi Arabic and the Iraqi Civic Context},
  year   = {2026},
  doi    = {10.5281/zenodo.22714865}
}