nawaralseelawi/mizan-iraqi-arabic-benchmark
Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set Mizan is the first comprehensive, originally-authored evaluation benchmark for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2 public development set: 340 originally-authored, dually-reviewed items across two tracks (MSA baseline / Iraqi) and six axes. 📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865 🏆 Live leaderboard: https://mizan-bench.onrender.com 💻 Code… See the full description on the dataset page: https://huggingface.co/datasets/nawaralseelawi/mizan-iraqi-arabic-benchmark.
Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set
Mizan is the first comprehensive, originally-authored evaluation benchmark for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2 public development set: 340 originally-authored, dually-reviewed items across two tracks (MSA baseline / Iraqi) and six axes.
- 📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865
- 🏆 Live leaderboard: https://mizan-bench.onrender.com
- 💻 Code (Apache-2.0): https://github.com/nawaralseelawi/mizan-platform
Composition
340 items = 265 Iraqi-track + 75 MSA-track:
Answer positions in multiple-choice items are balanced and audited (chi-square ≤ 1.20, df = 3 for every track–axis group). All official documents are fully simulated; no real records or personal data. Safety items describe harm categories without instantiating harmful content.
Fields
Each line of pilot-0.2-all.jsonl is one item: item_id, track (arabic/iraqi), axis, question_format (multiplechoice / extraction / opengeneration), tier (all public_dev in this release), a payload with the question, choices/gold or document text and expected fields, and a dialect-region tag on Iraqi items.
Intended use & contamination notice
This is the development tier, released for research transparency and reproduction. A sealed private test tier ships with v1.0. If you train on this data, please say so when reporting Mizan scores.
Citation
@misc{alseelawi2026mizan,
author = {Alseelawi, Nawar S. and Aljumaily, Mustafa S.},
title = {Mizan: A National Benchmark for Evaluating Large Language Models
on Iraqi Arabic and the Iraqi Civic Context},
year = {2026},
doi = {10.5281/zenodo.22714865}
}