syamjithnk/arabic-corpus-audit
Arabic Corpus Integrity Audit Author: Syamjith NK Date: 9 September 2026 · corrected 13 September 2026 Tool: arabic-lint 0.5.0 Correction, 13 September 2026. An earlier version of this card said the labels in Yousefmd/arabic_ocr_dataset were stored in visual order, and called that the full reshape + bidi signature. That was wrong. Only the shaping step ran; the words are in logical order and plain NFKC recovers them. What was measured, and stands, is that the labels store… See the full description on the dataset page: https://huggingface.co/datasets/syamjithnk/arabic-corpus-audit.
Add Zenodo DOI 10.5281/zenodo.22733934
Correct the Yousefmd finding: only shaping ran, not bidi
Correct the visual-order claim about Yousefmd/arabic_ocr_dataset; add a corrections policy
Correct the visual-order claim about Yousefmd/arabic_ocr_dataset; add a corrections policy
Measure the tokenizer cost: one stray glyph makes an Arabic word [UNK] on AraBERT and mBERT
Add severity per dataset: 16 stray, 4 partial, 1 reshaped
Arabic corpus integrity audit: 341 datasets, 276 readable, one fully corrupted
initial commit
