datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KITAB_pdf_to_markdown_reviewed
KITAB_pdf_to_markdown_reviewed (Corrected KITAB-Bench PDF→Markdown)
Short description. A carefully reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic document OCR evaluation. We fixed ground-truth errors (hallucinated text, missing page numbers, omissions of small-font text) and standardized formatting to provide a reliable benchmark for model comparison.
TL;DR
✅ Human-verified ground truth for Arabic PDF→Markdown
✅ Removes hallucinations and fills… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/KITAB_pdf_to_markdown_reviewed.pdf-to-markdown-test-corpus
PDF-to-Markdown Test Corpus
18 small PDFs, each built to break a PDF-to-Markdown converter in one specific
way, plus the measured output of one converter against all of them.
If you are writing a converter, or choosing one, the hard part is not the happy
path. It is knowing what happens when a document has two columns, or a heading
that is only bold, or a scanned page in the middle. This corpus is meant to make
that testable in about a minute, and to give you somewhere to point… See the full description on the dataset page: https://huggingface.co/datasets/BayKolm/pdf-to-markdown-test-corpus.
