datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MDPBench
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
We introduce Multilingual Document Parsing Benchmark, the first benchmark for multilingual digital and photographed document parsing. Document parsing has made remarkable strides, yet almost exclusively on clean, digital, well-formatted pages in a handful of dominant languages. No systematic benchmark exists to evaluate how models perform on digital and photographed documents across diverse scripts and… See the full description on the dataset page: https://huggingface.co/datasets/Delores-Lin/MDPBench.mdpbench-doc-ocr-sft
mdpbench-doc-ocr-sft
Qwen 4B (Qwen3-VL) MDPBench 한국어/일본어 문서 OCR/파싱 SFT 학습 데이터(공개 가능분).
재현 코드: https://github.com/sionic-ai/qwen-mdpbench-ocr
Configs
config
rows
내용
상태
ko
2,407
한국 지자체 소식지(신문형) 이미지 + Markdown GT
공개
jp
—
일본어 (추후)
예정
스키마
image (Image, bytes 임베드) — 문서 페이지
markdown (string) — GT 전사(Markdown, 표/읽기순서 보존)
source (string) — snvision / seocho / gongju
lang (string) — ko
doc_type (string) — newspaper
출처 & 라이선스… See the full description on the dataset page: https://huggingface.co/datasets/sionic-ai/mdpbench-doc-ocr-sft.
