SoMarkAI/DocParsingBench
DocParsingBench is a document intelligence benchmark of 1,400 images, systematically collected and annotated from real business workflows. It is the first dataset to systematically catalogue the document elements most frequently encountered in enterprise settings, covering five major domains: finance, legal, scientific research, manufacturing, and education. 🆕 Latest Updates [2026.04.17] DocParsingBench evaluation toolkit release. Unified scoring is now available for the… See the full description on the dataset page: https://huggingface.co/datasets/SoMarkAI/DocParsingBench.
docs: update latest news with evaluation toolkit release (#3)
Update markdowns with gt_2026-3-17: improve LaTeX and chemistry notation
Remove .DS_Store from tracking and add .gitignore
Add ModelScope badge (lost during rebase)
Update README: fix badges and add cross-platform download instructions
Update README.md
Update README.md
Update README.md
Update README.md
File markdowns/legal_standard_00076.md has been reverted due to sensitive file content or message
Fix image file loss
File images/research_academic_paper_00091.png has been reverted due to sensitive file content or message
File images/education_scanned_00056.png has been reverted due to sensitive file content or message
File images/education_scanned_00087.png has been reverted due to sensitive file content or message
Add *.png to Git LFS tracking and include missing images
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
File images/education_scanned_00056.png has been reverted due to sensitive file content or message
File images/research_academic_paper_00091.png has been reverted due to sensitive file content or message
File images/education_scanned_00087.png has been reverted due to sensitive file content or message
upload dataset folder to repo (batch 6/6)
upload dataset folder to repo (batch 5/6)
upload dataset folder to repo (batch 4/6)
upload dataset folder to repo (batch 3/6)
upload dataset folder to repo (batch 2/6)
upload dataset folder to repo (batch 2/6)
upload dataset folder to repo (batch 1/6)
upload dataset folder to repo (batch 1/6)
Update README.md
Update README.md
Commit .gitattributes README.md file(s) in datasets-DocParsingBench
