piushorn/pdf-parse-bench
PDF Parse Bench Benchmark for evaluating how effectively PDF parsing solutions extract mathematical formulas and tables from documents. We generate synthetic PDFs with diverse formatting scenarios, parse them with different parsers, and score the extracted content using LLM-as-a-Judge. This semantic evaluation approach substantially outperforms traditional metrics in agreement with human judgment. Leaderboard (2026-Q1) Results are based on two benchmark… See the full description on the dataset page: https://huggingface.co/datasets/piushorn/pdf-parse-bench.
Update eval.yaml
Upload README.md with huggingface_hub
Update citation
Upload 2026-q1-formulas-only/test.jsonl with huggingface_hub
Upload 2026-q1-tables-only/test.jsonl with huggingface_hub
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Upload eval.yaml with huggingface_hub
Upload README.md with huggingface_hub
initial commit
