rochanaro/hf-arxiv-url-bench
arXiv URL Extraction Benchmark & Longitudinal Corpus Dataset Description This repository hosts datasets designed to assess format-specific coverage gaps in URL extraction across the arXiv corpus. The collection facilitates large-scale reproducibility studies and evaluates how different document formats (LaTeX, HTML, XML, Markdown, TXT, PNG) impact automated extraction pipelines. The repository is divided into two primary corpora: The 200-Paper Benchmark: A… See the full description on the dataset page: https://huggingface.co/datasets/rochanaro/hf-arxiv-url-bench.
0192
