Trungdaik/Visual_retrieval
Dataset Card for IRPAPERS ArXiv Link: https://arxiv.org/pdf/2602.17687 Dataset Description IRPAPERS is a collection of 166 Information Retrieval papers spanning 3,230 pages. Each page in the dataset is jointly represented as a base64 encoded string of the page image as well as an OCR-derived text transcription. IRPAPERS also contains 180 needle-in-the-haystack queries. Retrieval Leaderboard ๐ Rank Retriever Type Recall@1 Recall@5 Recall@20โฆ See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_retrieval.
Dataset Card for IRPAPERS
ArXiv Link: https://arxiv.org/pdf/2602.17687
Dataset Description
IRPAPERS is a collection of 166 Information Retrieval papers spanning 3,230 pages. Each page in the dataset is jointly represented as a base64 encoded string of the page image as well as an OCR-derived text transcription. IRPAPERS also contains 180 needle-in-the-haystack queries.
Retrieval Leaderboard ๐
Question Answering Leaderboard ๐ฌ
Citation
Please consider citing our paper if you find this work useful:
@misc{shorten2026,
title={IRPAPERS: A Visual Document Benchmark for Scientific Retrieval and Question Answering},
author={Connor Shorten and Augustas Skaburskas and Daniel M. Jones and Charles Pierse and Roberto Esposito and John Trengrove and Etienne Dilocker and Bob van Luijt},
year={2026},
eprint={2602.17687},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/pdf/2602.17687},
}