datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Visual_information_retrieval
GDZ Scientific Document Retrieval Benchmark
A needle‑in‑a‑haystack benchmark for scientific document retrieval, built from historical volumes of the Göttinger Digitalisierungszentrum (GDZ). This dataset explicitly adapts the IRPAPERS methodology onto a real‑world, multilingual corpus to evaluate both text-based and visual document retrieval models.
Dataset Structure
The dataset is divided into two operational configurations:
1. queries
Contains the… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_information_retrieval.Visual_retrieval
Dataset Card for IRPAPERS
ArXiv Link: https://arxiv.org/pdf/2602.17687
Dataset Description
IRPAPERS is a collection of 166 Information Retrieval papers spanning 3,230 pages. Each page in the dataset is jointly represented as a base64 encoded string of the page image as well as an OCR-derived text transcription. IRPAPERS also contains 180 needle-in-the-haystack queries.
Retrieval Leaderboard 🔎
Rank
Retriever
Type
Recall@1
Recall@5
Recall@20… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_retrieval.
