datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
document-visual-retrieval-test
Model Card: Document Visual Retrieval Test (internal)
Dataset Overview
This dataset is designed to evaluate the performance of visual retrievers by testing their ability to match a query to a relevant image. Each of the three examples in this dataset contains a text query and an associated image, which is a scanned page from the foundational "Attention is All You Need" paper. The purpose of this dataset is to facilitate the evaluation of visual retrievers, where the… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/document-visual-retrieval-test.Visual_information_retrieval
GDZ Scientific Document Retrieval Benchmark
A needle‑in‑a‑haystack benchmark for scientific document retrieval, built from historical volumes of the Göttinger Digitalisierungszentrum (GDZ). This dataset explicitly adapts the IRPAPERS methodology onto a real‑world, multilingual corpus to evaluate both text-based and visual document retrieval models.
Dataset Structure
The dataset is divided into two operational configurations:
1. queries
Contains the… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_information_retrieval.Visual_retrieval
Dataset Card for IRPAPERS
ArXiv Link: https://arxiv.org/pdf/2602.17687
Dataset Description
IRPAPERS is a collection of 166 Information Retrieval papers spanning 3,230 pages. Each page in the dataset is jointly represented as a base64 encoded string of the page image as well as an OCR-derived text transcription. IRPAPERS also contains 180 needle-in-the-haystack queries.
Retrieval Leaderboard 🔎
Rank
Retriever
Type
Recall@1
Recall@5
Recall@20… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_retrieval.
