datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-passagebeir-corpus
BeIR Corpus
msmarco-passage-corpuswiki-ss-nqscifact
SciFact
AgentIR-dataThis dataset contains the DR-Synth generated WebShaper data for training AgentIR-4B.
Paper: AgentIR: Reasoning-Aware Retrieval for Deep Research Agents
Code: https://github.com/texttron/AgentIR
Project Page: https://texttron.github.io/AgentIR/
Model: AgentIR-4B
Dataset Details
Each instance contains:
query_id: {webshaper_query_id}_turn{i}, where webshaper_query_id is the original id in WebShaper, and i is the turn number during agent rollout when constructing the data.
query:… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/AgentIR-data.reasonir-data-hnThis is the dataset forked from https://huggingface.co/datasets/reasonir/reasonir-data with hard negatives expanded by BM25. Please follow the original License.
docmatix-ir
Docmatix-IR
Docmatix is originally a large dataset designed for fine-tuning large vision-language models on Visual Question Answering tasks. It contains a substantial collection of PDF images (2.4M) and a vast set of questions (9.5M) related to these images. However, many of the questions in the Docmatix dataset are not suitable for open-domain question answering.
To address this, we have converted Docmatix into Docmatix-IR, a training set suitable for training document visual… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/docmatix-ir.scifact-tevatrontevatron_fiqa_qrel_testtevatron_fiqa_qrelsreason-embed-10k-single-pos-tevatron-format-bridgetevatron_fiqa_queriestevatron_fiqa_corpusreason-embed-10k-single-pos-tevatron-format
