datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
browsecomp-plus
BrowseComp-Plus
BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep-Research agents. The benchmark sources challenging, reasoning-intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp-Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both human-verified evidence documents… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus.browsecomp-plus-corpus
BrowseComp-Plus
Project Page | Paper | Code
BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep-Research agents. The benchmark sources challenging, reasoning-intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp-Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-corpus.wiki-ss-corpusmsmarco-passagewiki-ss-corpus-newbeir-corpus
BeIR Corpus
wikipedia-nqmsmarco-passage-corpuswikipedia-triviaaudiocaps-corpuscolpali-corpus
Overview
This dataset is a transformed version of the ColPali (vidore/colpali_train_set) dataset, formatted to be compatible with Tevatron.
beirwiki-ss-nqcornstack-php-v1-tevatron-1Mbge-ir-promptbge-irscifact
SciFact
wikipedia-curatedAgentIR-dataThis dataset contains the DR-Synth generated WebShaper data for training AgentIR-4B.
Paper: AgentIR: Reasoning-Aware Retrieval for Deep Research Agents
Code: https://github.com/texttron/AgentIR
Project Page: https://texttron.github.io/AgentIR/
Model: AgentIR-4B
Dataset Details
Each instance contains:
query_id: {webshaper_query_id}_turn{i}, where webshaper_query_id is the original id in WebShaper, and i is the turn number during agent rollout when constructing the data.
query:… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/AgentIR-data.cornstack-6-langs-v1-tevatron-6Mwikipedia-wqbrowsecomp-plus-md-toc-gpt5.4-nano
BrowseComp-Plus Structured 100k Corpus
This dataset is a drop-in, official-format variant of the BrowseComp-Plus 100k corpus used in our RISE Agent experiments. It keeps the same row count, document ids, URLs, and column names as the original BrowseComp-Plus corpus, but replaces each document's text field with a structured version that adds a generated table of contents and section headings.
Files
data.parquet: the corpus in the same three-column schema as the… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-md-toc-gpt5.4-nano.pixmo-docs-corpusreasonir-data-hnThis is the dataset forked from https://huggingface.co/datasets/reasonir/reasonir-data with hard negatives expanded by BM25. Please follow the original License.
docmatix-ir
Docmatix-IR
Docmatix is originally a large dataset designed for fine-tuning large vision-language models on Visual Question Answering tasks. It contains a substantial collection of PDF images (2.4M) and a vast set of questions (9.5M) related to these images. However, many of the questions in the Docmatix dataset are not suitable for open-domain question answering.
To address this, we have converted Docmatix into Docmatix-IR, a training set suitable for training document visual… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/docmatix-ir.cornstack-python-v1-tevatron-1Mcornstack-go-v1-tevatron-1Mmsrvtt-corpusbrightThis is the dataset forked from https://huggingface.co/datasets/xlangai/BRIGHT. Please follow the original License.
webshaper-fineweb-1m-corpus
