browsecomp
Qwen3-1.7B-BrowseComp-Worker-SFT-iter1366-0916Qwen3-1.7B-BrowseComp-Worker-Judger-RL-iter149-0916Qwen3-1.7B-BrowseComp-Worker-OPD-iter149-0917Qwen3-1.7B-BrowseComp-Worker-Base-GRPO-iter149-0916Qwen3-4B-BrowseComp-Worker-SFT-0915Qwen3-4B-BrowseComp-Worker-RL-i149-0915acm-browsecompplus-qwen3.5-9b-opd-iter3Qwen3-4B-BrowseComp-Worker-RL-i129-0915
browsecomp-plus
BrowseComp-Plus
BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep-Research agents. The benchmark sources challenging, reasoning-intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp-Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both human-verified evidence documents… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus.browsecomp-plus-corpus
BrowseComp-Plus
Project Page | Paper | Code
BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep-Research agents. The benchmark sources challenging, reasoning-intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp-Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-corpus.browse_compbrowsecomp-plus-trajectoriesBrowseCompLongContext
BrowseComp Long Context
BrowseComp Long Context is a dataset based on BrowseComp to benchmark LLM’s capability to retrieve relevant information from noisy data in its context. It converts the agentic question answering tasks from Browsecomp into long context tasks.
For each of the questions in a subset of BrowseComp, a list of urls are attached. Each url will be paired with an indicator indicating whether the content of the web page is required to answer the question or is… See the full description on the dataset page: https://huggingface.co/datasets/openai/BrowseCompLongContext.browsecomp
BrowseComp answer and judge rows
This dataset contains Ergon-native sharded rollout-card exports for the paper artifact.
Evidence type: One-step trace.
This export contains one-step prompt/response evidence; it is not a full rollout trace.
Source URL: https://github.com/openai/simple-evals
License: See upstream BrowseComp/simple-evals releases
Redistribution class: metadata-plus-fetch
Card claim class: paper-claim
Public evidence class: one_step_completion_trace
Reducer paper… See the full description on the dataset page: https://huggingface.co/datasets/annon124816/browsecomp.
