CoolFace
19 results

browsecomp

Tevatron /browsecomp-plus BrowseComp-Plus BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep-Research agents. The benchmark sources challenging, reasoning-intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp-Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both human-verified evidence documents… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus.textquestion-answeringn<1K37 likes43k downloads9mo agoHugging FaceTevatron /browsecomp-plus-corpus BrowseComp-Plus Project Page | Paper | Code BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep-Research agents. The benchmark sources challenging, reasoning-intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp-Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-corpus.textquestion-answering100K<n<1M18 likes30k downloads1y agoHugging Facesmolagents /browse_comptext1K<n<10K7 likes9.4k downloads1y agoHugging Facetimchen0618 /browsecomp-plus-trajectoriestabular10K<n<100K0 likes5.6k downloads6mo agoHugging Faceopenai /BrowseCompLongContext BrowseComp Long Context BrowseComp Long Context is a dataset based on BrowseComp to benchmark LLM’s capability to retrieve relevant information from noisy data in its context. It converts the agentic question answering tasks from Browsecomp into long context tasks. For each of the questions in a subset of BrowseComp, a list of urls are attached. Each url will be paired with an indicator indicating whether the content of the web page is required to answer the question or is… See the full description on the dataset page: https://huggingface.co/datasets/openai/BrowseCompLongContext.textquestion-answeringn<1K54 likes4k downloads1y agoHugging Faceannon124816 /browsecomp BrowseComp answer and judge rows This dataset contains Ergon-native sharded rollout-card exports for the paper artifact. Evidence type: One-step trace. This export contains one-step prompt/response evidence; it is not a full rollout trace. Source URL: https://github.com/openai/simple-evals License: See upstream BrowseComp/simple-evals releases Redistribution class: metadata-plus-fetch Card claim class: paper-claim Public evidence class: one_step_completion_trace Reducer paper… See the full description on the dataset page: https://huggingface.co/datasets/annon124816/browsecomp.0 likes4k downloads2mo agoHugging Face