deep search
deepsearchqa
DeepSearchQA
A 900-prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi-step information-seeking tasks across 17 different fields.
▶ Google DeepMind Release Blog Post▶ DeepSearchQA Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code
Benchmark
DeepSearchQA is a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/google/deepsearchqa.observation-masking-eval-logs
Eval Logs
Paper | Code
This repository contains model evaluation logs for four deep-research / web-agent benchmarks. Each run directory contains evaluated.jsonl judge results and node_0_shard_*.jsonl trajectory logs. Plot files and local bookkeeping files are intentionally excluded.
CM denotes the observation mask context management setting used in the paired run.
Data Access
You can download all released evaluation data, including tasks and… See the full description on the dataset page: https://huggingface.co/datasets/i-DeepSearch/observation-masking-eval-logs.MTA-Vision-DeepSearchDeepSearch
xbench-evals
🌐 Website | 📄 Paper | 🤗 Dataset
Evergreen, contamination-free, real-world, domain-specific AI evaluation framework
xbench is more than just a scoreboard — it's a new evaluation framework with two complementary tracks, designed to measure both the intelligence frontier and real-world utility of AI systems:
AGI Tracking: Measures core model capabilities like reasoning, tool-use, and memory
Profession Aligned: A new class of evals grounded in workflows, environments… See the full description on the dataset page: https://huggingface.co/datasets/xbench/DeepSearch.DeepSearch-2510
xbench-evals
🌐 Website | 📄 Paper | 🤗 Dataset
Evergreen, contamination-free, real-world, domain-specific AI evaluation framework
xbench is more than just a scoreboard — it's a new evaluation framework with two complementary tracks, designed to measure both the intelligence frontier and real-world utility of AI systems:
AGI Tracking: Measures core model capabilities like reasoning, tool-use, and memory
Profession Aligned: A new class of evals grounded in workflows, environments… See the full description on the dataset page: https://huggingface.co/datasets/xbench/DeepSearch-2510.DeepSearch-World-Env
