CoolFace
20 results

deep search

google /deepsearchqa DeepSearchQA A 900-prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi-step information-seeking tasks across 17 different fields. ▶ Google DeepMind Release Blog Post▶ DeepSearchQA Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark DeepSearchQA is a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/google/deepsearchqa.textquestion-answeringn<1K132 likes26k downloads9mo agoHugging Facei-DeepSearch /observation-masking-eval-logs Eval Logs Paper | Code This repository contains model evaluation logs for four deep-research / web-agent benchmarks. Each run directory contains evaluated.jsonl judge results and node_0_shard_*.jsonl trajectory logs. Plot files and local bookkeeping files are intentionally excluded. CM denotes the observation mask context management setting used in the paired run. Data Access You can download all released evaluation data, including tasks and… See the full description on the dataset page: https://huggingface.co/datasets/i-DeepSearch/observation-masking-eval-logs.other2 likes24k downloads2mo agoHugging FaceSalesforce /MTA-Vision-DeepSearchimagen<1K0 likes1k downloads4mo agoHugging Facexbench /DeepSearch xbench-evals 🌐 Website | 📄 Paper | 🤗 Dataset Evergreen, contamination-free, real-world, domain-specific AI evaluation framework xbench is more than just a scoreboard — it's a new evaluation framework with two complementary tracks, designed to measure both the intelligence frontier and real-world utility of AI systems: AGI Tracking: Measures core model capabilities like reasoning, tool-use, and memory Profession Aligned: A new class of evals grounded in workflows, environments… See the full description on the dataset page: https://huggingface.co/datasets/xbench/DeepSearch.textn<1K13 likes729 downloads1y agoHugging Facexbench /DeepSearch-2510 xbench-evals 🌐 Website | 📄 Paper | 🤗 Dataset Evergreen, contamination-free, real-world, domain-specific AI evaluation framework xbench is more than just a scoreboard — it's a new evaluation framework with two complementary tracks, designed to measure both the intelligence frontier and real-world utility of AI systems: AGI Tracking: Measures core model capabilities like reasoning, tool-use, and memory Profession Aligned: A new class of evals grounded in workflows, environments… See the full description on the dataset page: https://huggingface.co/datasets/xbench/DeepSearch-2510.textn<1K4 likes513 downloads11mo agoHugging FaceOrnamentt /DeepSearch-World-Env2 likes460 downloads4mo agoHugging Face