datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DDRBench_10K_trajectory
10K Agent Trajectories Dataset
Project Page | Paper | Code
Overview
This dataset contains agent trajectories from the Deep Data Research (DDR) project's 10-K financial analysis task, as presented in the paper "Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models".
DDR-Bench is a large-scale benchmark designed to evaluate "investigatory intelligence" in LLM agents—the autonomy to set goals and explore raw data without explicit queries. This… See the full description on the dataset page: https://huggingface.co/datasets/thinkwee/DDRBench_10K_trajectory.DDRBench_10K
DDRBench: Deep Data Research Benchmark
📊 Leaderboard & Demo | 📄 Paper (Arxiv)
Overview
DDRBench (Deep Data Research Benchmark) is a comprehensive evaluation framework designed to assess the capabilities of Large Language Model (LLM) agents in performing complex, multi-turn data research and reasoning tasks. Unlike traditional Q&A benchmarks, DDRBench focuses on scenarios requiring deep interaction with structured databases, tool usage, and long-context reasoning.
This… See the full description on the dataset page: https://huggingface.co/datasets/thinkwee/DDRBench_10K.
