datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DDRBench_10K_trajectory
10K Agent Trajectories Dataset
Project Page | Paper | Code
Overview
This dataset contains agent trajectories from the Deep Data Research (DDR) project's 10-K financial analysis task, as presented in the paper "Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models".
DDR-Bench is a large-scale benchmark designed to evaluate "investigatory intelligence" in LLM agents—the autonomy to set goals and explore raw data without explicit queries. This… See the full description on the dataset page: https://huggingface.co/datasets/thinkwee/DDRBench_10K_trajectory.ddro-msmarco-doc-dataset-300k
DDRO — MS MARCO Top-300K Processed Dataset
This dataset contains the preprocessed MS MARCO Top-300K document corpus used to train and evaluate the DDRO generative retrieval models from:
Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval (SIGIR 2025)
Files
File
Description
Size
msmarco-docs-sents.top.300k.json
Top-300K documents selected by click frequency, with sentence tokenization (JSONL format)
~2 GB… See the full description on the dataset page: https://huggingface.co/datasets/kiyam/ddro-msmarco-doc-dataset-300k.ddro-testsets
