datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-doc-2025
Dataset Card for Multi-Doc-2025
Dataset Summary
Multi-Doc-2025 is a financial question-answering benchmark built from SEC Form 10-K annual reports of S&P 500 companies. It is designed to evaluate retrieval-augmented generation (RAG) and financial QA systems under three reasoning settings that are not jointly covered by existing financial QA benchmarks: cross-company reasoning, cross-year reasoning, and hybrid-modal reasoning over both text and tables.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-Team-HC-RAG/Multi-doc-2025.Agent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.MemoReason
MemoReason
MemoReason evaluates how entity familiarity affects document-grounded reasoning in language models. It pairs factual passages with fictitious variants that preserve task structure and specified reasoning operations.
The benchmark contains 100 document templates, 12 questions per template, and 109,200 question-answer examples across ten settings. Questions cover extraction, arithmetic, temporal reasoning, and inference.
Subsets
Split
Description… See the full description on the dataset page: https://huggingface.co/datasets/memoreason-anonymous/MemoReason.SciRec
SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction
This dataset contains multimodal question-answering examples grounded in
textbook figures. Records in the figure-grounded configurations are filtered to
include only examples whose referenced image files are present in this release.
Configurations
visual: 13791 figure-grounded visual questions with resolved images.
knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousNu/SciRec.insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.OmniMemBench
Contents
data/ — 103 benchmark samples, each with 128K/256K/512K/1M token tiers
Usage
Use the download script in the code repo:
python download_data.py
Data Format
Each data/{run_id}/benchmark_{tier}.json contains:
character_profile: persona and conversation style
multi_session_dialogues: multi-session conversation history with multimodal references
QAs: evaluation questions with ground truth answers, evidence chains, and clues… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-release-username/OmniMemBench.DocHopQA_Dataset
DocHopQA Dataset
Paper: https://arxiv.org/abs/2508.15851
Overview
We introduce DocHop-QA, a large-scale benchmark comprising 11,379 QA instances for multimodal, multi-document, multi-hop question answering. Constructed from publicly available scientific documents sourced from PubMed Central, DocHop-QA is domain-agnostic and incorporates diverse information formats, including textual passages, tables, and structural layout cues.
Unlike existing datasets, DocHop-QA does not… See the full description on the dataset page: https://huggingface.co/datasets/anonymousaaai123/DocHopQA_Dataset.Data-Prep-Bench
Data-Prep-Bench
Dataset Overview
This dataset is a comprehensive resource built for Supervised Fine-Tuning (SFT) and evaluation of Large Language Models (LLMs), covering six domains: Finance, Medicine, Law, Mathematics, Science, and General.
A key feature of this dataset is that we employed 12 different data generation methods (including Agent-based methods, DataFlow series, pure LLM-based generation, and a SKILL method) using multiple cutting-edge models (such as GPT-5… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-Data-Preparation-Bench/Data-Prep-Bench.IO-Bench
IO-Bench
IO-Bench is a 155-example evaluation dataset for mathematical economics reasoning. Each record contains a standalone economics question, a reference answer, a machine-comparable answer field, symbolic answer metadata where applicable, and review status metadata.
Unless otherwise noted, the dataset materials in this repository are licensed under the Creative Commons Attribution-NoDerivatives 4.0 International License (CC BY-ND 4.0). See LICENSE for details.… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-zxcvbnm/IO-Bench.cogbench-passages
CogBench Passages
A curated collection of 120 short academic passages across 8 subjects, used as the source text for the CogBench benchmark — an evaluation of whether LLMs can generate questions that satisfy specified Bloom's-Taxonomy cognitive levels under deterministic, code-checkable constraints.
Anonymized for double-blind review (NeurIPS 2026 Evaluations & Datasets Track).
Author and institutional metadata will be added after acceptance.
Evaluative role
This… See the full description on the dataset page: https://huggingface.co/datasets/cogbench-anonymous/cogbench-passages.anti-guessing-olympiad
Algorithmic Anti-Guessing Olympiad Math Benchmark
A memorization-robust olympiad math benchmark constructed via a three-power-tier algorithmic anti-guessing pipeline. The pipeline operationalizes the per-problem verification gap (answer-only accuracy minus solution-correctness accuracy across a fixed target-model set) as both a construction criterion and an evaluation metric.
Three released subsets
This dataset ships three complementary subsets to support headline… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-antiguessing-2026/anti-guessing-olympiad.BRUNCH
BRUNCH: A Utility-Centric Benchmark for Evaluating Deep Research
Dataset Description
BRUNCH is a benchmark dataset for evaluating whether deep research agents can produce surveys that are genuinely useful for downstream reasoning.
The motivation behind BRUNCH is that many existing evaluations of deep research systems focus on surface-level qualities, such as whether a report cites relevant papers, whether it looks well structured, or whether it matches a reference answer.… See the full description on the dataset page: https://huggingface.co/datasets/anonymousneuripssubmission/BRUNCH.DECADE
DECADE: Dataset for Evolving Context And Dialogue Evaluation
DECADE is a benchmark for evaluating long-term memory reasoning in personalized conversational AI. It simulates a decade (2016–2026) of user interactions across 500 QA instances, each paired with a personal conversation history of up to 1,047 sessions.
Task
Given a user's long conversation history (haystack) and a question posed from a future date, a system must retrieve the relevant sessions and synthesize an… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-penguin/DECADE.horizonmath
HorizonMath
HorizonMath is a benchmark of research-level mathematical problems for measuring progress in reasoning toward mathematical discovery with automatic verification.
Files
data/problems_full.json
data/problems_full.jsonl
data/baselines.json
data/baselines.jsonl
croissant.json
Loading
from datasets import load_dataset
problems = load_dataset(
"anonymousAIresearcher/horizonmath",
name="problems",
split="train",
)
baselines = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/anonymousAIresearcher/horizonmath.mobiplant
Dataset Card for MoBiPlant
Dataset Summary
MoBiPlant is a multiple-choice question-answering dataset curated by plant molecular biologists worldwide. It comprises two merged versions:
Expert MoBiPlant: 565 expert-level questions authored by leading researchers.
Synthetic MoBiPlant: 1,075 questions generated by large language models from papers in top plant science journals.
Each example consists of a question about plant molecular biology, a set of answer options, and… See the full description on the dataset page: https://huggingface.co/datasets/anonymousSub10/mobiplant.backtrader-mcq-base-pool-all-strategies
Backtrader MCQ Benchmark
This dataset contains multiple-choice questions for evaluating whether a model can reason about trading-strategy behavior using the Backtrader backtesting framework. Each question provides a complete backtest configuration and asks for a single answer choice in the format <<< X >>>, where X is one of A, B, C, or D.
The primary evaluation file used in the paper is:
backtrader_mcq_balanced_30_all_strategies.jsonl
The larger supporting pool is:… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-678/backtrader-mcq-base-pool-all-strategies.TSQA
TSQA: Time-Sensitive Question Answering Benchmark
TSQA is a benchmark designed to evaluate a model’s ability to handle time-aware factual knowledge. Unlike standard static QA datasets, TSQA tests whether models can identify facts whose correct answers change over time.
Dataset Overview
Name: TSQA (Time-Sensitive Question Answering)
Years Covered: 2013–2024
Number of Questions: 10,063
Choices per Question: 4 (one correct, three distractors)
Temporal Context: Each… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-789/TSQA.
