datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EDGAR_FILINGS_DATASET
SFD: SEC Filings Dataset (v1)
SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation.
This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in:
The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.M4BenchCaReCoS
CaReCoS
A medical acoustic question-answering dataset for reasoning over mel spectrograms
of heart, lung, and cough sounds. Each record provides a clinical question, the
mel-spectrogram image of a recording, a ground-truth answer, and the
recording's clinical metadata.
The task is purely visual: a model receives the spectrogram image together with the
question and must reason over the spectrogram to produce the answer. The raw audio is
not used as model input - the original .wav… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-dataset-1/CaReCoS.Multi-doc-2025
Dataset Card for Multi-Doc-2025
Dataset Summary
Multi-Doc-2025 is a financial question-answering benchmark built from SEC Form 10-K annual reports of S&P 500 companies. It is designed to evaluate retrieval-augmented generation (RAG) and financial QA systems under three reasoning settings that are not jointly covered by existing financial QA benchmarks: cross-company reasoning, cross-year reasoning, and hybrid-modal reasoning over both text and tables.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-Team-HC-RAG/Multi-doc-2025.Agent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.jurisbenchomni
JurisBenchOmni — Model Predictions
This repository hosts the per-sample model predictions and the
analysis-ready aggregated CSVs that accompany our paper JurisBenchOmni:
A Judicial-Practice-Oriented Benchmark for Legal Omni-Modal Evaluation.
What you get here lets you read off, slice, and re-aggregate the
per-task / per-dimension / pipeline-level numbers we cite in the paper,
without having to re-run inference. Companion code that produces the
aggregated tables from the per-sample… See the full description on the dataset page: https://huggingface.co/datasets/jurisbenchomni-anonymous/jurisbenchomni.ST-Bench
ST-Bench: Spatial-Temporal Reasoning Benchmark
ST-Bench is a comprehensive benchmark dataset for training and evaluating spatial-temporal reasoning capabilities in large language models. It includes data with raw time series, text descriptions, and image visualizations.
📊 Dataset Overview
Default Data (with time_series key)
Subset
Description
Files
Total Size
ST-Align
Alignment data for initial training
3 files
~3.2GB
ST-Causal
Causal reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Time-HD-Anonymous/ST-Bench.reviewarena
ReviewArena
ReviewArena accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review.
This release is a large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of each paper PDF where available.
51,529 papers
196,099 reviews
558,785 OCR'd PDF pages (markdown inlined per… See the full description on the dataset page: https://huggingface.co/datasets/anonymousNeurIPS2026submission4281/reviewarena.MemoReason
MemoReason
MemoReason evaluates how entity familiarity affects document-grounded reasoning in language models. It pairs factual passages with fictitious variants that preserve task structure and specified reasoning operations.
The benchmark contains 100 document templates, 12 questions per template, and 109,200 question-answer examples across ten settings. Questions cover extraction, arithmetic, temporal reasoning, and inference.
Subsets
Split
Description… See the full description on the dataset page: https://huggingface.co/datasets/memoreason-anonymous/MemoReason.SciRec
SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction
This dataset contains multimodal question-answering examples grounded in
textbook figures. Records in the figure-grounded configurations are filtered to
include only examples whose referenced image files are present in this release.
Configurations
visual: 13791 figure-grounded visual questions with resolved images.
knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousNu/SciRec.physcorp-a
PhysCorp-A — Audited Physics Training Corpus (6,432 records)
The audit-clean training corpus released alongside the Physics-R1 paper (NeurIPS 2026 D&B Track submission). Subset of the 14,294-record pre-audit pool that survives the joint two-stage contamination audit against all six paper-canonical eval splits.
Lineage
14,294 pre-audit pool (PhysCorp-pre-audit) aggregating nine source families.
Construction audit (Stage-1 5-gram Jaccard ≥ 0.4 union Stage-2… See the full description on the dataset page: https://huggingface.co/datasets/physics-r1-anonymous/physcorp-a.insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.OmniMemBench
Contents
data/ — 103 benchmark samples, each with 128K/256K/512K/1M token tiers
Usage
Use the download script in the code repo:
python download_data.py
Data Format
Each data/{run_id}/benchmark_{tier}.json contains:
character_profile: persona and conversation style
multi_session_dialogues: multi-session conversation history with multimodal references
QAs: evaluation questions with ground truth answers, evidence chains, and clues… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-release-username/OmniMemBench.GRE30KDocHopQA_Dataset
DocHopQA Dataset
Paper: https://arxiv.org/abs/2508.15851
Overview
We introduce DocHop-QA, a large-scale benchmark comprising 11,379 QA instances for multimodal, multi-document, multi-hop question answering. Constructed from publicly available scientific documents sourced from PubMed Central, DocHop-QA is domain-agnostic and incorporates diverse information formats, including textual passages, tables, and structural layout cues.
Unlike existing datasets, DocHop-QA does not… See the full description on the dataset page: https://huggingface.co/datasets/anonymousaaai123/DocHopQA_Dataset.MUSE-Bench
MUSE-Bench: Memory Utilization Evaluation Benchmark
Official dataset for the paper "Beyond Memorization: Benchmarking Memory
Utilization in Conversational LLM Agents."
Anonymous release. This repository is an anonymized copy provided for
double-blind peer review. Author and affiliation information is withheld
until the review process concludes.
Motivation
LLM agents increasingly rely on persistent cross-session memory to support
long-horizon and personalized… See the full description on the dataset page: https://huggingface.co/datasets/anonymous111111111/MUSE-Bench.Data-Prep-Bench
Data-Prep-Bench
Dataset Overview
This dataset is a comprehensive resource built for Supervised Fine-Tuning (SFT) and evaluation of Large Language Models (LLMs), covering six domains: Finance, Medicine, Law, Mathematics, Science, and General.
A key feature of this dataset is that we employed 12 different data generation methods (including Agent-based methods, DataFlow series, pure LLM-based generation, and a SKILL method) using multiple cutting-edge models (such as GPT-5… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-Data-Preparation-Bench/Data-Prep-Bench.submission14717_fictionalqa_reformatted_triviaqa
Reformatted TriviaQA for use alongside FictionalQA
Repository: omitted
Paper: omitted
Dataset Description
This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we create a few versions of the resulting data for use as… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_reformatted_triviaqa.IO-Bench
IO-Bench
IO-Bench is a 155-example evaluation dataset for mathematical economics reasoning. Each record contains a standalone economics question, a reference answer, a machine-comparable answer field, symbolic answer metadata where applicable, and review status metadata.
Unless otherwise noted, the dataset materials in this repository are licensed under the Creative Commons Attribution-NoDerivatives 4.0 International License (CC BY-ND 4.0). See LICENSE for details.… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-zxcvbnm/IO-Bench.CHIMERA
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation.
Total: 9,225 problems
Subjects: 8
Topics: 1,179
Overview
Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.cogbench-passages
CogBench Passages
A curated collection of 120 short academic passages across 8 subjects, used as the source text for the CogBench benchmark — an evaluation of whether LLMs can generate questions that satisfy specified Bloom's-Taxonomy cognitive levels under deterministic, code-checkable constraints.
Anonymized for double-blind review (NeurIPS 2026 Evaluations & Datasets Track).
Author and institutional metadata will be added after acceptance.
Evaluative role
This… See the full description on the dataset page: https://huggingface.co/datasets/cogbench-anonymous/cogbench-passages.physr1corp
PhysR1Corp — Closed-form RL Training Pool (2,268 records)
The closed-form (numeric / MCQ-gradable) RL training pool used by Physics-R1 (NeurIPS 2026 D&B Track submission). Carved out of PhysCorp-A (the audited 6,432-record corpus) by dropping open-ended questions, then decontaminated against MMMU-Pro Physics (−87 records) and against PhyX-mini + PhysUniBench-en at cos ≥ 0.85 (−78 additional records: 69 PhyX-mini near-duplicates and 9 PhysUniBench-en template duplicates). One… See the full description on the dataset page: https://huggingface.co/datasets/physics-r1-anonymous/physr1corp.anti-guessing-olympiad
Algorithmic Anti-Guessing Olympiad Math Benchmark
A memorization-robust olympiad math benchmark constructed via a three-power-tier algorithmic anti-guessing pipeline. The pipeline operationalizes the per-problem verification gap (answer-only accuracy minus solution-correctness accuracy across a fixed target-model set) as both a construction criterion and an evaluation metric.
Three released subsets
This dataset ships three complementary subsets to support headline… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-antiguessing-2026/anti-guessing-olympiad.sciconbench
SciConBench
SciConBench is a benchmark dataset for evaluating large language models (LLMs) on their ability to synthesize scientific conclusions.
Dataset Summary
Each record corresponds to one published Cochrane (CDSR) review article, identified by its DOI. For each review, the dataset provides:
A clinical question derived from the review's Objective section
A set of atomic facts decomposed from the review's conclusions
Atomic fact pairs mapping each original conclusion… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-sciconbench/sciconbench.SEABED
SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning
SEABED (SouthEast Asian Benchmark for Evaluating Audio Reasoning) covers
six audio-reasoning tasks: speech emotion recognition, speech affective
interpretation, dialect and language identification, dialectal speech
comprehension, prosodic ambiguity resolution, and long-form audio
reasoning. This repository releases a stratified 10% sample (541 of 5,404
records) of its QA data for anonymous peer review, so reviewers… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-1/SEABED.STRIDE-Bench
STRIDE Benchmark
Dataset Description
STRIDE Benchmark is an evaluation benchmark for assessing the behavioral realism of crowd trajectory generation and simulation models. Rather than comparing trajectories point-by-point, it evaluates whether generated trajectories exhibit behaviors consistent with a given scenario description — measuring trajectory-context consistency through decomposed behavioral questions.
Dataset Summary
The dataset is distributed as three… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1ads34/STRIDE-Bench.BRUNCH
BRUNCH: A Utility-Centric Benchmark for Evaluating Deep Research
Dataset Description
BRUNCH is a benchmark dataset for evaluating whether deep research agents can produce surveys that are genuinely useful for downstream reasoning.
The motivation behind BRUNCH is that many existing evaluations of deep research systems focus on surface-level qualities, such as whether a report cites relevant papers, whether it looks well structured, or whether it matches a reference answer.… See the full description on the dataset page: https://huggingface.co/datasets/anonymousneuripssubmission/BRUNCH.VeRA
VeRA: Reasoning Benchmarks as Executable Specifications
NeurIPS 2026 - Evaluations & Datasets Track (Anonymous Submission)
What is VeRA?
Most reasoning benchmarks are static: the same problems are reused indefinitely, enabling memorization, format exploitation, and saturation. VeRA redefines a benchmark as an executable specification - a triple (template, generator, verifier) that can be sampled post-training to produce unlimited fresh instances with programmatically… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-NeurIPS26-VeRA/VeRA.submission14717_fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Description
This dataset is a derivative of the main dataset. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for the associated paper. The primary purpose of this dataset repository is for transparency and to help… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_training_splits.submission14717_fictionalqa
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.
