datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
video2mentaltoc_bench
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models
TOC-Bench is a diagnostic benchmark for evaluating whether Video Large Language Models maintain object identity, state, persistence, and temporal relations throughout a video. It focuses on object-centric phenomena including occlusion, disappearance, reappearance, repeated events, event order, temporal location, duration, conditional state, and relative movement.
Anonymous-review notice. This… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-video-benchmark/toc_bench.ship-dataset
ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning
ShipBench is a metadata-grounded vision-language benchmark on parametrically-generated ship structural drawings. Six commercial ship types × nine drawing-grounded sub-tasks × deterministic ground truth derived directly from the generator's input dictionary (no human annotation, no rule-citation labels).
Quick reference
Total candidates: 6{,}450 across 6 ship types (Tanker, VLCC, BULKC, CNTR, LNGC… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1383/ship-dataset.autocode-fresh-cf
AutoCode-RL fresh-CF
Executable training problems for AutoCode-RL: Reinforcement Learning for
Code with Verifiable Synthetic Data. A frozen GPT-5.5 setter constructs
harder and easier variants and verification packages; a separate GPT-OSS-20B
solver learns from binary program-execution rewards.
View
Problems
Description
originals
226
Source Codeforces tasks with generated verification packages
enhance
84
Harder generated variants
simplify
63
Easier generated… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1926/autocode-fresh-cf.Sympatheia-18k
Sympatheia-18k
Sympatheia-18k is an emotion-aware spoken dialogue dataset for empathetic speech synthesis research.
It contains 18,000 query–response pairs across 12 emotion categories, each accompanied by
synthesized audio and text transcripts.
Dataset Structure
Subset
Unique Queries
Responses
Description
Emotional
8,400 train / 3,600 eval
8,400 train / 3,600 eval
Emotional queries with emotionally-matched responses
Neutral
350 train / 150 eval
4,200 train /… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2222/Sympatheia-18k.repnoise_beavertailVL-DocIR
Abstract
VL-DocIR is a page-level benchmark for vision-based long-document retrieval built from 29,641 documents rendered into 388,548 page images from Wikipedia, arXiv, PubMed, and SEC proxy statements. The benchmark contains 271,760 questions over 23 domains and six query types, covering single-page, multi-page, and cross-document evidence configurations. Questions are grounded to rendered pages and HTML element identifiers, then filtered with a cleaning pipeline that targets… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-8421/VL-DocIR.Multi-doc-2025
Dataset Card for Multi-Doc-2025
Dataset Summary
Multi-Doc-2025 is a financial question-answering benchmark built from SEC Form 10-K annual reports of S&P 500 companies. It is designed to evaluate retrieval-augmented generation (RAG) and financial QA systems under three reasoning settings that are not jointly covered by existing financial QA benchmarks: cross-company reasoning, cross-year reasoning, and hybrid-modal reasoning over both text and tables.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-Team-HC-RAG/Multi-doc-2025.devbenchagentmorph-bugs-v0.1
AgentMorph
AgentMorph is a trajectory-level metamorphic testing benchmark for tool-using
LLM agents. Instead of requiring a labeled correct answer for every task,
AgentMorph mutates a task in a way that should preserve the user's intent,
reruns the agent, and checks whether the original and mutated trajectories
preserve a rule-specific invariant.
This repository is the anonymous review artifact for the AgentMorph paper. It
contains synthetic e-commerce agent trajectories… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous2535k/agentmorph-bugs-v0.1.Agent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.pushupbench
PushUpBench: Video Repetition Counting Benchmark
Project Page | Paper | GitHub
PushUpBench is a benchmark for evaluating vision-language models (VLMs) on their ability to count exercise repetitions in videos. It was introduced in the paper "PushupBench: Your VLM is not good at counting pushups". The dataset consists of 446 long-form clips (averaging 36.7s) designed to test temporal reasoning and repetition counting beyond simple pattern recognition.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymousatom/pushupbench.anonymous-datasetAgricultural_pests_and_diseases_instruction_tuning_dataSciRec
SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction
This dataset contains multimodal question-answering examples grounded in
textbook figures. Records in the figure-grounded configurations are filtered to
include only examples whose referenced image files are present in this release.
Configurations
visual: 13791 figure-grounded visual questions with resolved images.
knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousNu/SciRec.resource
AnchorSIPS 🧠
AnchorSIPS is a synthetic dataset and evaluation resource for evidence-supported psychosis-risk symptom measurement. It is workflow-aligned to an observable structured interview flow and includes transcript-linked evidence targets for benchmark evaluation. This release is intended for public anonymous-review access with research-only methodological use.
Plain-Language Summary
AnchorSIPS contains fully synthetic interview conversations. No real patient… See the full description on the dataset page: https://huggingface.co/datasets/anonymousxxxy/resource.agent-failure-dynamics
AgentHazard: Process-Centric Benchmark for AI Agent Trajectory Analysis
AgentHazard is the first benchmark designed specifically for process-level analysis of AI coding agent trajectories. Unlike existing benchmarks that evaluate only final outcomes (pass/fail), AgentHazard provides standardized edit-level annotations, hazard estimation protocols, and stopping-policy evaluation tasks with unified evaluation across 85,050 trajectories from 6+ agent families.
Why… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousblind/agent-failure-dynamics.CTSpinoPelvic1K
CTSpinoPelvic1K
A fused spine + pelvis 3D CT segmentation dataset built by patient-level
crosswalk between three public sources:
TCIA CT COLONOGRAPHY — DICOM CT volumes (prone + supine per patient)
CTSpine1K (COLONOG subset) — VerSe-convention vertebral label masks
CTPelvic1K dataset2 — sacrum + bilateral hip label masks
Annotations are placed onto the TCIA CT volume with the highest bone
coverage (HU > 200), separately per anatomy. For ~650 patients both
annotations land on the… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips-ED/CTSpinoPelvic1K.Mol2Pro-Binder-DatasetBinder-Dataset as defined in "Generalise or Memorise? Benchmarking Ligand-Conditioned Protein Generation from Sequence-Only Data"
Our dataset is curated from the following sources:
BindingDB (Liu et al., 2025)
Drug Target Commons (DTC) (Tanoli et al., 2018)
AlphaFill (Hekkelman et al., 2023)
BioLiP (Yang et al., 2012)
Citation
Anonymised for double-blind review
Instruction_recall_dataset
CanaryBench-PII
Frequency-aware canary injection benchmark for auditing memorization
in finetuned language models, built on the AI4Privacy PII reconstruction
task.
Dataset Description
This dataset is part of CanaryBench, a benchmark for evaluating
memorization in finetuned language models across repetition tiers
and privacy regimes.
Frequency tiers: 1×, 10×, 50×
PII types: EMAIL, PHONE
Member canaries: 770
Reference canaries: 1000
Tasks: PII detection, secret… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/Instruction_recall_dataset.lig-rebuttal-data
LIG Rebuttal Data
Anonymous data release for COLM 2026 paper: The Latent Intelligence Gap
Tiers
Tier
Contents
Size
Sufficient for
1 Results
Aggregate JSONs
~5 MB
Verify every number in rebuttal
2 Matched
Per-problem trajectories (K=32)
~700 MB
Reproduce all baselines
3 Embeddings
Last-layer hidden states
~5 GB
Reproduce IWC, verifier
Models
Model
Parameters
Benchmarks
Qwen2.5-7B-Instruct
7B
GSM8K, GPQA, AIME24, AIME25… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousblind/lig-rebuttal-data.MemGUI-3K
MemGUI-3K
Anonymous Project Page | Anonymous Code | Model
MemGUI-3K is a memory-intensive mobile GUI agent trajectory dataset for training and analyzing agents that proactively manage long-horizon context. It contains teacher rollouts from MemGUI-Agent using the ConAct Context-as-Action paradigm, where the agent emits both GUI actions and context actions for history folding and UI memory management.
Code, data processing scripts, model training scripts, and evaluation tools are… See the full description on the dataset page: https://huggingface.co/datasets/memgui-agent-anonymous/MemGUI-3K.OmniMemBench
Contents
data/ — 103 benchmark samples, each with 128K/256K/512K/1M token tiers
Usage
Use the download script in the code repo:
python download_data.py
Data Format
Each data/{run_id}/benchmark_{tier}.json contains:
character_profile: persona and conversation style
multi_session_dialogues: multi-session conversation history with multimodal references
QAs: evaluation questions with ground truth answers, evidence chains, and clues… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-release-username/OmniMemBench.arxiv-metadataHalluCompass
HalluCompass
A direction-aware diagnostic benchmark and protocol for vision-language model (VLM) hallucination evaluation
NeurIPS 2026 Datasets & Benchmarks Track
Unified release: 2,200 images across MS-COCO + AMBER + NoCaps + VizWiz, 10,000 queries, two-pass annotation (GPT-4o-mini 1st-pass + 5-human verification, Fleiss' κ = 0.72 inter-annotator agreement on a stratified 250-image validation subset, substantial agreement per Landis & Koch 1977). A balanced POPE-compatible… See the full description on the dataset page: https://huggingface.co/datasets/anonymous80934/HalluCompass.RIG-Bench
RIG-bench
Anonymous submission to the NeurIPS 2026 Evaluations & Datasets (E&D) Track.
A benchmark for reasoning-driven image generation: given visual context (images + instruction + optional demonstration pairs), the model must produce the answer as a single image.
2,000 samples
4 task families × 11 subtasks
~1.4 GB
Files
RIG-bench/
├── README.md
├── samples.jsonl # 2,000 records
└── images/<sample_id>/
├── input_<order>.<ext>
├──… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-RIG-bench/RIG-Bench.deka-retrieval
Loading
from datasets import load_dataset
corpus = load_dataset("anonymousresearch123/deka-retrieval", "corpus", split="train")
queries = load_dataset("anonymousresearch123/deka-retrieval", "queries", split="train")
labels = load_dataset("anonymousresearch123/deka-retrieval", "labels", split="train")
``
insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.mypcbench-baselineGenixerForShikra-DatasetsPaper: Genixer: Empowering Multimodal Large Language Model as a Powerful Data Generator (ECCV 2024)
Arxiv: https://arxiv.org/abs/2312.06731
Description: syn_lcs_filtered60.jsonl and syn_sbu_filtered60.jsonl are two synthetic datasets produced by our Genixer_S model for advancing grounding-based multimodal understanding.
