datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EDGAR_FILINGS_DATASET
SFD: SEC Filings Dataset (v1)
SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation.
This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in:
The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.autocode-fresh-cf
AutoCode-RL fresh-CF
Executable training problems for AutoCode-RL: Reinforcement Learning for
Code with Verifiable Synthetic Data. A frozen GPT-5.5 setter constructs
harder and easier variants and verification packages; a separate GPT-OSS-20B
solver learns from binary program-execution rewards.
View
Problems
Description
originals
226
Source Codeforces tasks with generated verification packages
enhance
84
Harder generated variants
simplify
63
Easier generated… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1926/autocode-fresh-cf.Agent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.reviewarena
ReviewArena
ReviewArena accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review.
This release is a large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of each paper PDF where available.
51,529 papers
196,099 reviews
558,785 OCR'd PDF pages (markdown inlined per… See the full description on the dataset page: https://huggingface.co/datasets/anonymousNeurIPS2026submission4281/reviewarena.reviewarena-eval
ReviewArena-Eval
ReviewArena-Eval is the benchmark slice that accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review. It evaluates LLMs on the LLM-as-reviewer task: given a full paper and the structured review form used at that paper's venue-year, the model must produce overall rating, confidence, sub-scores, and free-text fields that are then compared against the actual human reviews on… See the full description on the dataset page: https://huggingface.co/datasets/anonymousNeurIPS2026submission4281/reviewarena-eval.divdata
divdata
Every simulation run behind our heterogeneous-LLM social-simulation work, consolidated into
one dataset indexed by run_id and step.
1,296 simulation runs · 89,171 posts · 652,588 comments · 14,020,992 impressions · 748,285 agent activations · 32 model variants.
Agents with distinct personas post and comment on a shared message board built on
OASIS. Each agent is driven by one of ~10 different
LLMs, so a single board mixes model families. The corpus supports asking which… See the full description on the dataset page: https://huggingface.co/datasets/anonymousfileupload/divdata.insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.econ_eval
The Price of Progress: Benchmark-Level LLM Inference Cost Dataset
Dataset Summary
This dataset combines historical LLM inference prices with benchmark performance scores to construct the largest publicly available benchmark-level LLM price dataset we are aware of. It covers 100+ models across three major benchmarks (GPQA-Diamond, SWE-bench Verified, and AIME) over a two-year window from April 2024 to April 2026, with varying coverage per benchmark.
The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-noname/econ_eval.ChemSafetyBench
Dataset Card for ChemSafetyBench
Dataset Summary
ChemSafetyBench is a regulatory-grounded benchmark dataset of 32,614
chemical substances for multi-label GHS (Globally Harmonized System)
hazard prediction and LLM safety reliability evaluation. Unlike prior
molecular benchmarks constructed by querying pharmaceutical databases,
ChemSafetyBench is seeded from a curated hazardous materials registry,
ensuring coverage of real-world industrial and safety-critical chemicals… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-07/ChemSafetyBench.bird-train-gemini3-flash
Dataset Card for Think2SQL-SFT
This dataset is a distilled Supervised Fine-Tuning (SFT) dataset designed to improve the reasoning capabilities of models in Text-to-SQL tasks.
It contains high-quality reasoning traces and SQL queries generated by Gemini 3 Flash.
Paper: Think2SQL: Blueprinting Reward Density and Advantage Scaling for Effective Text-To-SQL Reasoning
Base Benchmark: BIRD-Train
Dataset Description
The dataset consists of 9,428 high-quality traces, of… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-2321/bird-train-gemini3-flash.SafeChem
Dataset Card for SafeChem
Dataset Summary
SafeChem is a regulatory-grounded benchmark dataset of 32,211
chemical substances for multi-label GHS (Globally Harmonized System)
hazard prediction and LLM safety reliability evaluation. Unlike prior
molecular benchmarks constructed by querying pharmaceutical databases,
SafeChem is seeded from a curated hazardous materials registry,
ensuring coverage of real-world industrial and safety-critical chemicals
including solvents… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-07/SafeChem.IndicMMLU-Pro
IndicMMLU Dataset
This dataset contains the following languages:
punjabi
hindi
urdu
telugu
gujrati
kannada
tamil
marathi
bengali
UPLOAD
Cite our work.
This dataset is also described in IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding.
@dataset{kj2024indicmmlupro,
author = {Kj, Sankalp and Kumar, Ashutosh and Balaji, Laxmaan and Kotecha, Nikunj and Jain, Vinija and Chadha, Aman and Bhaduri, Sreyoshi},
title =… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1069/IndicMMLU-Pro.IO-Bench
IO-Bench
IO-Bench is a 155-example evaluation dataset for mathematical economics reasoning. Each record contains a standalone economics question, a reference answer, a machine-comparable answer field, symbolic answer metadata where applicable, and review status metadata.
Unless otherwise noted, the dataset materials in this repository are licensed under the Creative Commons Attribution-NoDerivatives 4.0 International License (CC BY-ND 4.0). See LICENSE for details.… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-zxcvbnm/IO-Bench.Strudel-Synth
Strudel-Synth
Strudel-Synth is a synthetic corpus of 21,174 (MIDI, Strudel) pairs for
training and evaluating MIDI-to-Strudel decompilation, introduced in
Decomposer: Learning to Decompile Symbolic Music to Programs.
🎹 Live demo: anonymousgiraffe/decomposer-demo
🤗 Model: anonymousgiraffe/Decomposer-Qwen3-8B
Each pair consists of a Strudel program distilled from Claude-Opus-4.6 (conditioned on independently sampled musical and code-style seeds) and the MIDI produced by… See the full description on the dataset page: https://huggingface.co/datasets/anonymousgiraffe/Strudel-Synth.StereoTales
Multilingual Story-Generation Bias Samples
A multilingual evaluation dataset for probing demographic biases in LLM
story generation. Each sample instructs a model to write a ~200-word story
about a character carrying a given demographic attribute value (age, gender,
ethnicity, religion, disability status, immigration status, ...) placed into a
specific life scenario, with the goal of surfacing socio-economic and
demographic biases in the generated narratives.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-authors/StereoTales.STRIDE-Bench
STRIDE Benchmark
Dataset Description
STRIDE Benchmark is an evaluation benchmark for assessing the behavioral realism of crowd trajectory generation and simulation models. Rather than comparing trajectories point-by-point, it evaluates whether generated trajectories exhibit behaviors consistent with a given scenario description — measuring trajectory-context consistency through decomposed behavioral questions.
Dataset Summary
The dataset is distributed as three… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1ads34/STRIDE-Bench.submission14717_fictionalqa
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.straj_linuxarena
straj_linuxarena (public-env subset)
Adversarial sabotage agentic SWE benchmark trajectories from the linuxarena project, run as part of the no-CoT time-horizons paper.
The dataset viewer above shows the per-cell outcome records (precomputed_results.csv, 257 rows). Full Inspect .eval trajectories for the 13 public-environment tasks (~2.4 GB, 150 files) are stored under evals/ — see "Inspect trajectories" below.
Per-cell schema
Each row in precomputed_results.csv (and… See the full description on the dataset page: https://huggingface.co/datasets/anonymouslinuxarena/straj_linuxarena.Narrative-Infilling
Dataset Card for NarrativeInfilling Benchmark
Dataset Details
Dataset Description
The Narrative Infilling Benchmark is a large-scale evaluation dataset for
narrative infilling the task of generating a missing span within a
narrative while maintaining consistency with both the preceding and following
context. The benchmark spans four narrative domains and contains 9,142 instances
with systematic variation in blank position and span length, enabling… See the full description on the dataset page: https://huggingface.co/datasets/anonymous12341952/Narrative-Infilling.prompt-sensitivity-codegen
Anonymous Prompt Sensitivity Dataset
This package contains model generations and evaluation outcomes for an anonymized
submission on prompt sensitivity in few-shot code generation.
What is included
prompt_sensitivity_dataset.jsonl: one row per generated sample
prompt_sensitivity_dataset.csv: tabular view of the same rows
prompt_sensitivity_dataset.parquet: columnar copy when parquet support is available
prompt_variant_spec.json: machine-readable description of the prompt… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl26/prompt-sensitivity-codegen.
