datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LM-SimBench
LM-SimBench
Dataset Description
LM-SimBench is a large-scale training-performance profiling dataset for large language models. The dataset is collected from training runs based on the MindSpeed-LLM framework and the Ascend NPU development stack, covering multiple model families, context lengths, and distributed parallel configurations.
Each model is sampled under feasible combinations of data parallelism (DP), tensor parallelism (TP), pipeline parallelism (PP), context… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench.SIIB-Time
1- Scope
The increasing penetration of inverter-based resources (IBRs), e.g, renewable and energy storage systems, is fundamentally reshaping power grid dynamics. Unlike conventional resources, IBRs interact with the grid through power electronics operating at microsecond timescales, introducing ultrafast dynamic phenomena that conventional time-domain simulation methods, e.g., RMS techniques, fail to capture [1]. Electromagnetic transient (EMT) simulations can capture these fast… See the full description on the dataset page: https://huggingface.co/datasets/neurips26-PSML/SIIB-Time.LM-SimBench_example
LM-SimBench (Example Snapshot)
Dataset Description
This repository distributes a compact example snapshot of LM-SimBench, the structured CSV release of large-scale LLM training-performance profiling data. The snapshot is provided so reviewers and readers can inspect file layout, schemas, and representative records without downloading the multi–tens-of-gigabyte full release.
The profiling methodology, software stack, and field definitions are the same as in the complete… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench_example.SASS-Bench
SASS-Bench
Hardware-grounded benchmark for execution prediction on NVIDIA GPU assembly (SASS).
Given a compiled SASS kernel, launch configuration, raw inputs, and a list of
checkpoint queries, a model must predict exact uint32 register values at
specified instructions and the bit-pattern of the final output buffers.
Ground truth is captured from a Blackwell sm_120 GPU via NVBit dynamic
binary instrumentation. The initial release is 50 kernels with 2--3 input-regime
variants each… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips-ed-3552/SASS-Bench.causalverify-neurips2026
🎯 CausalVerify
An Execution-Grounded Benchmark for LLM Causal Inference Workflows
NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission
💡 TL;DR
A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.DIVE
Dataset Card for DIVE
This dataset contains safety ratings for image and text inputs.
It contains 1000 adversarial prompts and 5 attention check prompts
There are 35164 safety annotations from high-quality raters and 3246 safety annotations from low-quality raters
The total number of ratings in the dataset is 38410 equal of the number of rows in this dataset.
All the ratings in this dataset are provided by 707 demographically diverse raters - 637 are deemed high-quality and… See the full description on the dataset page: https://huggingface.co/datasets/neurips-dataset-1211/DIVE.LR-Transfer-Trajectory
Dataset Documentation
Overview
This dataset captures per-step training and validation metrics from training runs of a 12-layer GPT-style decoder-only transformer. Each run is stored as a single .csv file in which every row corresponds to one logged step, and several columns hold per-parameter measurements encoded as JSON.
The dataset is designed to support post-hoc analysis of:
Loss curves (train / val)
Throughput and step latency
Per-layer / per-parameter dynamics:… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026lrtransfer/LR-Transfer-Trajectory.NeurIPS-ED-2026-SubID-956
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
👁️Overview
This repository contains the official dataset of VisFactor, a novel benchmark derived from the Factor-Referenced Cognitive Test (FRCT) that digitizes 20 vision-centric subtests from established cognitive psychology assessments. Our work systematically investigates the gap between human visual cognition and state-of-the-art Multimodal Large Language Models (MLLMs).
🎯 Key… See the full description on the dataset page: https://huggingface.co/datasets/for-anonymous-submission/NeurIPS-ED-2026-SubID-956.IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
Anonymized review copy — NeurIPS 2026 Evaluations & Datasets Track, Submission 3390 (under review).
IndustryBench is a benchmark for evaluating the industrial procurement knowledge of large language models. It comprises 2,049 QA pairs grounded in Chinese national standards (GB/T) and structured industrial product records, with item-aligned renderings in Chinese, English, Russian, and Vietnamese.
Erratum note:… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips26-3390/IndustryBench.beyond_the_lab_neurips_paperNeurIPS-Dataset
NeurIPS Papers Dataset
This dataset contains information about NeurIPS conference paper submissions including peer reviews, author rebuttals, and decision outcomes across multiple years.
Files
dataset.csv: Main dataset file containing all paper submission data
Dataset Structure
The CSV file contains the following columns:
title: Paper title
paper_decision: Decision outcome (Accept/Reject with specific categories)
review_1, review_2, etc.: Peer reviews from… See the full description on the dataset page: https://huggingface.co/datasets/MajorTimberWolf/NeurIPS-Dataset.checkpoint-zoo-metadata
LowRankArena Checkpoint Zoo Metadata Catalog
This repository is a companion metadata catalog for the checkpoint zoo hosted in the
linked Hugging Face Model repository.
Each artifact-catalog record points to a checkpoint or reconstruction artifact at a fixed
source commit. This repository does not duplicate or modify the referenced artifacts, and its
CC BY 4.0 license does not apply to checkpoint files or upstream model weights.
Source commit and coverage
The full… See the full description on the dataset page: https://huggingface.co/datasets/neurips-ed2026-anon-checkpoints/checkpoint-zoo-metadata.retrieval-conditional-neurips2026
Dataset Release — Retrieval-Conditional NeurIPS 2026
This bundle accompanies the NeurIPS 2026 D&B Track submission
"To Retrieve or Not to Retrieve? Most of the Benefit is Structural, Not Semantic."
Contents
File
Config name
Description
data/per_task_outcomes.csv
per_task_outcomes (default)
Per-(backbone × env × condition × task) success/failure labels. 3,064 rows.
data/stats_per_cell.csv
stats_per_cell
54-cell aggregate success rates and pairwise contrasts.… See the full description on the dataset page: https://huggingface.co/datasets/anon123312/retrieval-conditional-neurips2026.SIIB-Time-subsetThis is a subset of the data provided at https://huggingface.co/datasets/neurips26-PSML/SIIB-Time. The subset includes the first scenario for both grid forming and grid following trajectories. See the datacard of the original dataset for more details.
GoalPrefBench
GoalPref-Bench: Goal-Preference Alignment Benchmark
GoalPref-Bench evaluates how AI systems handle conflicts between users' long-term goals and immediate preferences. The benchmark tests whether models prioritize helping users achieve their stated objectives or instead accommodate conflicting preferences that undermine those goals.
Overview
This benchmark addresses a core question in AI alignment — when a user's immediate preference conflicts with their long-term goal… See the full description on the dataset page: https://huggingface.co/datasets/neurips26-sycophancy/GoalPrefBench.
