datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anonymous-datasetdelibsim-bench
delibsim-bench
A benchmark of multi-agent LLM deliberation simulations. Measures the
Deliberative Reason Index (DRI) before and after a structured group
deliberation, plus per-turn AQuA discourse-quality scores, across:
11 model setups (single-model and mixed-model, with/without reasoning):
GPT-5.1, Gemini-3-Pro-Preview, DeepSeek-V3.2-Exp, Kimi-K2-Thinking, Claude
Opus 4.5.
12 policy topics (acp, auscj, bep, biobanking_wa, ccps, energy_futures,
fnqcj, forestera, fremantle… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submissions/delibsim-bench.anonymous-submission-neurips26-2831
DroneBench Dataset (clean release)
Wildfire monitoring layouts for benchmarking sensor placement and drone routing
(NeurIPS 2026 Datasets & Benchmarks anonymous submission).
Summary
Attribute
Value
Layouts in data archive
56 (49 + 7 tables23 extras)
Tables 2/3 split
471 scenarios / 12 layouts
Fire-spread
JPG in Satellite_Images_Mask/
Risk maps
burn_map.npy (oracle), static_risk_bp2024.npy (BP)
Large archives… See the full description on the dataset page: https://huggingface.co/datasets/anonymoussubmission2/anonymous-submission-neurips26-2831.AnonymousSubmission
Anonymous Submission — Series 1 (Motion Preference) demos
10 LIBERO missions × 100 successful scripted-policy rollouts each, stored as
HDF5 under demos/Mission_<id>.hdf5. The viewer below lists per-mission
metadata (preference axis, task instruction, file size, episode count); the
actual trajectory data is in the HDF5 files.
Quick-inspect samples (< 4 GB each)
For reviewers who want a fast look without downloading the full 62 GB:
Mission
Preference
Size… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSubmissionAccount/AnonymousSubmission.submission14717_fictionalqa
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.tactile-mnist-touch-real-seq-t256-320x240Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
NeurIPS-ED-2026-SubID-956
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
👁️Overview
This repository contains the official dataset of VisFactor, a novel benchmark derived from the Factor-Referenced Cognitive Test (FRCT) that digitizes 20 vision-centric subtests from established cognitive psychology assessments. Our work systematically investigates the gap between human visual cognition and state-of-the-art Multimodal Large Language Models (MLLMs).
🎯 Key… See the full description on the dataset page: https://huggingface.co/datasets/for-anonymous-submission/NeurIPS-ED-2026-SubID-956.llama-alpaca
Dataset Card for Alpaca-Llama3.1-KD
Dataset Summary
This dataset is a distilled version of the classic tatsu-lab/alpaca dataset. It utilizes Meta-Llama-3.1-8B-Instruct as an answer generation model to generate high-quality, instruction-following responses for the original 52,000 instructions.
The primary goal of this dataset is to provide a set of responses that align with the Llama 3.1 distribution for model realignment after compression
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSubmission9090/llama-alpaca.VERDICTS
VERDICTS: Verified Expert Response Dataset for Identifying Correctness over Text Style
This dataset contains 3600 raw human-expert annotations of correctness over 1200 unique LLM responses to questions from BFF-Bench and CMT-Bench.
qid
Question ID
cid
Conversation ID
turn
Turn in the conversation
model
Name of the model that generated the response
label
One of Correct, Incorrect or Not sure
timeThe approximate time, in seconds, to perform the annotation… See the full description on the dataset page: https://huggingface.co/datasets/anonymoussubmission764/VERDICTS.
