datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TSP_EXECUTION_RUNSllbench-dataset
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences
Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute.
LL-Bench is a large-scale, human-preference benchmark for evaluating low-level
vision restoration in the era of large generative models (LGMs). It compares
10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.LM-SimBench
LM-SimBench
Dataset Description
LM-SimBench is a large-scale training-performance profiling dataset for large language models. The dataset is collected from training runs based on the MindSpeed-LLM framework and the Ascend NPU development stack, covering multiple model families, context lengths, and distributed parallel configurations.
Each model is sampled under feasible combinations of data parallelism (DP), tensor parallelism (TP), pipeline parallelism (PP), context… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench.High_Dimensional_Time_Series
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Time-HD-Anonymous/High_Dimensional_Time_Series.M4BenchNeapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.egomonth-dataset
EgoMonth Dataset
Overview
EgoMonth is a month-level egocentric video question-answering benchmark for evaluating long-term spatiotemporal memory in multimodal large language models. The dataset focuses on daily-life first-person videos and QA tasks that require temporal indexing, spatial grounding, multi-video reasoning, and long-horizon memory.
This repository provides QA metadata, structured annotations, representative anonymized sample videos, and baseline… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-egomonth/egomonth-dataset.anonymous-working-histories
Structured Anonymous Career Paths extracted from Resumes
Dataset Summary
This dataset contains 2164 anonymous career paths across 24 differend industries.
Each work experience is tagger with their corresponding ESCO occupation (ESCO v1.1.1).
Languages
We use the English version of ESCO.
All resume data is in English as well.
Dataset Structure
Each working history contains up to 17 experiences.
They appear in order, and each experience has a title… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/anonymous-working-histories.MoleculeCLA
Overview
We present MoleculeCLA: a large-scale dataset consisting of approximately 140,000 small molecules derived from computational ligand-target binding analysis, providing nine properties that cover chemical, physical, and biological aspects.
Aspect
Glide Property (Abbreviation)
Description
Molecular Characteristics
Chemical
glide_lipo (lipo)
Hydrophobicity
Atom type, number
glide_hbond (hbond)
Hydrogen bond formation propensity
Atom type, number
Physical… See the full description on the dataset page: https://huggingface.co/datasets/anonymousxxx/MoleculeCLA.LM-SimBench_example
LM-SimBench (Example Snapshot)
Dataset Description
This repository distributes a compact example snapshot of LM-SimBench, the structured CSV release of large-scale LLM training-performance profiling data. The snapshot is provided so reviewers and readers can inspect file layout, schemas, and representative records without downloading the multi–tens-of-gigabyte full release.
The profiling methodology, software stack, and field definitions are the same as in the complete… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench_example.video_saliency
SalTempto — Saliency Eye-Tracking Dataset
A video eye-tracking dataset of 224 naturalistic video stimuli with gaze recordings from multiple subjects, designed for training and evaluating visual saliency models.
Dataset Overview
Videos
224 (1920×1080, ~30 fps)
Train / Val / Test
204 / 10 / 10
Subjects per video
Train: 1–3 (mean ≈ 2.5) · Val: 15–16 · Test: held out
Gaze recordings for the test split are intentionally held out as a hidden… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-saltempto-submission/video_saliency.AU8go-mo-dataset
GO-MO, a massive Graph agumented Open urban MObility dataset
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).
Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.Ambig-DS-T
Ambig-DS-T: Target Ambiguity Benchmark
A benchmark for measuring how well data-science agents handle ambiguous prediction targets in tabular Kaggle competitions.
Each task is a Kaggle competition derived from DSBench. For every task we provide two prompt variants — one in which the target column is named, and one in which the target is hidden behind two candidate columns. The agent must select and predict the true target; submissions are graded by the original competition metric… See the full description on the dataset page: https://huggingface.co/datasets/anonymous222bit/Ambig-DS-T.prosite_functional_motif_scaffolding_benchmark
PROSITE-derived Functional Motif Benchmark
This archive contains an anonymized dataset artifact for a systematically derived benchmark of structurally conserved functional motif-scaffolding cases from PROSITE-linked experimental protein structures.
The benchmark is intended for static motif-scaffolding evaluation with standard MotifBench-style pipelines. Cases are derived from PROSITE motif-pattern entries, mapped to experimentally resolved PDB structures, filtered for recurrent… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-motif-scaffolding/prosite_functional_motif_scaffolding_benchmark.ChemSafetyBench
Dataset Card for ChemSafetyBench
Dataset Summary
ChemSafetyBench is a regulatory-grounded benchmark dataset of 32,614
chemical substances for multi-label GHS (Globally Harmonized System)
hazard prediction and LLM safety reliability evaluation. Unlike prior
molecular benchmarks constructed by querying pharmaceutical databases,
ChemSafetyBench is seeded from a curated hazardous materials registry,
ensuring coverage of real-world industrial and safety-critical chemicals… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-07/ChemSafetyBench.evocodebench
EvoDev-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions
This is the anonymous review bundle for EvoDev-Bench. It separates the benchmark,
the anonymized source repository, the technical blog, the submitted-paper results,
and the supplemental official-format results so that each evidence source can be
inspected independently.
Start here
Material
What it contains
Entry point
Benchmark
26 task chains and 227 evaluated steps in Harbor's… See the full description on the dataset page: https://huggingface.co/datasets/anonymousee8/evocodebench.jurisbenchomni
JurisBenchOmni — Model Predictions
This repository hosts the per-sample model predictions and the
analysis-ready aggregated CSVs that accompany our paper JurisBenchOmni:
A Judicial-Practice-Oriented Benchmark for Legal Omni-Modal Evaluation.
What you get here lets you read off, slice, and re-aggregate the
per-task / per-dimension / pipeline-level numbers we cite in the paper,
without having to re-run inference. Companion code that produces the
aggregated tables from the per-sample… See the full description on the dataset page: https://huggingface.co/datasets/jurisbenchomni-anonymous/jurisbenchomni.Ambig-DS-M
Ambig-DS-M: Metric Ambiguity Benchmark
A benchmark for measuring how well ML engineering agents handle ambiguous evaluation metrics in Kaggle-style competitions.
Each task is a Kaggle competition from MLE-bench (OpenAI, 2024). For every task we provide two prompt variants — one in which the true evaluation metric is named, and one in which it is redacted. The agent must produce a submission CSV that is graded against the true metric using MLE-bench's grading infrastructure.
The… See the full description on the dataset page: https://huggingface.co/datasets/anonymous222bit/Ambig-DS-M.AptaBench_dataset
AptaBench
AptaBench is a benchmark for aptamer–small-molecule interaction prediction. It contains curated DNA/RNA aptamer–ligand pairs with standardized sequences, canonical SMILES, experimentally grounded active/inactive labels, quantitative affinity values where available, and fixed leakage-aware evaluation splits.
This repository is provided for anonymous peer review. Author identities, affiliations, acknowledgements, citation information, and non-anonymous project… See the full description on the dataset page: https://huggingface.co/datasets/aptabench-anonymous/AptaBench_dataset.SkillGenBench
SkillGenBench
SkillGenBench is a benchmark for evaluating LLM skill generation from explicit repository- and document-grounded corpora. Each benchmark instance exposes visible generation materials and an instance-specific evaluation bundle. The official v1 task set contains 187 enabled tasks across three source types: 123 Code Repo tasks, 28 Code Doc tasks, and 36 Domain Knowledge Doc tasks.
Repository Layout
data/task_manifest.csv: the Hugging Face-loadable task index.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-skillgenbench/SkillGenBench.econ_eval
The Price of Progress: Benchmark-Level LLM Inference Cost Dataset
Dataset Summary
This dataset combines historical LLM inference prices with benchmark performance scores to construct the largest publicly available benchmark-level LLM price dataset we are aware of. It covers 100+ models across three major benchmarks (GPQA-Diamond, SWE-bench Verified, and AIME) over a two-year window from April 2024 to April 2026, with varying coverage per benchmark.
The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-noname/econ_eval.timewarp
The TimeWarp Dataset
Tasks for the TimeWarp benchmark, which
evaluates the robustness of web agents to temporal changes in web UI across three
environments (Wiki, News, Shop) rendered in six UI eras.
Files
File
Rows
Description
train.csv
128
Human-facing tasks: Set, Goal, Answer, Plan, Visual_Information.
test.csv
103
Same schema, held-out split.
data.json
231
The full runnable benchmark (train + test): each task carries its intent, sites, start_url… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-1827/timewarp.GRE30Kgenerative-world-renderer-clips
A Scaling Recipe for Generative World Renderer
NeurIPS 2026 Evaluations & Datasets Track — anonymous submission.
Reviewer Sample — NeurIPS 2026 Submission.
This repository currently hosts a 40-clip flattened reviewer sample (≈ 2.8 GB) so that NeurIPS 2026 reviewers can inspect data quality and format without per-request gated access. The full dataset described in the accompanying paper — approximately 4 M frames, 40 hours of playtime, 720p / 30 FPS, ~11 k sub-clips, with five… See the full description on the dataset page: https://huggingface.co/datasets/anonymous111111/generative-world-renderer-clips.sen_legal_graphrag_dataSafeChem
Dataset Card for SafeChem
Dataset Summary
SafeChem is a regulatory-grounded benchmark dataset of 32,211
chemical substances for multi-label GHS (Globally Harmonized System)
hazard prediction and LLM safety reliability evaluation. Unlike prior
molecular benchmarks constructed by querying pharmaceutical databases,
SafeChem is seeded from a curated hazardous materials registry,
ensuring coverage of real-world industrial and safety-critical chemicals
including solvents… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-07/SafeChem.ms_marco_cocondensertoxifrench-anonymous
ToxiFrench: Large-Scale French Toxicity Dataset
Author: Phantom-Researcher
Affiliations: Anomymized for peer review
Email: anonymized for peer review
Dataset Overview
While English toxicity detection is well-established, French models often lack a deep grasp of cultural nuances and coded toxicity. ToxiFrench provides the necessary data to bridge this gap.
Total Examples: 53,622 native French comments (2011–2025).
Core Feature: Rich Chain-of-Thought (CoT) explanations… See the full description on the dataset page: https://huggingface.co/datasets/phantom-researcher/toxifrench-anonymous.AnonymousSubmission
Anonymous Submission — Series 1 (Motion Preference) demos
10 LIBERO missions × 100 successful scripted-policy rollouts each, stored as
HDF5 under demos/Mission_<id>.hdf5. The viewer below lists per-mission
metadata (preference axis, task instruction, file size, episode count); the
actual trajectory data is in the HDF5 files.
Quick-inspect samples (< 4 GB each)
For reviewers who want a fast look without downloading the full 62 GB:
Mission
Preference
Size… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSubmissionAccount/AnonymousSubmission.
