datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pickapic_v1
Dataset Card for "pickapic_v1"
More Information needed
CrowdEvalTSP_EXECUTION_RUNSEDGAR_FILINGS_DATASET_2022_2026H1EDGAR_FILINGS_DATASET
SFD: SEC Filings Dataset (v1)
SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation.
This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in:
The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.llbench-dataset
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences
Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute.
LL-Bench is a large-scale, human-preference benchmark for evaluating low-level
vision restoration in the era of large generative models (LGMs). It compares
10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.structured-file-audit-benchmark
Paper Data Release
This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them.
Contents
datasets/
Benchmark data and per-task manifests for the three paper-facing splits.
datasets/sc_flat/data
SC-Flat is derived from DaBench, augmented with a replayable perturbation
injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.LM-SimBench
LM-SimBench
Dataset Description
LM-SimBench is a large-scale training-performance profiling dataset for large language models. The dataset is collected from training runs based on the MindSpeed-LLM framework and the Ascend NPU development stack, covering multiple model families, context lengths, and distributed parallel configurations.
Each model is sampled under feasible combinations of data parallelism (DP), tensor parallelism (TP), pipeline parallelism (PP), context… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench.High_Dimensional_Time_Series
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Time-HD-Anonymous/High_Dimensional_Time_Series.EDGAR_FILINGS_DATASET_2016_2021HABIT
HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation
⚠️ Anonymous release. Authors and institutional information are intentionally withheld. This dataset card will be updated when these details become available.
TL;DR: HABIT is a large-scale robot demonstration dataset for human-present environments, designed to teach robot policies human-aware behaviors.
Keywords: Robot Manipulation Dataset, Human-Robot Interaction, Vision-Language-Action… See the full description on the dataset page: https://huggingface.co/datasets/habit-anonymous/HABIT.Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.egomonth-dataset
EgoMonth Dataset
Overview
EgoMonth is a month-level egocentric video question-answering benchmark for evaluating long-term spatiotemporal memory in multimodal large language models. The dataset focuses on daily-life first-person videos and QA tasks that require temporal indexing, spatial grounding, multi-video reasoning, and long-horizon memory.
This repository provides QA metadata, structured annotations, representative anonymized sample videos, and baseline… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-egomonth/egomonth-dataset.anonymous-working-histories
Structured Anonymous Career Paths extracted from Resumes
Dataset Summary
This dataset contains 2164 anonymous career paths across 24 differend industries.
Each work experience is tagger with their corresponding ESCO occupation (ESCO v1.1.1).
Languages
We use the English version of ESCO.
All resume data is in English as well.
Dataset Structure
Each working history contains up to 17 experiences.
They appear in order, and each experience has a title… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/anonymous-working-histories.Agent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.autocode-fresh-cf
AutoCode-RL fresh-CF
Executable training problems for AutoCode-RL: Reinforcement Learning for
Code with Verifiable Synthetic Data. A frozen GPT-5.5 setter constructs
harder and easier variants and verification packages; a separate GPT-OSS-20B
solver learns from binary program-execution rewards.
View
Problems
Description
originals
226
Source Codeforces tasks with generated verification packages
enhance
84
Harder generated variants
simplify
63
Easier generated… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1926/autocode-fresh-cf.pnp_20260904_063146_langThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousMouse404/pnp_20260904_063146_lang.MoleculeCLA
Overview
We present MoleculeCLA: a large-scale dataset consisting of approximately 140,000 small molecules derived from computational ligand-target binding analysis, providing nine properties that cover chemical, physical, and biological aspects.
Aspect
Glide Property (Abbreviation)
Description
Molecular Characteristics
Chemical
glide_lipo (lipo)
Hydrophobicity
Atom type, number
glide_hbond (hbond)
Hydrogen bond formation propensity
Atom type, number
Physical… See the full description on the dataset page: https://huggingface.co/datasets/anonymousxxx/MoleculeCLA.OneOcean_Environment_Dataset
oneocean_public_env
This folder is prepared for Zenodo upload.
Contents
Dataset files
schema.json: variable/dimension schema
hf_sample.parquet: small table sample for Hugging Face dataset schema detection
checksums.sha256: SHA256 for all files in this folder
Summary
Time: 2025-01-01T00:00:00.000000000 -> 2025-01-31T00:00:00.000000000 (n=31)
BBox: lat[30.0,40.0], lon[-72.0,-62.0]
Resolution (deg): dlat≈0.08264462809917461, dlon≈0.08264462809917461
Vars:… See the full description on the dataset page: https://huggingface.co/datasets/anonymous321123/OneOcean_Environment_Dataset.TabularMath
TabularMath
TL;DR. 114 tabular regression tasks, each compiled from a math word
problem into a Python (generator, verifier) pair that is validated
against the original seed answer. 2,048 rows per task, integer targets
y, zero label noise. Use it to diagnose whether your tabular model can
move from fitting to computing under controlled output extrapolation.
TabularMath is a program-verified tabular benchmark that probes whether
tabular machine-learning models can move from… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-NeurIPS26-TabularMath/TabularMath.R2R_Router_Training
Anonymous Review Only
Note that this dataset is only used for anonymous review.
AU8PawnsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.board": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousMouse404/Pawns.anonymous-ride-gold-lite
RIDE Gold Lite
RIDE Gold Lite is the smaller benchmark-ready release of the RIDE dataset. It contains fixed train/test snapshot splits, a canonical evaluation table, and model-ready representations for train delay prediction on Belgian passenger railway operations.
This release mirrors the structure and prediction task of RIDE Gold Standard, but uses fewer snapshots and rows for faster inspection, development, and lower-cost experimentation.
Links
Code repository:… See the full description on the dataset page: https://huggingface.co/datasets/ano6060/anonymous-ride-gold-lite.Pexels-Pairs-Masklets-330K
Pexels-Pairs-Masklets-330K (review sample)
Per-video masklet annotations stored as Parquet shards.
This repository is a small sample released for anonymous peer review. It contains
5 shards drawn from 5 different set_* directories of the full collection, which holds
roughly 330K shards across 301 sets. Contents are unmodified; only the number of shards
is reduced.
Layout
set_0000/<video_id>.mp4.parquet
set_0075/<video_id>.mp4.parquet… See the full description on the dataset page: https://huggingface.co/datasets/anonymousML123/Pexels-Pairs-Masklets-330K.go-mo-dataset
GO-MO, a massive Graph agumented Open urban MObility dataset
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).
Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.Ambig-DS-T
Ambig-DS-T: Target Ambiguity Benchmark
A benchmark for measuring how well data-science agents handle ambiguous prediction targets in tabular Kaggle competitions.
Each task is a Kaggle competition derived from DSBench. For every task we provide two prompt variants — one in which the target column is named, and one in which the target is hidden behind two candidate columns. The agent must select and predict the true target; submissions are graded by the original competition metric… See the full description on the dataset page: https://huggingface.co/datasets/anonymous222bit/Ambig-DS-T.anonymous-ride-gold-standard
RIDE Gold Standard
RIDE Gold Standard is the full benchmark-ready release of the RIDE dataset. It contains fixed train/test snapshot splits, a canonical evaluation table, and model-ready representations for train delay prediction on Belgian passenger railway operations.
This release is intended as the primary benchmark tier for RIDE. It is used for full-scale evaluation and comparison of models under the shared RIDE prediction task and evaluation protocol.
Links
Code… See the full description on the dataset page: https://huggingface.co/datasets/ano6060/anonymous-ride-gold-standard.prosite_functional_motif_scaffolding_benchmark
PROSITE-derived Functional Motif Benchmark
This archive contains an anonymized dataset artifact for a systematically derived benchmark of structurally conserved functional motif-scaffolding cases from PROSITE-linked experimental protein structures.
The benchmark is intended for static motif-scaffolding evaluation with standard MotifBench-style pipelines. Cases are derived from PROSITE motif-pattern entries, mapped to experimentally resolved PDB structures, filtered for recurrent… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-motif-scaffolding/prosite_functional_motif_scaffolding_benchmark.CTSpinoPelvic1K
CTSpinoPelvic1K
A fused spine + pelvis 3D CT segmentation dataset built by patient-level
crosswalk between three public sources:
TCIA CT COLONOGRAPHY — DICOM CT volumes (prone + supine per patient)
CTSpine1K (COLONOG subset) — VerSe-convention vertebral label masks
CTPelvic1K dataset2 — sacrum + bilateral hip label masks
Annotations are placed onto the TCIA CT volume with the highest bone
coverage (HU > 200), separately per anatomy. For ~650 patients both
annotations land on the… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips-ED/CTSpinoPelvic1K.
