datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llbench-dataset
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences
Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute.
LL-Bench is a large-scale, human-preference benchmark for evaluating low-level
vision restoration in the era of large generative models (LGMs). It compares
10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.SiliciclasticReservoirs
SiliciclasticReservoirs
1,000,000 synthetic 3D siliciclastic-reservoir geology cubes generated from rule-based sedimentological simulations (turbidite lobes + 6 fluvial-channel architectures + delta-fan distributary). Cubes are voxelized at (64, 64, 32) cells. Each sample carries facies, porosity, permeability, and a structured set of geological conditioning parameters.
Designed to train conditional generative models (flow matching, diffusion, etc.) of subsurface geology under… See the full description on the dataset page: https://huggingface.co/datasets/AnonymouScientist/SiliciclasticReservoirs.openbrush-anonymous-masters
OpenBrush Anonymous Masters
Unattributed works from OpenBrush-75K — anonymous old masters across centuries and styles. Useful for broad-style training without artist-specific bias.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 41,914 you actually want.
Why this subset
37% of the parent dataset is unattributed — a substantial… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-anonymous-masters.AstroAlertBench
AstroAlertBench
Dataset on the Hugging Face Hub: AnonymousUser16384/AstroAlertBench.
Vision–language benchmark on publicly available ZTF alerts brokered by ALeRCE. Each example pairs tabular metadata (ZTF-style candidate fields) with a single RGB stamp montage (science, reference, difference) as one PNG.
Dataset structure
Path
Description
data/manifest_benchmark_final.csv
Primary benchmark metadata — 1500 rows (300 per class: SN, AGN, VS, asteroid, bogus)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousUser16384/AstroAlertBench.PHOEBI
PHOEBI
Phase-contrast Optical bEnchmark for Bacterial Identification — a benchmark and framework for open-world identification of mixed bacterial cultures from optical phase-contrast microscopy.
At a glance
The release contains two subsets that share the same imaging pipeline:
Subset
Species
Combinations
Images
PHOEBI-6 (primary)
6 (bs, bt, fj, ka, mx, pf)
40
~120,000
PHOEBI-4 (legacy)
4 (b, f, k, p)
14
~14,000
Both subsets:
1024×1024×3 JPEG images at… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousResearchTiger/PHOEBI.MPCFire
FireMPC: A Pan-Canadian Wildfire Forecasting Benchmark
FireMPC is a pan-Canadian wildfire risk benchmark covering approximately one
billion hectares across all fifteen Canadian terrestrial ecozones at 1 km daily
resolution from 2000 to 2025, integrating 55 drivers across fuel, terrain,
anthropogenic, and meteorological families.
This release provides four pre-built sample caches that share the same
underlying data cube but differ in their training and test sample construction… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousData4NeurIPS/MPCFire.tactile-mnist-touch-real-single-t256-320x240Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
InSpect
InSpect
InSpect is a curated natural history collection dataset for visual insect specimen understanding. It contains digitized insect specimen images with aligned crops, hierarchical taxonomy, label-derived structured metadata, and fine-grained anatomical part annotations.
Files
specimen_benchmark_metadata.csv: main metadata table. Each row corresponds to one specimen image and includes split information, taxonomic labels, image/crop paths, and label-derived… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-dataset/InSpect.tactile-mnist-touch-syn-single-t32-64x64Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
tactile-mnist-touch-syn-single-t32-320x240Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
tactile-mnist-touch-starstruck-syn-single-t32-64x64Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
tactile-mnist-touch-real-single-t256-64x64Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
GenScale
GenScale Benchmark (Anonymized Review Release)
This repository contains an anonymized, machine-readable release of the GenScale benchmark for double-blind review.
Files
GenScale_Benchmark_anonymous_clean.json: cleaned benchmark JSON with relative paths only.
data/genscale_entries.parquet: one row per benchmark entry.
data/genscale_pairs.parquet: one row per pairwise scale evaluation unit.
data/genscale_task3_edit_plans.parquet: precise scale-correction plans for S5.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2049/GenScale.tactile-mnist-touch-starstruck-syn-single-t32-320x240Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
vision-benchmark
GravCal: Large-Scale Orientation-Diverse Dataset for IMU Gravity Calibration
NeurIPS 2026 Evaluations & Datasets Track
Dataset Description
GravCal is a large-scale dataset specifically designed for single-image IMU gravity calibration. The dataset addresses a critical gap in existing visual-inertial datasets, which exhibit severe upright-pose bias with most frames captured near canonical orientations.
Key Features
148,000+ frames with diverse camera… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-vision-bench/vision-benchmark.
