datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llbench-dataset
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences
Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute.
LL-Bench is a large-scale, human-preference benchmark for evaluating low-level
vision restoration in the era of large generative models (LGMs). It compares
10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.M4BenchAU8evocodebench
EvoDev-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions
This is the anonymous review bundle for EvoDev-Bench. It separates the benchmark,
the anonymized source repository, the technical blog, the submitted-paper results,
and the supplemental official-format results so that each evidence source can be
inspected independently.
Start here
Material
What it contains
Entry point
Benchmark
26 task chains and 227 evaluated steps in Harbor's… See the full description on the dataset page: https://huggingface.co/datasets/anonymousee8/evocodebench.GRE30Kmultimodal_query_rewrites
ReVision: Visual Instruction Rewriting Dataset
Dataset Summary
The ReVision dataset is a large-scale collection of task-oriented multimodal instructions, designed to enable on-device, privacy-preserving Visual Instruction Rewriting (VIR). The dataset consists of 39,000+ examples across 14 intent domains, where each example comprises:
Image: A visual scene containing relevant information.
Original instruction: A multimodal command (e.g., a spoken query referencing visual… See the full description on the dataset page: https://huggingface.co/datasets/anonymoususerrevision/multimodal_query_rewrites.ViewRecDB-100K
ViewRecDB-100K
Dataset Summary
ViewRecDB-100K is a 3D viewpoint recommendation dataset for AI photography. It contains 100K training samples and 1K test samples. Each sample consists of a paired suboptimal and optimal image, along with the corresponding 3D viewpoint change annotation. Given a suboptimal image, the task is to predict the 3D viewpoint change toward the optimal image.
The dataset is automatically constructed from the Unsplash Full Dataset. The… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-12314/ViewRecDB-100K.InSpect
InSpect
InSpect is a curated natural history collection dataset for visual insect specimen understanding. It contains digitized insect specimen images with aligned crops, hierarchical taxonomy, label-derived structured metadata, and fine-grained anatomical part annotations.
Files
specimen_benchmark_metadata.csv: main metadata table. Each row corresponds to one specimen image and includes split information, taxonomic labels, image/crop paths, and label-derived… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-dataset/InSpect.GREval-Benchtest-csv-conversion
license: cc-by-nc-nd-4.0
reva_datasets
