datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nrvbench-review
NR Video Editing Benchmark
This repository contains two non-rigid video editing benchmark subsets for evaluating instruction-driven video editing methods. Each row in metadata.csv corresponds to one editing instruction for a source video, with relative paths to the source video, extracted frames, binary masks, prompts, and evaluation questions.
The dataset card is written without author or institution identifiers so it can be used for anonymous review uploads. Before a non-anonymous… See the full description on the dataset page: https://huggingface.co/datasets/NRVBench/nrvbench-review.Amazon-Reviews-DatasetThis dataset provides a free trial sample of best-selling products and their customer reviews from a leading e-commerce platform, designed to support product intelligence, sentiment analysis, and market trend evaluation. This sample is provided for evaluation purposes only. It includes a curated subset of the full dataset.
To access the complete dataset, request additional attributes, or explore alternative product segments, please contact the data provider directly.
Key Features
2… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/Amazon-Reviews-Dataset.bangla-english-and-code-mixed-ecommerce-review-dataset
BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce
Description
The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.review-dataset
MSIR-Bench Review Dataset
This repository contains an anonymized review snapshot of MSIR-Bench, a benchmark for identity-preserving style image retrieval.
Dataset Description
Each source identity is represented by an anonymous five-digit ID. Images are organized by split and identity folder. File names follow either <id>_<Style>.png, <id>_original.png, or legacy original.jpg for original reference images.
The dataset is intended for evaluating whether a retrieval… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-review-dataset-2026/review-dataset.agent_paper_reviewdataset_for_review UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
👀 UNO-Bench Overview
Multimodal Large Languages models have been progressing from uni-modal understanding toward unifying visual, audio and language modalities, collectively termed omni models. However, the correlation between uni-modal and omni-modal remains unclear, which requires comprehensive evaluation to drive omni model's intelligence evolution. In… See the full description on the dataset page: https://huggingface.co/datasets/blue-tundra-42/dataset_for_review.multimodal-product-reviews-lazada
A Multimodal Product Reviews Dataset Collected from Lazada (2024)
Description
This dataset comprises:
Product Information: including both product images and textual descriptions
User-Generated Reviews: including review images and textual reviews
Language: Vietnamese
The data was collected from the Lazada e-commerce platform in 2024.
Example
{
"product_id": "1086202",
"product_information": {
"product_images": [… See the full description on the dataset page: https://huggingface.co/datasets/trucmtnguyen/multimodal-product-reviews-lazada.segvqa-review
SegVQA — Human-Validation Review Set (Kvasir-SEG, n=100)
Human-validation subset of the SegVQA Kvasir-SEG benchmark
(segvqa_benchmark_kvasir.json), 100 verified items sampled with seed 42,
stratified by category (multi-hop/comparison slightly oversampled).
Category
n
morphology
16
multi_hop
18
localization
16
counting
16
comparison
18
existence_detection
16
Review protocol
For each row, mark:
Q — is the question sensible and answerable… See the full description on the dataset page: https://huggingface.co/datasets/AiventraLab/segvqa-review.Generation-Reviewqiuli-collected-works-ocr-reviewed-20260906
裘李全集逐页审核输出
当前只有以下 2页 按百度首扫、千问疑难/全页盘点、GPT本地原图全文审核分工验收并冻结。千问盘点不等于千问全文校对。
LXQ-01 PDF第6页:revision3,原冻结包保留;PDF7仅跨页证据。
QXG-4 PDF第30页(书页25):revision3;PDF31仅跨页证据。
逐页canonical JSON是唯一当前真值,原始证据、模型版本/SHA与绝对路径映射均保留。源PDF仍在独立source-only仓。本仓不表示36册均审核完成,不额外授予原书版权许可。
locus-bench
LOCUS-Bench (anonymized release for peer review)
A benchmark for embodied multi-robot task planning that grades two
difficulties separately. Axis S (state judgment): S0 no judgment; S1 whether
a single named target is already in place; S2 which of 2 to 3 conditional
candidates are absent; S3 which of 3 to 6 quantified instances are unsatisfied;
S4 whether invisible implies absent under occlusion, with single-frame fallback
planning. Axis M (mechanical structure): M0 none; M1 an… See the full description on the dataset page: https://huggingface.co/datasets/review-artifacts/locus-bench.skin-gan-review-baseline
Dataset Card for dermatological diagnoses
This dataset is intended for the study of skin lesion images.
Its main goal is to support research on dermatological image synthesis, data augmentation, and the improvement of models for skin lesion analysis and computer-aided diagnosis.
This resource is designed to support comparative studies on synthetic image generation, class balancing, and dataset expansion in the dermatology domain, contributing to the development of artificial… See the full description on the dataset page: https://huggingface.co/datasets/z72pepee/skin-gan-review-baseline.amazon-reviewsskin-gan-review-224x224
Dataset Card for dermatological diagnoses
This dataset is intended for the study of GAN-based methods applied to the generation and augmentation of skin lesion datasets.
Its main goal is to support research on dermatological image synthesis, data augmentation, and the improvement of deep learning models for skin lesion analysis and computer-aided diagnosis.
This resource is designed to support comparative studies on synthetic image generation, class balancing, and dataset… See the full description on the dataset page: https://huggingface.co/datasets/z72pepee/skin-gan-review-224x224.review
🛡️ LabShield: A Multimodal Benchmark for Laboratory Safety
Official dataset for the paper: "LabShield: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories".
[Project Website (Coming Soon)] | [Paper (NeurIPS 2026 Submission)] | [Code (Coming Soon)]
📌 Introduction
LabShield is a rigorous, multi-view benchmark designed to assess the safety awareness and decision-making reliability of Multimodal Large Language Models (MLLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/LabShield-Review/review.Banana_reviewer3amazon-reviews-10MNL3D-Review-Sample
NL3D anonymized review sample
This repository contains an anonymized review-time sample of the NL3D real benchmark. It is intended to let reviewers inspect the data organization and annotations during the double-blind review period; it is not the complete real benchmark.
The sample contains two metallic and two transmissive scenes:
real_sample/metal/scene_fish1
real_sample/metal/scene_statue
real_sample/transmissive/scene_bird
real_sample/transmissive/scene_banana
Each scene… See the full description on the dataset page: https://huggingface.co/datasets/NL3D/NL3D-Review-Sample.Gpt1_reviewer4AU8review
🛡️ LabShield: A Multimodal Benchmark for Laboratory Safety
Official dataset for the paper: "LabShield: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories".
[Project Website (Coming Soon)] | [Paper (NeurIPS 2026 Submission)] | [Code (Coming Soon)]
📌 Introduction
LabShield is a rigorous, multi-view benchmark designed to assess the safety awareness and decision-making reliability of Multimodal Large Language Models (MLLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/LabShield/review.womens-clothing-reviews-eda
Women's Clothing E-Commerce Reviews – EDA Project
This project explores customer reviews from a women's clothing e-commerce store.
The goal was to clean, analyze, and visualize the data to uncover key insights about customer satisfaction and behavior.
Dataset Overview
Source: Public dataset of women’s clothing reviews
Size: 23,486 reviews and 11 features
Target variable: Recommended IND
Data Cleaning
Removed index column
Dropped duplicates
Handled missing… See the full description on the dataset page: https://huggingface.co/datasets/Oriminkowski/womens-clothing-reviews-eda.review-dataset-index
MANUS-Bench — anonymous release bundle (WACV 2027 submission #1065)
MANUS-Bench (Multimodal Annotated Naturalistic Hand Understanding) is a
benchmark for geometry-conditioned hand-gesture synthesis in natural scenes:
aligned full-scene and localized hand conditions over an official
condition-complete protocol of 88,294 samples (64,931 train / 23,363 test)
drawn from an 88,312-record archive, with fixed in-domain, near-domain, and
hand-object OOD split roles.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/anon-review-artifact-7k3p9/review-dataset-index.adpd-review
Agile Drone Pose Dataset (ADPD)
ADPD is an indoor stereo benchmark for close-range agile drone trajectories.
It provides synchronized stereo images, MoCap-derived 6-DoF annotations,
official sequence-level splits, detector annotations, and camera calibration.
It does not require LiDAR, point clouds, or dense scene reconstruction.
Dataset Summary
Split
Sequences
Stereo pairs
Images
Train
29
13,517
27,034
Validation
8
3,621
7,242
Test
4
1,700
3,400… See the full description on the dataset page: https://huggingface.co/datasets/adpd-anonymous-review/adpd-review.dgfe_reviewtdbench-review
TDBench: Benchmarking Vision-Language Models on Top-Down Images
Note (Anonymous Review Version). This dataset card accompanies a NeurIPS 2026 Evaluations & Datasets double-blind submission. Author identifiers, institutional affiliations, project pages, and external repository links have been removed for the review period. The full set of public artifacts and the final citation will be restored upon decision.
Overview
TDBench is a benchmark for evaluating… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-tdbench/tdbench-review.HEPROBench-CRC-CODEX-review-demo
HEPROBench CRC-CODEX reviewer demo
This is a small real-data software-verification subset for HEPROBench. It
contains paired, registered H&E and CRC-CODEX patches derived from Schürch et
al., Coordinated cellular neighborhoods orchestrate antitumoral immunity at
the colorectal cancer invasive front, created by Christian Schürch, Mendeley Data
10.17632/mpjzbtfgfr.1, licensed under
CC BY 4.0.
The bundle has 2 patches from each of 2
anonymized FOVs per split (4 train,
4 validation… See the full description on the dataset page: https://huggingface.co/datasets/u3011706/HEPROBench-CRC-CODEX-review-demo.wildbox-reviewgeodml-emnlp-2026-reviewer
GEODML — EMNLP 2026 reviewer pack
A condensed companion to the EMNLP 2026 submission "Causal
Analysis of LLM Search Rerankers via Double/Debiased Machine
Learning." Designed for fast verification, not full reproduction.
This pack (5.6 MB) — every paper claim as a CSV + every figure
as PDF/PNG + a one-shot verify.py claim-checker.
Full reproducibility dataset (1.8 GB) →
ValerianFourel/geodml-emnlp-2026
— raw LLM rerank outputs, page features, DFS confounders, and the
full… See the full description on the dataset page: https://huggingface.co/datasets/ValerianFourel/geodml-emnlp-2026-reviewer.b300-lora-review
