datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
document-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.imagenet_hard_review_data_r2drug-reviewsencoded_drug_reviewsS-EMBER
S-EMBER
S-EMBER is a benchmark for streaming episodic memory over long egocentric
(wearable-camera) video. Given a video and a natural-language question about
something that happened earlier in the recording, a model must recall the
relevant moment and answer.
This repository is an anonymized mirror provided for peer review. It
contains the complete benchmark. Author, institution, and provenance
information has been intentionally omitted for double-blind review.… See the full description on the dataset page: https://huggingface.co/datasets/paper-review-only/S-EMBER.datause-displacement-reviewed
datause-displacement-reviewed
The Luna-reviewed subset of
rafmacalaba/datause-displacement:
only spans that received a v2.3 Luna verdict (band review + drop-side rescue,
source == luna_review). Every span carries the binary label plus
usage_type / drop_reason / specificity, and is traceable via key
(split:row:start:end) to the verdict records in
extraction_analysis/band_review/.
Configs
config
fields
gliner_reviewed
tokenized_text, corpus, origin… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-displacement-reviewed.google-play-reviewMLVU-Attention-Review
MLVU Attention Review 下载与解压说明
本仓库提供一个未压缩 TAR,里面是完整静态 HTML 展示、所有页面所需图片、QA/GT、模型原始回答、attention 元数据与未归一化 grids.npz,并附离线查看和逐文件 SHA256 校验脚本。归档不含模型权重、完整源视频或推理环境。
容量要求
下载和解压期间需同时存放 TAR 与解压内容,请预留至少归档大小约 2.2 倍的可用空间。
文件系统必须支持大于 4 GB 的单文件(NTFS、exFAT、ext4、APFS 等;FAT32 不可用)。
Windows 建议在较短路径中操作,例如 D:\reviews\MLVU,避免长路径限制。
下载
安装最新版 Hugging Face CLI:
pip install -U huggingface_hub
hf download GraciaChen/MLVU-Attention-Review MLVU-Attention-Review.tar SHA256SUMS… See the full description on the dataset page: https://huggingface.co/datasets/GraciaChen/MLVU-Attention-Review.french_book_reviews
Dataset Card for French book reviews
I-Dataset Summary
The majority of review datasets are in English. There are datasets in other languages, but not many. Through this work, I would like to enrich the datasets in the French language(my mother tongue with Arabic).The data was retrieved from two French websites: Babelio and Critiques LibresLike Wikipedia, these two French sites are made possible by the contributions of volunteers who use the Internet to share their… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/french_book_reviews.SciCode-Runnable-Benchmark-Reviewediclr2026_real_reviewspi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-diff-review.review
🛡️ LabShield: A Multimodal Benchmark for Laboratory Safety
Official dataset for the paper: "LabShield: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories".
[Project Website (Coming Soon)] | [Paper (NeurIPS 2026 Submission)] | [Code (Coming Soon)]
📌 Introduction
LabShield is a rigorous, multi-view benchmark designed to assess the safety awareness and decision-making reliability of Multimodal Large Language Models (MLLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/LabShield-Review/review.pad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.review
🛡️ LabShield: A Multimodal Benchmark for Laboratory Safety
Official dataset for the paper: "LabShield: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories".
[Project Website (Coming Soon)] | [Paper (NeurIPS 2026 Submission)] | [Code (Coming Soon)]
📌 Introduction
LabShield is a rigorous, multi-view benchmark designed to assess the safety awareness and decision-making reliability of Multimodal Large Language Models (MLLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/LabShield/review.gaming-input-review-100-media-20260910100 Gaming input-only preview clips, each 15 seconds. No visual filtering or semantic action labeling was performed for this sample. The sample is stratified by game and includes 30 GTA V clips. Use samples.jsonl for source intervals, checksums and input timing. Video paths are relative to this repo.
Companion review: https://huggingface.co/spaces/mikusama99/sunain-gaming-review-200-20260910
pi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each sessions/*.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-diff-review.VideoMMEv2-Attention-Review
VideoMMEv2 Attention Review 下载与解压说明
本仓库提供一个未压缩 TAR,里面是完整静态 HTML 展示、所有页面所需图片、QA/GT、模型原始回答、attention 元数据与未归一化 grids.npz,并附离线查看和逐文件 SHA256 校验脚本。归档不含模型权重、完整源视频或推理环境。
容量要求
下载和解压期间需同时存放 TAR 与解压内容,请预留至少归档大小约 2.2 倍的可用空间。
文件系统必须支持大于 4 GB 的单文件(NTFS、exFAT、ext4、APFS 等;FAT32 不可用)。
Windows 建议在较短路径中操作,例如 D:\reviews\VideoMMEv2,避免长路径限制。
下载
安装最新版 Hugging Face CLI:
pip install -U huggingface_hub
hf download GraciaChen/VideoMMEv2-Attention-Review… See the full description on the dataset page: https://huggingface.co/datasets/GraciaChen/VideoMMEv2-Attention-Review.PR_review_deepseek
Pull request review task
Given the original code snippet (may be truncated) and a pull request (in diff format), the model reviews the PR and decides whether it should be merged.
The answers are generated by Deepseek-V2
perekrestok-reviews
Dataset
Dataset of user reviews from "Перекрёсток/Perekrestok" shop.
Dataset Format
Dataset is in JSONLines format. Trivia:
product_id - Product internal ID (https://www.perekrestok.ru/cat/1/p/ID)
product_name - Product name
product_category - Category of product
product_price - Product price in RUB (decimal)
review_id - Review internal ID
review_author - Author of review
review_text - Text of review
rating - Review rating (decimal, from 0.0 to 5.0)
review_helpfulness_prediction
Dataset Card for Review Helpfulness Prediction (RHP) Dataset
Dataset Summary
The success of e-commerce services is largely dependent on helpful reviews that aid customers in making informed purchasing decisions. However, some reviews may be spammy or biased, making it challenging to identify which ones are helpful. Current methods for identifying helpful reviews only focus on the review text, ignoring the importance of who posted the review and when it was posted.… See the full description on the dataset page: https://huggingface.co/datasets/tafseer-nayeem/review_helpfulness_prediction.yelp_academic_dataset_reviewvocalcoachbench-review
VocalCoachBench
VocalCoachBench is a singing-audio benchmark for evaluating vocal coaching
judgments. This release contains expert annotations for 515 singing recordings:
free-form coaching feedback, atomic diagnosis/correction claims, Top-3 issue
labels, same-song triplet rankings, and segment-conditioned issue labels.
Subsets:
same_song / Dataset A: 207 Amazing Grace performances from DAMP-S-AG.
Audio is not redistributed; use audio_filename to match the official release.… See the full description on the dataset page: https://huggingface.co/datasets/vocalcoachbench/vocalcoachbench-review.amazon-review
Amazon Review Dataset
This dataset contains Amazon reviews from January 1, 2018, to June 30, 2018. It includes 2,245 sequences with 127,054 events across 18 category types. The original data is available at Amazon Review Data with citation information provided on the page. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
Update (2025-10-28): Added three timestamp fields (timestamp_event… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/amazon-review.MedHorizon
MedHorizon
MedHorizon is a long-context medical video benchmark for evaluating multimodal models on full-procedure clinical videos. The benchmark emphasizes two properties that are not captured by short-clip medical video datasets: extremely sparse evidence retrieval and multi-hop reasoning over observations distributed across a full procedure.
Dataset Contents
Videos: 340 full-procedure videos.
Questions: 1,253 multiple-choice QA pairs.
Evaluation split: test.
Video… See the full description on the dataset page: https://huggingface.co/datasets/mlvbench-review/MedHorizon.review_arcade
ACL ARR Reviews - Review Arcade Project
This dataset contains paper reviews from the ACL ARR (Association for Computational Linguistics - Annual Review of Research) program. The reviews are organized into splits corresponding to papers that were accepted or rejected for publication, as well as specific subsets used for the research analysis.
This dataset was introduced in the paper Review Arcade: On the Human Alignment and Gameability of LLM Reviews.
Code:… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/review_arcade.mobile-apps-user-sentiment-reviews
Top Mobile Apps User Sentiment & Review Corpus (Google Play)
Overview
This dataset contains clean, structured public data exported directly from production runs of Apify actors.
It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines.
Source Actor: captainhandsome/google-play-reviews-scraper
Dataset Page: Public sample and schema
Preconfigured Run Task: captainhandsome/instagram-1star-reviews… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/mobile-apps-user-sentiment-reviews.app-store-reviews-scraper
App Store Reviews Scraper
Scrape Apple App Store reviews, star ratings and app version history for any iOS app in any country storefront. No login, no API key.
Rows in this dataset
23,048
Fields
43
Collector runs behind it
88
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/app-store-reviews-scraper/ — 176 entity pages
Run the collector yourself
https://apify.com/reapx/app-store-reviews-scraper
What this is… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/app-store-reviews-scraper.hft12735-review-jh6gwynw
Model Candidate Review Report
Review Date: 2026-08-23
Total Candidates Reviewed: 12
Approved Count: 5
Rejected Count: 7
Top Approved Model: mod-010
Approved Model IDs: mod-001,mod-002,mod-003,mod-008,mod-010
Edit-Review
