datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-reward-bench
AgentRewardBench
💾Code
📄Paper
🌐Website
🤗Dataset
💻Demo
🏆Leaderboard
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor
Loading dataset
You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.nrvbench-review
NR Video Editing Benchmark
This repository contains two non-rigid video editing benchmark subsets for evaluating instruction-driven video editing methods. Each row in metadata.csv corresponds to one editing instruction for a source video, with relative paths to the source video, extracted frames, binary masks, prompts, and evaluation questions.
The dataset card is written without author or institution identifiers so it can be used for anonymous review uploads. Before a non-anonymous… See the full description on the dataset page: https://huggingface.co/datasets/NRVBench/nrvbench-review.REPID
REPID: Rendering Evaluation of Photographic Image Dataset
REPID (officially introduced as the Rendering Evaluation of Photographic Image Dataset) is a large-scale benchmark designed for Image Rendering Quality Assessment (IRQA) in paper Beyond distortions: a benchmark for subjective evaluation of image rendering quality.
Unlike traditional Image Quality Assessment (IQA) which focuses on technical degradations like noise or blur, REPID aims to model subjective human aesthetic… See the full description on the dataset page: https://huggingface.co/datasets/vsevolodpl/REPID.hnm-fashion-recommendations-data
Dataset Rekomendasi Fashion H&M
Dataset ini berisi data transaksi, atribut pelanggan, dan metadata produk yang telah dianonimkan dari H&M Group. Kumpulan data komprehensif ini memungkinkan pemodelan perilaku pembelian pelanggan secara mendalam.
Wawasan yang dihasilkan dapat dimanfaatkan untuk berbagai tujuan bisnis yang strategis, mulai dari meningkatkan personalisasi pengalaman berbelanja, mengoptimalkan manajemen inventaris untuk efisiensi produksi, hingga mendukung inisiatif… See the full description on the dataset page: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data.exercise-dataset
Exercise Dataset — Free Tier (RepDB)
A free, ready-to-use fitness exercise dataset: 601 exercises, each
illustrated with flat-style 512×512 WebP images (a start/peak pose pair, or
a single main pose for static holds and stretches), with target muscles,
equipment, MET values, and full instructions in English, German, and
Spanish.
This public snapshot is the free tier of RepDB. Free for personal
and commercial use inside applications, with attribution.
Need exercise… See the full description on the dataset page: https://huggingface.co/datasets/RepDB/exercise-dataset.rocketleague-analysis
Rocket League Analysis
Local Rocket League replay analysis using Ballchasing API exports and plain DuckDB.
The report is meant to answer one practical question: what should I work on next from my saved replay sample?
Quick Start
uv sync --locked
UV_CACHE_DIR=/tmp/rocketleague-uv-cache \
uv run --locked pytest -v
uv run --locked python scripts/analyze_scenarios.py \
--replay-dir /path/to/Rocket\ League/TAGame/Demos \
--limit 10
Start with CONTRIBUTING.md… See the full description on the dataset page: https://huggingface.co/datasets/edmundmiller/rocketleague-analysis.cbi-archive-raw
Central Bank of Ireland Archive: original source files
6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and
archive gathered from the Central Bank of Ireland's public website, stored by
content hash so that a search result can be turned back into the document a
human would actually read.
This is the raw tier. If you want the text, you almost certainly want
aditya487/cbi-archive-corpus
instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.robocurate-synth100
synth100 — 100 generated clips for validating Pre-Contact Level Filtering
100 episodes drawn (seed 20260824) from the 952-episode multi-object generation set, packaged so
Stage-5 filtering can be run on them without re-deriving anything. Every input the filter needs
travels with the package, in the space it is consumed in.
Read section 1 before using this. The single most important fact about this data is not in the
file layout, and getting it wrong invalidates any score… See the full description on the dataset page: https://huggingface.co/datasets/glory-hyeok/robocurate-synth100.E2AM_ResNet50
E2AM Ablation Results: ResNet-50
Energy-aware training ablation study for ResNet-50 across three image-classification datasets: CIFAR-10, CIFAR-100, and Tiny-ImageNet.
Each dataset has 15 training variants (8 individual-method M0..M7, 7 cumulative ablation C0..C6) at 50 epochs, plus a 5-variant deployment pipeline (FP32 baseline, structured pruning, pruning+finetune, INT8 quantization, pruned+INT8).
Status: 45 completed variants, 0 partial.
Quick links… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/E2AM_ResNet50.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.retina-age-analysis
Retina Age Analysis Dataset
Dataset Description
This dataset contains 9,857 retinal fundus images from 5,393 patients for age prediction tasks.
Dataset Summary
Task: Age prediction from retinal fundus images
Images: 9,857 high-quality retinal images
Patients: 5,393 unique patients
Age Range: 5-97 years
Image Format: JPEG
Average Image Size: ~1 MB
Supported Tasks
Regression: Predict continuous age (5-97 years)
Classification: Predict age group (5… See the full description on the dataset page: https://huggingface.co/datasets/ramankamran/retina-age-analysis.IS110_repo
IS110 ecology — data and analysis catalogue
Full-length transposon systematics + rearrangement + movement analysis
of the IS110 family across 10 bacterial species. All catalogues are
built from a common set of ~92k IS110 element records with confirmed
empty-vs-filled boundaries (Cross_reference_IS pipeline).
Author : Kuang Hu (kh36969@berkeley.edu) — pc_rubinlab, UC Berkeley
Species covered : Escherichia coli, Klebsiella pneumoniae,
Enterobacter hormaechei, Salmonella enterica… See the full description on the dataset page: https://huggingface.co/datasets/hukuang/IS110_repo.LabUtopia-Dataset🧪 LabUtopia-Dataset: Scientific Laboratory 3D Asset Library (OpenUSD)
LabUtopia-Dataset is a large-scale 3D asset library designed for simulating scientific laboratory environments.
It provides realistic lab scenes, scientific instruments, and environmental props, all stored in OpenUSD (.usd / .usdz) format for high interoperability and composability.
🧩 File Format: OpenUSD
Each asset is stored as a .usd or .usdz file.
You can load them directly in:
NVIDIA Omniverse (Create, Isaac Sim)… See the full description on the dataset page: https://huggingface.co/datasets/Ruinwalker/LabUtopia-Dataset.HUI360
HUI360
HUI360: A 360° Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation (IEEE FG 2026)
Open-access skeleton annotations for HUI360, a large-scale 360° egocentric dataset for human-robot interaction anticipation in the wild. This repository provides the annotations as tabular CSV files (one row per detection), ready for training and evaluation with HUI360-Baselines.
Related resources
Resource
Link
Project… See the full description on the dataset page: https://huggingface.co/datasets/rlorlou/HUI360.agent-reward-bench
AgentRewardBench
💾Code
📄Paper
🌐Website
🤗Dataset
💻Demo
🏆Leaderboard
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor
Loading dataset
You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/yuanyyaa/agent-reward-bench.llava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.riftbound-cards
Riftbound TCG Card Database
Machine-readable snapshot of every card in
Riftbound: The League of Legends TCG, auto-scraped
from the official Card Gallery and errata pages.
Snapshot date: 2026-09-24
Cards: 1188
Source repo (scraper + pipeline): https://github.com/LouisCourrian/riftbound-cards
Every GitHub release publishes the same three formats as attached assets and
mirrors them here.
Files
File
What
cards.csv
Full corpus, one row per card. Array fields… See the full description on the dataset page: https://huggingface.co/datasets/Wysme/riftbound-cards.cafo-iowa
Dataset Card for cafo-iowa
Dataset revision in progress. We are currently revising portions of this dataset. Researchers interested in using this data are encouraged to contact the authors.
cafo-iowa is a dataset of Concentrated Animal Feeding Operations (CAFOs) in Iowa, compiled using NAIP satellite imagery, manual annotation, Iowa DNR permit records, and ReGrid parcel data. It provides facility-level estimates of animal populations derived from manually annotated barn… See the full description on the dataset page: https://huggingface.co/datasets/reglab/cafo-iowa.RealDoc-Bench-Layout
RealDocBench-Layout
A 1,500-page document-layout benchmark for evaluating layout-detection
models on real-world documents. COCO-style annotations across 9 block
classes.
Contents
images/ — 1,500 page images (PNG / JPG / occasional WebP-as-PNG; see Caveats).
annotations/<pageId>.json — per-page COCO files, each with a single
image record, an annotations list, a categories list, and a
page_info block.
manifest.csv — pageId → domain + source URLs. The canonical row… See the full description on the dataset page: https://huggingface.co/datasets/Extend-AI/RealDoc-Bench-Layout.mac-app-store-apps-metadata
Dataset Card for Macappstore Applications Metadata
📌 Dataset status: static snapshot (no scheduled updates). The data was collected from the public iTunes Search API between December 2023 and January 2024 and reflects the Mac App Store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule.
Mac App Store Applications Metadata sourced by the public API.
Curated by: MacPaw Way Ltd.
Language(s) (NLP): Mostly EN, DE… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-metadata.hubble-8b-unlearning-resultsrebus-dataset
|🔄 🚍| Re-Bus: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
Understanding Rebus Puzzles requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, and multi-step reasoning, making this a challenging task for current Vision-Language Models. In this paper, we present Re-Bus, a large and diverse benchmark of 1,333 English Rebus Puzzles containing different artistic… See the full description on the dataset page: https://huggingface.co/datasets/TrishanuDas/rebus-dataset.REC_SYSriemann-clock-spectra
Riemann Clock spectra
This dataset accompanies the exploratory spectroscopy study in
maris205/riemann_clock, with analysis
base commit 321ae1cd48d69fe4d1091326167a8aaa65f2e597.
It mirrors public, previously reduced, extracted, or coadded quasar spectra and
the study's processed arrays. It contains no newly acquired observations and no
raw CCD frames. The project directory name data/raw/ means downloaded inputs.
This exploratory study reports no detection of Riemann-clock physics… See the full description on the dataset page: https://huggingface.co/datasets/dnagpt/riemann-clock-spectra.eligible-scroll-atlas-renders
Get one mesh in about twenty seconds
curl -sO https://raw.githubusercontent.com/rodriguescarson/eligible-scroll-atlas/main/scripts/atlas.py
python atlas.py list --ink-pass # the 5 meshes that pass the pre-registered screen
python atlas.py ink PHerc0125 z10544_w020 --preview # a downsampled ink map, about 12 KB
python atlas.py get PHerc0125 z10544_w020 # the surface volume, 31 planes, plane 15 is the surface
from atlas import meshes… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/eligible-scroll-atlas-renders.ridgelora-cross-sensor-sd302d-f-to-m-20260825
RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment
This public archive contains the leakage-controlled direct cross-sensor
experiment used to evaluate whether Stage-2 synthetic target-sensor images
help recognition on a physically different real sensor.
Locked protocol
Source/condition sensor: NIST SD302A device F.
Target sensor: NIST SD302D device M.
Identity: subject:finger-position; the same fingers exist across both
collections.
Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.BioVITA
Citation
@inproceedings{shinoda2026biovita,
title = {BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment},
author = {Risa Shinoda and Kaede Shiohara and Nakamasa Inoue and Kuniaki Saito and Hiroaki Santo and Fumio Okura},
booktitle = {CVPR},
year = {2026},
}
cdl-devai-results
CDL DevAI results — brain × interpretability × localisation, per model per checkpoint
Developmental analysis of 10 language-model families against the ds003604 auditory
language fMRI dataset. For every training checkpoint of every model we measured three
things and here report them side by side:
axis
what it asks
source tables
brain
does the model's representational geometry match the brain's?
brain_alignment
interp
how is the representation organised internally?… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results.rubber-tree-leaf-disease-ph-segmented
