datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RekaDaily-10k-raw
RekaDaily-10k (raw)
Raw, unscripted, first-person daily-life video, collected through
Claru, Reka's data collection marketplace — recorded by
paid collectors in their own homes and workplaces on head-mounted and handheld
phones, across multiple regions.
Videos are delivered as recorded — no cuts, no trimming, no editing, no
filtering beyond basic integrity checks. A processed tier (short clips with
machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.VDR_MEGA_MultiDomain_DocRetrieval
Visual Document Retrieval Dataset
Overview
This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks.
Dataset Structure
The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.RobotDesign1M
RobotDesign1M: A Large-scale Dataset for Robot Design Understanding
RobotDesign1M is a large-scale, multimodal dataset for robot design understanding, built from image–text data curated from scientific literature across a wide range of robotics domains. It is designed to support research on design-aware foundation models, including design image generation, visual question answering about designs, and design image retrieval.
📄 Paper: RobotDesign1M: A Large-scale Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Fsoft-AIC/RobotDesign1M.DSERT-RoLL
DSERT-RoLL Dataset
DSERT-RoLL is a road-scene dataset collected under diverse weather and lighting conditions (e.g., Clear, Fog, Rain, Snow) for robust perception research.
Project Page: https://jeongyh98.github.io/dsert-roll/
GitHub: https://github.com/jeongyh98/DSERT-RoLL-Dataset
Changelog
2026-05-28
Merged the previous Normal and Night categories into a single Clear category to align with the paper's weather condition taxonomy.
The dataset now… See the full description on the dataset page: https://huggingface.co/datasets/jeongyh98/DSERT-RoLL.RenderedTextThis dataset has been created by Stability AI and LAION.
This dataset contains 12 million 1024x1024 images of handwritten text written on a digital 3D sheet of paper generated using Blender geometry nodes and rendered using Blender Cycles. The text has varying font size, color, and rotation, and the paper was rendered under random lighting conditions.
Note that, the first 10 million examples are in the root folder of this dataset repository and the remaining 2 million are in ./remaining (due… See the full description on the dataset page: https://huggingface.co/datasets/wendlerc/RenderedText.catholic-resources
Vietnamese Catholic resources by v-bible
Data Structure
calendar: Generated Liturgical calendars using
v-bible/js-sdk.
misc/proper-names.json: Name translation from
ktcgkpv.org, generated by
v-bible/bible-scraper.
liturgical: Liturgical data from
The Lectionary for Mass (1998/2002 USA Edition),
compiled by Felix Just, S.J., Ph.D., and generated by
v-bible/bible-scraper.
books/bible: Generated Bible markdown data.
books/catechism-books: Official catechism… See the full description on the dataset page: https://huggingface.co/datasets/v-bible/catholic-resources.PhysicalAI-Robotics-Locomanipulation-GRAIL
📢 News
[2026-07-15] Released task-general tracking policy checkpoints trained on the released data. Follow the tracking doc to use them to track our released motion data.
[2026-07-14] Updated data/pickup_table and data/pickup_ground. If you downloaded them before this date, please re-download.
Dataset Overview
Tabletop Pickup
Ground Pickup
Tabletop Manipulation
Ground Manipulation
Sitting
Curb
Slope… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Locomanipulation-GRAIL.RGB-Event-ISP-Dataset3dfront_render_viewsolmocr-pre-rendered
olmOCR-bench Pre-Rendered
Pre-rendered PNG images of the olmOCR-bench benchmark dataset, ready for zero-setup evaluation of any OCR / vision model.
What This Is
The official olmOCR benchmark requires downloading 1,403 PDFs locally and rendering each page to a PNG image before sending it to a model. Every benchmark runner in the official repo does this same rendering step internally — see olmocr/data/renderpdf.py::render_pdf_to_base64png().
This dataset eliminates that… See the full description on the dataset page: https://huggingface.co/datasets/shhdwi/olmocr-pre-rendered.nerf-gs-datasetsI keep a collection compiled of existing datasets from various sources for training NeRFs or Splats. This dataset is most of that collection. All of the individual scenes also have a trained Gaussian Splat.
https://rishit-dagli.github.io/2025/03/28/nerf-gs-datasets.html
3d-front-rgbRDD2022
RDD2022: Multi-National Road Damage Detection Dataset (4-Class YOLO Export)
Unofficial redistribution of the RDD2022 multi-national road-damage dataset, reduced to the 4-class CRDDC2022 taxonomy and reformatted into a standardized YOLO-compatible directory layout, under the original CC BY-SA 4.0 license.
Disclaimer
This repository is not an official release of the RDD2022 dataset.
RDD2022 was created by Deeksha Arya, Hiroya Maeda, Sanjay Kumar Ghosh… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/RDD2022.random-imagesOmniRooms
UniSHARP:
Universal Sharp Monocular View Synthesis
Meixi Song1 ·
Dizhe Zhang1,* ·
Hao Ren1 ·
Ruiyang Zhang1 ·
Bo Du2 ·
Ming-Hsuan Yang3 ·
Lu Qi1,2,*
1Insta360 Research · 2Wuhan University · 3University of California, Merced
UniSHARP extends SHARP-style photorealistic monocular view synthesis to universal camera systems. Given a single image from a perspective, wide-FoV, fisheye, or panoramic camera, UniSHARP predicts a 3D Gaussian representation and… See the full description on the dataset page: https://huggingface.co/datasets/Insta360-Research/OmniRooms.gaia2_filesystem
GAIA2 Filesystem
This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset.
Dataset Link
https://huggingface.co/datasets/meta-agents-research-environments/gaia2
Contact Details
Publishing POC: Meta AI Research Team
Affiliation: Meta Platforms, Inc.
Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.RekaDaily-10k-processed
RekaDaily-10k (processed)
Short first-person clips cut from the RekaDaily-10k
recordings —
unscripted daily-life video collected through Claru, Reka's
data collection marketplace, recorded by paid collectors in their own homes and
workplaces on head-mounted and handheld phones, across multiple regions.
Every clip carries one dense caption and a multi-question Q&A exchange
written in the second person ("What am I doing in this video?"), so the corpus
drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.agent-reward-bench
AgentRewardBench
💾Code
📄Paper
🌐Website
🤗Dataset
💻Demo
🏆Leaderboard
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor
Loading dataset
You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.RealCQA
RealCQA: Real-World Complex Question Answering Dataset
This repository contains the dataset used in the paper "RealCQA: Scientific Chart Question Answering as a Test-Bed for First-Order Logic" (ICDAR 2023). The dataset is designed to facilitate research in complex question answering, involving a diverse set of real-world images and associated textual question-answer pairs.
Dataset Overview
The RealCQA dataset consists of 28,266 images, and corresponding 2 million… See the full description on the dataset page: https://huggingface.co/datasets/sal4ahm/RealCQA.RAG_EvalLLaVA-ReCap-CC12Mimagenet-r
ImageNet-R
This repo is made to facilitate the evaluation of various pretraining models. It's constructed from the source file provided by official implementation.
Usage
from datasets import load_dataset
dataset = load_dataset('axiong/imagenet-r')
Dataset Summary
ImageNet-R(endition) contains art, cartoons, deviantart, graffiti, embroidery, graphics, origami, paintings, patterns, plastic objects, plush objects, sculptures, sketches, tattoos, toys, and video… See the full description on the dataset page: https://huggingface.co/datasets/axiong/imagenet-r.imagenet_1k_resized_256
Dataset Card for "imagenet_1k_resized_256"
Dataset summary
The same ImageNet dataset but all the smaller side resized to 256.
A lot of pretraining workflows contain resizing images to 256 and random cropping to 224x224, this is why 256 is chosen.
The resized dataset can also be downloaded much faster and consume less space than the original one.
See here for detailed readme.
Dataset Structure
Below is the example of one row of data. Note that the labels in… See the full description on the dataset page: https://huggingface.co/datasets/evanarlian/imagenet_1k_resized_256.ViFailback-Dataset
ViFailback Dataset: Real-World Robotic Manipulation Failure Dataset with Visual Symbol Guidance
A real-world dataset for diagnosing, correcting, and learning from robotic manipulation failures via visual symbols.
ViFailback is a large-scale, real-world robotic manipulation failure dataset introduced in the CVPR 2026 paper "Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols". It introduces visual… See the full description on the dataset page: https://huggingface.co/datasets/sii-rhos-ai/ViFailback-Dataset.KAIST-Multispectral-Pedestrian-Detection-Datasetkaz-vision-50kRoboBench
RoboBench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
📋 Overview
RoboBench is a comprehensive evaluation benchmark designed to assess the capabilities of Multimodal Large Language Models (MLLMs) in embodied intelligence tasks. This benchmark provides a systematic framework for evaluating how well these models can understand and reason about robotic scenarios.
This repository contains the released RoboBench… See the full description on the dataset page: https://huggingface.co/datasets/LeoFan01/RoboBench.fractal_rawD3HRgw-tti-playground-queue
