datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imgBoxFusiongdpval-gpt5
GDPval with GPT-5 Execution Results
This dataset contains the OpenAI GDPval benchmark with comprehensive execution results from GPT-5, demonstrating AI capabilities across real-world professional tasks.
🎯 Dataset Overview
This is an enhanced version of the original OpenAI GDPval dataset with actual AI model execution results and professional deliverables.
📊 Key Statistics
Total tasks: 220
Tasks with AI deliverables: 87 (39.5%)
Professional files generated:… See the full description on the dataset page: https://huggingface.co/datasets/kevindenight/gdpval-gpt5.gdpval-gpt5-fork
GDPval Fork Dataset with GPT-5 Results
🏆 A comprehensive evaluation dataset featuring GPT-5 execution results on real-world professional tasks
This is an enhanced fork of the original OpenAI GDPval dataset with complete GPT-5 execution results, including actual deliverable files created by the AI model.
📊 Dataset Overview
Metric
Value
Total Tasks
220
AI-Completed Tasks
87 (39.5%)
Deliverable Files
492+ professional documents
Occupations
44
Industry… See the full description on the dataset page: https://huggingface.co/datasets/kevindenight/gdpval-gpt5-fork.ScreenSpotReBOTabularMath
📊 TabularMath
TabularMath is a tabular mathematical reasoning benchmark introduced in TabularMath: Understanding Math Reasoning over Tables with Large Language Models. It is built via AUTOT2T, a neuro-symbolic pipeline that automatically transforms math word problems into verified tabular reasoning tasks, enabling scalable evaluation without manual table annotation.
TabularMath jointly assesses reasoning accuracy, information retrieval over complex table structures, and… See the full description on the dataset page: https://huggingface.co/datasets/kevin715/TabularMath.frame_envThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 300,
"total_frames": 57506,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60.0,
"splits": {
"train": "0:300"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kevin-ys-zhang/frame_env.PennEngineeringPhotos
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/kevkevhan/PennEngineeringPhotos.structure_wildfire_damage_classification
Dataset Card for Structures Damaged by Wildfire
Homepage: Image Dataset of Structures Damaged by Wildfire in California 2020-2022
Dataset Summary
The dataset contains over 18,000 images of homes damaged by wildfire between 2020 and 2022 in California, USA, captured by the California Department of Forestry and Fire Protection (Cal Fire) during the damage assessment process. The dataset spans across more than 18 wildfire events, including the 2020 August Complex Fire, the… See the full description on the dataset page: https://huggingface.co/datasets/kevincluo/structure_wildfire_damage_classification.test_obsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 200,
"total_frames": 62462,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60.0,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kevin-ys-zhang/test_obs.block_envThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 300,
"total_frames": 61779,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60.0,
"splits": {
"train": "0:300"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kevin-ys-zhang/block_env.egocentric-100k-sample-framesProteinFoldingfasth3-pr-assetsMUCARrlbench-lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "rlbench_panda",
"total_episodes": 1200,
"total_frames": 155851,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1000",
"test": "1000:1200"
},
"data_path":… See the full description on the dataset page: https://huggingface.co/datasets/KEVIN04087/rlbench-lerobot.rlbench-lerobot-train-close_boxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "rlbench_panda",
"total_episodes": 100,
"total_frames": 19044,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/KEVIN04087/rlbench-lerobot-train-close_box.Vincent-van-Goghwill_datasetSintelreplicatoxicity-dataset-reasoningEnhancing_Intent_Understanding
Image Prompt Dataset
This repository contains image datasets used in the paper "Enhancing Intent Understanding for Ambiguous Prompts: A Human-Machine Co-Adaption Strategy" by Yangfan He, Jianhui Wang, et al.
Overview
This dataset was created to support research on human-machine co-adaptation in text-to-image generation systems. It contains various categories of images that can be used for training and evaluating models that aim to better understand user intent in… See the full description on the dataset page: https://huggingface.co/datasets/Kevin3777/Enhancing_Intent_Understanding.IDRnD-Replayoriginal dataset source: https://www.kaggle.com/datasets/nurmukhammed7/idrd-train-set
MirrorSentinel-Elevator
MirrorSentinel Reflective Elevator Dataset
MirrorSentinel-Elevator is a synchronized RGB-LiDAR-IMU ROS 2 dataset for
studying severe mirror and glass interference in compact elevator
environments. It contains eight short traversals, E01-E08, collected across
seven measured physical elevator footprints. E02 and E03 repeat the same
physical cabin and are grouped as one footprint for inferential statistics.
The release contains the exact raw bags used by the accompanying paper. It… See the full description on the dataset page: https://huggingface.co/datasets/KevinWong216/MirrorSentinel-Elevator.LeafScan-CornDefoliation2025
# LeafScan-CornDefoliation2025-V1.0 Dataset
A multi-level dataset for corn leaf defoliation assessments research. Provides corn leaves processed in both video and image form. Paired to the LeafScan research project at https://github.com/KevynAngueira/LeafScan.
Fields: 7 sampled sites across 3 states (IA, IN, OH)
Plants: 18 total plants
Leaves: 149 leaves total, with leaf number 7-21
Media: 1000+ videos & images across healthy, defoliated, and simulated conditions
Zenodo:… See the full description on the dataset page: https://huggingface.co/datasets/KevynAngueira/LeafScan-CornDefoliation2025.PokeFA-pokemon-fanart-captioned
PokeFA — Pokémon fan-art metadata with relevance/aesthetic scores and hybrid captions
PokeFA is a large-scale Pokémon fan-art dataset released as metadata + URLs only (no image bytes).~30,000 candidate images are collected across 1,025 Pokémon using a popularity-banded budget with following curation pipeline:
NSFW filtering → OCR localization & inpainting → resizing → relevance & aesthetic scoring (GPT-5-mini vision) → near-duplicate removal → quality filtering to the top ~16… See the full description on the dataset page: https://huggingface.co/datasets/Kev0208/PokeFA-pokemon-fanart-captioned.flint_images_600_600
Dataset Card for "flint_images_600_600"
More Information needed
libero100_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 5000,
"total_frames": 807133,
"total_tasks": 100,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/kevin-ys-zhang/libero100_lerobot.
