datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eden_objaverseGenECGGenECG is an image-based ECG dataset which has been created from the PTB-XL dataset (https://physionet.org/content/ptb-xl/1.0.3/).
The PTB-XL dataset is a signal-based ECG dataset comprising 21799 unique ECGs.
GenECG is divided into the following subsets:
-Dataset A: ECGs without imperfections (Dataset_A_ECGs_without_imperfections) - This subset includes 21799 ECG images that have been generated directly from the original PTB-XL recordings, free from any visual imperfections.
-Dataset B: ECGs… See the full description on the dataset page: https://huggingface.co/datasets/edcci/GenECG.gpt-edit-simplerGPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.dartlab-mediaEditReward-CompassHQ-Edit
Dataset Card for HQ-EDIT
HQ-Edit, a high-quality instruction-based image editing dataset with total 197,350 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3.
HQ-Edit’s high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/HQ-Edit.seggpt-example-dataeddmpython-mediaedenszero
Bangumi Image Base of Edens Zero
This is the image base of bangumi Edens Zero, we detected 201 characters, 16692 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/edenszero.NHR-Edit
NoHumanRequired (NHR) Dataset for image editing
🌐 NHR Website |
📜 NHR Paper on arXiv |
💻 GitHub Repository |
🤗 NHR-Edit Dataset (part2) |
🤗 BAGEL-NHR-Edit |
❗️ Important: This is Part 1 of the Dataset ❗️
Please be aware that this repository contains the first part of the full NHR-Edit dataset. To have the complete training data, you must also download the second part.
➡️ Click Here to Access Part 2 on… See the full description on the dataset page: https://huggingface.co/datasets/iitolstykh/NHR-Edit.godot_rl_FlyByA RL environment called FlyBy for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_FlyBy
Wake-Vision
Dataset Card for Wake Vision
Dataset Description
"Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of
current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally,
it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age,
subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.Complex-Edit
Complex-Edit: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
📃Arxiv | 🌐Project Page | 💻Github | 📚Dataset | 📄HF Paper
We introduce Complex-Edit, a comprehensive benchmark designed to systematically evaluate instruction-based image editing models across instructions of varying complexity. To develop this benchmark, we harness GPT-4o to automatically collect a diverse set of editing instructions at scale.
Our approach follows a well-structured… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Complex-Edit.godot_rl_BallChaseA RL environment called BallChase for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_BallChase
edit3d-bench
Dataset Card for edit3d-bench
This is a FiftyOne dataset with 300 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from huggingface_hub import snapshot_download
# Download the dataset snapshot to the current working directory
snapshot_download(
repo_id="Voxel51/edit3d-bench",
local_dir=".",
repo_type="dataset"
)
# Load dataset from current directory using FiftyOne's… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/edit3d-bench.multi_reference_image_editing
Multi-Reference Instruction-Based Image Editing Dataset
Overview
This dataset contains 20,000 high-resolution image pairs and multi-modal instructions designed for training advanced image-to-image editing models. It combines two complementary example types: 10,000 reference-grounded edits, where structural or stylistic changes are driven by up to three provided visual reference images, and 10,000 occlusion-based inpainting/outpainting edits, where the model must… See the full description on the dataset page: https://huggingface.co/datasets/molbal/multi_reference_image_editing.Edit3D-Bench
Edit3D-Bench
Paper | Project Page | Code
Edit3D-Bench is a benchmark for 3D editing evaluation, introduced in the paper VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space.
This dataset comprises 100 high-quality 3D models, with 50 selected from Google Scanned Objects (GSO) and 50 from PartObjaverse-Tiny.
For each model, we provide 3 distinct editing prompts. Each prompt is accompanied by a complete set of annotated 3D assets, including
original 3D asset… See the full description on the dataset page: https://huggingface.co/datasets/huanngzh/Edit3D-Bench.rocketleague-analysis
Rocket League Analysis
Local Rocket League replay analysis using Ballchasing API exports and plain DuckDB.
The report is meant to answer one practical question: what should I work on next from my saved replay sample?
Quick Start
mise install
mise run setup
mise run test
mise exec -- python scripts/analyze_scenarios.py \
--replay-dir /path/to/Rocket\ League/TAGame/Demos \
--limit 10
Start with CONTRIBUTING.md before changing the pipeline.
Replay files and… See the full description on the dataset page: https://huggingface.co/datasets/edmundmiller/rocketleague-analysis.UIEBEpiBench-NeurIPS2026
EpiBench
Anonymous release for NeurIPS 2026 Evaluations & Datasets Track review (paper ID 1899). All methodology, ablations, and analyses are in the companion paper; this card lists only what reviewers and downstream users need to load the data.
A 25,737-patient ILAE-aligned multimodal epilepsy benchmark derived from PubMed Central case reports + 192 EpiRAG textbook vignettes.
6 tasks: epilepsy_type, seizure_type, ez_localization, aed_response, surgery_outcome, status_epilepticus… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-1899-ED-2026/EpiBench-NeurIPS2026.mmconflict-editable-values-1k
MMConflict Editable Values 2K
This dataset contains 2,000 source images with visible atomic values for
multimodal conflict research. It has 100 images in each of 20 categories. Every
image comes from a photograph, scan, captured website, software screenshot, or
page of a source document. The dataset does not contain generated images or
project-rendered examples.
Each row records the source, source URL, license, attribution, visible value,
question, and a candidate box around the… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/mmconflict-editable-values-1k.Omni-Edu
Omni-Edu — Core V6 SFT mixture
69,999 supervised instruction examples (~158M characters) covering K-12 subject
competence, curriculum grounding, diagnostic reasoning, pedagogical action and
general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image
referenced by the JSONL ships in this repository under images/.
This is the system-prompted assembly of the v6 core mixture: every row carries
an explicit system message, and the non-system turns are byte-identical… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Omni-Edu.EditReward-Bench
News |
Quick Start |
Benchmark Usage |
Citation
EditScore is a series of state-of-the-art open-source reward models (7B–72B) designed to evaluate and enhance instruction-guided image editing.
✨ Highlights
State-of-the-Art Performance: Effectively matches the performance of leading proprietary VLMs. With a self-ensembling strategy, our largest model surpasses even GPT-5 on our comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/EditScore/EditReward-Bench.godot_rl_JumperHardA RL environment called JumperHard for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_JumperHard
v1godot_rl_3DCarParkingA RL environment called 3DCarParking for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_3DCarParking
nornikel-metallurgy-rag-index
Nornikel Metallurgy RAG Knowledge Base (для векторного индекса)
Дедуплицированный корпус текстовых фрагментов (и связанных изображений) из
технической базы знаний по металлургии/горному делу/обогащению — источник
для построения векторного индекса (Annoy) в пайплайне RAG.
Индекс НЕ включён в этот репозиторий — эмбеддинги и Annoy-индекс
строятся во время выполнения ноутбука Google Colab (на GPU, это быстрее,
чем на CPU), используя corpus.jsonl как исходные данные. Модель… See the full description on the dataset page: https://huggingface.co/datasets/brics-edtech/nornikel-metallurgy-rag-index.Prostate-Anatomical-Edge-Cases
Prostate-Anatomical-Edge-Cases
Stress-Testing Pelvic Autosegmentation Algorithms Using Anatomical Edge Cases —
a TCIA collection of pelvic radiotherapy planning CT with manually contoured
organs at risk, curated so that most cases contain anatomy known to break
autosegmentation algorithms (Kanwar et al., Phys Imaging Radiat Oncol 2023).
Read before using — the name is misleading in two ways:
This is CT, not MRI. Despite "Prostate" in the name it is not a prostate
mpMRI/zonal… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Prostate-Anatomical-Edge-Cases.
