datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpatialEdit-500K
SpatialEdit-500K
SpatialEdit-500K is a synthetic training dataset for fine-grained image spatial editing. It is built for learning geometry-aware edits such as object moving, object rotation, and camera viewpoint change.
The dataset was introduced in the paper SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing. It is generated with a controllable rendering pipeline to provide structured spatial transformations at scale.
Project Resources
GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/EasonXiao-888/SpatialEdit-500K.GPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.FootballMolmo2-ER-RoboPoint
Molmo2-ER · wentao-yuan/robopoint-data
1.43M robotics affordance instruction-tuning examples (pointing + detection + VQA).
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
Upstream source
Original dataset: wentao-yuan/robopoint-data
Paper: RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics (arXiv:2406.10721)
License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-RoboPoint.wds_vtab-eurosat3D_Visual_Illusion_Depth_Estimation
3D Visual Illusion Depth Estimation Dataset
Dataset Summary
The 3D Visual Illusion Depth Estimation Dataset is designed for research on stereo and monocular depth estimation in 3D visual illusion scenes.It contains left and right stereo images, depth maps estimated from DepthAnything V2, and illusion-region masks.
Dataset Structure
Each sample in the dataset includes:
left: Left-view RGB image
right: Right-view RGB image
depth: Monocularly estimated depth… See the full description on the dataset page: https://huggingface.co/datasets/AdamYao/3D_Visual_Illusion_Depth_Estimation.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Trainset.Uni-Edit-Train-Data
Uni-Edit Training Data: Uni-Edit-148k
Project Page | GitHub Repository | Paper
👀 Intro
We introduce Uni-Edit, an intelligent image editing task that serves as the first general task for Unified Multimodal Model (UMM) tuning. Unlike conventional mixed multi-task training that suffers from inherent task conflicts and requires complex multi-stage pipelines, Uni-Edit breaks this paradigm. It achieves true mutual reinforcement by improving image… See the full description on the dataset page: https://huggingface.co/datasets/Uni-Edit/Uni-Edit-Train-Data.EgoVid_framesNeMo
NVIDIA NeMo Speech
Checkout our HuggingFace🤗 collection for the latest open
weight checkpoints and demos!
Updates
2026-03: Nemotron 3 VoiceChatis now released in Early Access. Built on the Nemotron Nano v2 LLM backbone with Nemotron speech and TTS decoder, VoiceChat delivers full-duplex, natural, interruptible conversations with low latency. Try out the demo and apply for early access.
2026-03: Nemotron-Speech-Streaming v2603 has been
updated. It has been… See the full description on the dataset page: https://huggingface.co/datasets/echodict/NeMo.sdxl_images_easy_prompts-artists-seed1LOVE-Agibot-BetaDL3DV-Evaluation
DL3DV Testing Split Download Instructions
This repo contains all 55 scenes for evaluation. Note: it is an independent dataset, and none of its scenes overlap with those in DL3DV-10K. Have a galance on the preview page: https://dl3dv-10k.github.io/DL3DV-Testing-Split-Preview/.
Download
As the whole benchmark dataset is ~500G, a python script to download and untar files.
Environment Setup
The download script relies on huggingface hub, tqdm. You can download by… See the full description on the dataset page: https://huggingface.co/datasets/DL3DV/DL3DV-Evaluation.egoxtreme
EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
📖 Dataset Information
EgoXtreme is a novel large-scale dataset designed for robust egocentric 6D object pose estimation under extreme environmental conditions. The dataset comprises approximately 1.3 million frames with a total duration of 775.5 minutes (~12.9 hours). It was captured at 30 fps using Aria glasses, providing high-resolution 1408 x 1408 raw fisheye RGB… See the full description on the dataset page: https://huggingface.co/datasets/taegyoun88/egoxtreme.dynamic_earthnetDynamic EarthNet dataset redistributed from https://mediatum.ub.tum.de/1650201 and https://cvg.cit.tum.de/webshare/u/toker/dynnet_training_splits/ under a common tarball for simpler download speeds.
Individual zip files were replaced with tarballs instead.
In the mediatum server version the following directories have the wrong name compared to the given split txt files:
/labels/5111_4560_13_38S
/labels/6204_3495_13_46N
/labels/7026_3201_13_52N
/labels/7367_5050_13_54S
/labels/2459_4406_13_19S… See the full description on the dataset page: https://huggingface.co/datasets/torchgeo/dynamic_earthnet.shadow-eo
Dataset Card for S-EO: A Large-Scale Dataset for Geometry-Aware Shadow Detection in Remote Sensing Applications
Project page
We introduce the S-EO dataset: a large-scale, high-resolution dataset designed to advance geometry-aware shadow detection. Collected from diverse public-domain sources, including challenge datasets and government providers such as USGS, our dataset comprises 702 georeferenced tiles across the USA, each covering 500 × 500 meters. Each tile includes multi-date… See the full description on the dataset page: https://huggingface.co/datasets/emasquil/shadow-eo.EyecandiesEmoArt-130k
EmoArt: A Large-Scale Emotion-Annotated Artistic Dataset
Overview
EmoArt is a comprehensive, large-scale emotion-annotated artistic dataset containing 132,664 high-resolution artworks spanning 56 painting styles across 7 thematic categories. This dataset bridges the gap between visual art and emotional computing, enabling groundbreaking research in emotion-aware AI systems.
Key Statistics
📊 132,664 artworks with rich emotional annotations
🎨 56 distinct… See the full description on the dataset page: https://huggingface.co/datasets/printblue/EmoArt-130k.ESpeech-webinars2
Webinar Audio Dataset
Dataset Description
This dataset contains 850 hours processed webinar audio segments with corresponding metadata. Each audio file represents a segment extracted from webinar recordings, processed at 44.1kHz sample rate.
Dataset Summary
Language: Russian
Task: TTS, ASR, Quality Asessment
Audio format: MP3, 44.1kHz sample rate
Structure: Segmented audio files with JSON metadata
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-webinars2.Molmo2-ER-RefSpatial
Molmo2-ER · JingkunAn/RefSpatial
2.5M spatial-referring corpus (web + indoor + simulated) covering 31 spatial relations.
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
Upstream source
Original dataset: JingkunAn/RefSpatial
Paper: RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics (arXiv:2506.04308)… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-RefSpatial.e621-2024-webp-4MpixelDataset Description:
This is a processed version of the https://huggingface.co/datasets/boxingscorpionbagel/e621-2024 dataset, primarily prepared for personal use in future projects.
Therefore, for licensing and other legal information, please refer to the original project.
You can directly download tar file,or use https://deepghs.github.io/hfutils/main/api_doc/index/fetch.html#hf-tar-file-download to download anything .webp file you want.
The following modifications have been made to the… See the full description on the dataset page: https://huggingface.co/datasets/NebulaeWis/e621-2024-webp-4Mpixel.ImagePulseV2-Edit-Structure
ImagePulseV2 Dataset - Image Structure
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It comprises multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Structure.shitspotter
Dataset Card for ShitSpotter ("ScatSpotter")
ShitSpotter (or "ScatSpotter" in formal settings) is an open dataset of images containing dog feces.
This dataset contains full-resolution smartphone images of dog feces ("poop") collected in urban outdoor
environments taken using a "before/after/negative" protocol.
It includes thousands of polygon annotations of feces in varied lighting, seasonal, and terrain conditions.
The dataset is designed for training and evaluating object… See the full description on the dataset page: https://huggingface.co/datasets/erotemic/shitspotter.Echo-4o-Image
Echo-4o-Image Dataset
Paper | Project Page | Code
Introduction
Echo-4o-Image is a 180K-scale synthetic dataset generated by GPT-4o, designed to advance open-source models in image generation. While real-world image datasets are valuable, synthetic images offer crucial advantages, especially in addressing blind spots in real-world coverage:
Complementing Rare Scenarios: Synthetic data can generate examples for scenarios less represented in real-world datasets, such as… See the full description on the dataset page: https://huggingface.co/datasets/Yejy53/Echo-4o-Image.Swift-OpenX-Embodiment-wrist-imageshistorical-appraisals-ocr-ml
Digitized Historical Property Valuations and Building Feature Data
A dataset of historical property valuations and contemporary parcel/building feature data created by the process described in this ACM COMPASS paper: TODO add link when published
Data description
hamilton_county
Data for properties in Hamilton County, Ohio (primarily Cincinnati). All the data sources below are linked using the land parcel identifier parcelid.
property_cards/: folder with tar balls of… See the full description on the dataset page: https://huggingface.co/datasets/eruka-cmu-housing/historical-appraisals-ocr-ml.Semi-Truths-Evalset
Semi-Truths: The Evaluation Sample
Recent efforts have developed AI-generated image detectors claiming robustness against various augmentations, but their effectiveness remains unclear. Can these systems detect varying degrees of augmentation?
To address these questions, we introduce Semi-Truths, featuring 27,600 real images, 245,300 masks, and 850,200 AI-augmented images featuring varying degrees of targeted and localized edits, created using diverse augmentation methods… See the full description on the dataset page: https://huggingface.co/datasets/semi-truths/Semi-Truths-Evalset.Inter-Edit-Test
Inter-Edit-Test
Official test benchmark release for the CVPR 2026 paper:
Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing
This repository hosts the public release of Inter-Edit-Test, a human-annotated benchmark for the Interactive Instruction-based Image Editing (I^3E) task.
Each sample contains:
a source image,
a coarse user-style interaction mask,
a concise editing instruction,
and a ground-truth edited image.
To simplify large-scale distribution on… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Inter-Edit-Test.emonet-face-big
EmoNet-Face: A Fine-Grained, Expert-Annotated Benchmark for Facial Emotion Recognition
Dataset Summary
EmoNet-Face is a comprehensive benchmark suite designed to address critical gaps in facial emotion recognition (FER). Current benchmarks often have a narrow emotional spectrum, lack demographic diversity, and use uncontrolled imagery. EmoNet-Face provides a robust foundation for developing and evaluating AI systems with a deeper, more nuanced understanding of human… See the full description on the dataset page: https://huggingface.co/datasets/laion/emonet-face-big.imagenet-gen-sd1.5
ImageNet Generated using Stable Diffusion v1.5
The following repository mimics the size and class structure of the original ImageNet database. The classes can be found in the classes.txt file.
This dataset contains approximately 1300 images per class over 1000 classes for a total of 1.3 million images.
Here is an excerpt from classes.txt:
0 tench, Tinca tinca
1 goldfish, Carassius auratus
2 great white shark, white shark, man-eater, man-eating shark, Carcharodon caharias
3 tiger… See the full description on the dataset page: https://huggingface.co/datasets/ek826/imagenet-gen-sd1.5.
