datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen-Image-Bench
Qwen-Image-Bench
A creator-centric benchmark for evaluating Text-to-Image models beyond semantic alignment.
Links
Resource
Link
📑 Paper
http://arxiv.org/abs/2605.28091
📊 Benchmark Dataset (HuggingFace)
https://huggingface.co/datasets/Qwen/Qwen-Image-Bench
📊 Benchmark Dataset (ModelScope)
https://www.modelscope.cn/datasets/Qwen/Qwen-Image-Bench
💻 GitHub
https://github.com/QwenLM/Qwen-Image-Bench
🧑⚖️ Q-Judger Model… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/Qwen-Image-Bench.Nemotron-Image-Training-v3
Nemotron Image Training v3
Versions
Date
Commit
Changes
2026-04-28
HEAD
Initial commit.
Dataset Description
Nemotron Image Training v3 is a collection of image-centric multimodal training data for vision–language models. Similar to Nemotron-VLM-Dataset v2, it was curated as a large-scale, multi-subdataset release where each subset ships a standardized conversation JSONL alongside a dataset card describing sources, licensing, and media layout.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Image-Training-v3.imagenet_hard_review_data_r2ShareGPT-4o-Image
📚 ShareGPT-4o-Image
ShareGPT-4o-Image is a large-scale and high-quality image generation dataset, where all images are produced by GPT-4o’s image generation capabilities. This dataset is designed to align open multimodal models with GPT-4o’s strengths in visual content creation. It includes 45K text-to-image and 46K text-and-image-to-image samples, making it a useful resource for enhancing multimodal models in both image generation and editing tasks.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ShareGPT-4o-Image.sat-image-boundingbox-sft-full
NU-TONIC raw SFT Full
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2)
Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.ImageEval-ArabicNLP26
ImageEval-ArabicNLP26 👁️
ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026.
It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation.
The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.Nemotron-Image-Training-v3
Nemotron Image Training v3
Versions
Date
Commit
Changes
2026-04-28
HEAD
Initial commit.
Dataset Description
Nemotron Image Training v3 is a collection of image-centric multimodal training data for vision–language models. Similar to Nemotron-VLM-Dataset v2, it was curated as a large-scale, multi-subdataset release where each subset ships a standardized conversation JSONL alongside a dataset card describing sources, licensing, and media layout.… See the full description on the dataset page: https://huggingface.co/datasets/liyah1616/Nemotron-Image-Training-v3.imagenet_hard_review_dataImagewebdsh-images
webdsh-images
Disk images for the emulated machines webdsh
offers.
Why this exists
v86 can run about a hundred and twenty-five
machines in a browser, and every one of them is the same emulator with a
different disk. What copy.sh/v86 has that a fork does
not is a CDN with the disks on it: its own host, i.copy.sh, refuses browser
requests from anywhere else — deliberately, and it is their bandwidth to
protect.
So webdsh's catalog was complete and its machines were… See the full description on the dataset page: https://huggingface.co/datasets/AndyZijianZhang/webdsh-images.ramanv-image-editing
ramanv-image-editing
Image editing dataset for training FLUX.1-Kontext / InstructPix2Pix style models.
Size
592,141 total editing pairs
Sources: ultraedit
Schema
Each shard tar contains {uid}_src.jpg, {uid}_edit.jpg, {uid}_mask.png (where available).
Metadata per record: instruction, prompt, edit_type, caption_before/after, license, sha256.
Licenses
MagicBrush, InstructPix2Pix, Pico-Banana, HumanEdit: CC-BY-4.0
UltraEdit, AnyEdit… See the full description on the dataset page: https://huggingface.co/datasets/lingamvamshikrishnareddy/ramanv-image-editing.sat-image-boundingbox-sft
NU-TONIC raw SFT init
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning LFM-VL (leap-finetune vlm_sft format). JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (default HF source: stochastic/random_streetview_images_pano_v0.0.2) via download_geoguessr_poi_imagery.py.
Optical: multispectral optical COGs from a public STAC catalog… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft.ramanv-image-textrenderImageCoDe
Dataset Card for ImageCoDe
To get started quickly, load descriptions via:
from datasets import load_dataset
examples = load_dataset('BennoKrojer/ImageCoDe')
And download image_sets.zip for all images sets (each directory consisting of 10 images).
Dataset Summary
We introduce ImageCoDe, a vision-and-language benchmark that requires contextual language understanding in the form of pragmatics, temporality, long descriptions and visual nuances. The task: Given a detailed… See the full description on the dataset page: https://huggingface.co/datasets/BennoKrojer/ImageCoDe.ImagePulseV2-Edit-Change
ImagePulseV2 Dataset - Foreground Editing
The ImagePulseV2 dataset is a custom-built dataset we created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Change.zoya-image-1-experiments
ZOYA IMAGE-1 — Reproducible GGUF Experiments
Purpose
This dataset stores reproducible ZOYA IMAGE-1 image-generation
experiments together with the exact generation parameters,
model identities, SHA256 fingerprints, and validation reports.
The package is designed for controlled comparisons where the
tested variable is changed explicitly and all other relevant
variables remain fixed.
Current baseline
Experiment ID: ZOYA_PHASE0_BASELINE_00001… See the full description on the dataset page: https://huggingface.co/datasets/tigerking009/zoya-image-1-experiments.imagenet_sketch_resizedImagePulseV2-Edit-AddRemove
ImagePulseV2 Dataset - Local Add/Delete
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Models:… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-AddRemove.med-imagesVideoEspresso_train_multi_image
VideoEspresso
This dataset is the multi-image version.
Leaderboard
Model
Params
Frames
Overall
Narrative Analysis
Event Dynamic
Preparation Steps
Causal Analysis
Theme Analysis
Contextual Analysis
Influence Analysis
Role Analysis
Interaction Analysis
Behavior Analysis
Emotion Analysis
Cooking Process
Traffic Analysis
Situation Analysis
LLaVA-Video
72B
64
66.3%
68.4%
66.2%
74.5%
62.7%
62.3%
71.6%
62.5%
63.5%
67.7%
63.2%
60.0%
75.5%
76.7%
74.0%
LLaVA-OneVision… See the full description on the dataset page: https://huggingface.co/datasets/hshjerry0315/VideoEspresso_train_multi_image.PyVision-Image-SFT-Data
PyVision-Image-RL-Data
Project Page | Paper | GitHub
This repository contains the reinforcement learning (RL) data used to train PyVision-Image-RL, as presented in the paper "PyVision-RL: Forging Open Agentic Vision Models via RL".
PyVision-RL is a reinforcement learning framework for open-weight multimodal models designed to stabilize training and sustain interaction, preventing interaction collapse and encouraging multi-turn tool use in agentic tasks.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/PyVision-Image-SFT-Data.arxiv-chandra-ocr-2-include-images-first50-20260415
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Source paper IDs in input list: 27,584
Processed IDs recorded in state/processed_ids.txt: 50
Successes: 50
Partial successes: 0
Errors: 0
Next shard index: 10
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.imagenet-metric-refsMMH3_Image_Edit_WorkflowThis is just an example of using MiniMax H3 as an image editor. The actual workflow that I use requires several custom nodes, some of which are not published, so this one is simply a bare bones demonstration.
This uses the hybrid MiniMax H3 model from here: https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/tree/main
It uses the custom VAE from here: https://huggingface.co/Mamad8/MiniMax-H3-Image-VAE/tree/main
It uses the LoRA from here:… See the full description on the dataset page: https://huggingface.co/datasets/fizzlepoof/MMH3_Image_Edit_Workflow.wikipedia-image-zh-tw
中文維基百科圖文配對資料集
一圖一列,附圖說、替代文字、所在條目與章節。繁體(wiki_images_dataset.jsonl)
與簡體(wiki_images_dataset_CN.jsonl)各一份。
📅 目前版本:2609(維基百科 dump 日期:2026/9/1)
HF 的預設 train split 會讀入繁簡兩個檔案。只需要其中一種字體時,請明確指定:
from datasets import load_dataset
tw = load_dataset("yuhuanstudio/wikipedia-image-zh-tw", data_files="wiki_images_dataset.jsonl", split="train")
cn = load_dataset("yuhuanstudio/wikipedia-image-zh-tw", data_files="wiki_images_dataset_CN.jsonl", split="train")
欄位
欄位
說明… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/wikipedia-image-zh-tw.ImagePulseV2-Edit-Pose
ImagePulseV2 Dataset - Pose Adjustment
The ImagePulseV2 dataset is a custom-built dataset created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Pose.ImagePulseV2-Edit-Style
ImagePulseV2 Dataset - Style Transfer
The ImagePulseV2 dataset is a custom-built dataset we created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Style.LoRA-Merge-ImagesPyVision-Image-RL-Data
PyVision-Image-RL-Data
Project Page | Paper | GitHub
This repository contains the Reinforcement Learning (RL) training data used to train PyVision-Image-RL, as presented in the paper "PyVision-RL: Forging Open Agentic Vision Models via RL".
Dataset Summary
PyVision-RL is a reinforcement learning framework for open-weight multimodal models designed to stabilize training and sustain interaction in agentic tasks. This dataset specifically supports the training of… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/PyVision-Image-RL-Data.Qwen-Image-Bench
Qwen-Image-Bench
A creator-centric benchmark for evaluating Text-to-Image models beyond semantic alignment.
Links
Resource
Link
📑 Paper
http://arxiv.org/abs/2605.28091
📊 Benchmark Dataset (HuggingFace)
https://huggingface.co/datasets/Qwen/Qwen-Image-Bench
📊 Benchmark Dataset (ModelScope)
https://www.modelscope.cn/datasets/Qwen/Qwen-Image-Bench
💻 GitHub
https://github.com/QwenLM/Qwen-Image-Bench
🧑⚖️ Q-Judger Model… See the full description on the dataset page: https://huggingface.co/datasets/KilersBotz/Qwen-Image-Bench.
