CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Qwen /Qwen-Image-Bench Qwen-Image-Bench A creator-centric benchmark for evaluating Text-to-Image models beyond semantic alignment. Links Resource Link 📑 Paper http://arxiv.org/abs/2605.28091 📊 Benchmark Dataset (HuggingFace) https://huggingface.co/datasets/Qwen/Qwen-Image-Bench 📊 Benchmark Dataset (ModelScope) https://www.modelscope.cn/datasets/Qwen/Qwen-Image-Bench 💻 GitHub https://github.com/QwenLM/Qwen-Image-Bench 🧑‍⚖️ Q-Judger Model… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/Qwen-Image-Bench.imageimage-to-text1K<n<10K49 likes7.3k downloads4mo agoHugging Face02nvidia /Nemotron-Image-Training-v3 Nemotron Image Training v3 Versions Date Commit Changes 2026-04-28 HEAD Initial commit. Dataset Description Nemotron Image Training v3 is a collection of image-centric multimodal training data for vision–language models. Similar to Nemotron-VLM-Dataset v2, it was curated as a large-scale, multi-subdataset release where each subset ships a standardized conversation JSONL alongside a dataset card describing sources, licensing, and media layout.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Image-Training-v3.textvisual-question-answering1M<n<10M82 likes7.1k downloads5mo agoHugging Face03taesiri /imagenet_hard_review_data_r2tabular1K<n<10K0 likes6.1k downloads3y agoHugging Face04FreedomIntelligence /ShareGPT-4o-Image 📚 ShareGPT-4o-Image ShareGPT-4o-Image is a large-scale and high-quality image generation dataset, where all images are produced by GPT-4o’s image generation capabilities. This dataset is designed to align open multimodal models with GPT-4o’s strengths in visual content creation. It includes 45K text-to-image and 46K text-and-image-to-image samples, making it a useful resource for enhancing multimodal models in both image generation and editing tasks. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ShareGPT-4o-Image.texttext-to-image10K<n<100K102 likes6k downloads1y agoHugging Face05NuTonic /sat-image-boundingbox-sft-full NU-TONIC raw SFT Full Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters. Provenance Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2) Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8. Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.imageimage-text-to-text100K<n<1M14 likes5.2k downloads5mo agoHugging Face06QCRI /ImageEval-ArabicNLP26 ImageEval-ArabicNLP26 👁️ ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026. It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation. The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.audio10K<n<100K4 likes2.5k downloads22d agoHugging Face07liyah1616 /Nemotron-Image-Training-v3 Nemotron Image Training v3 Versions Date Commit Changes 2026-04-28 HEAD Initial commit. Dataset Description Nemotron Image Training v3 is a collection of image-centric multimodal training data for vision–language models. Similar to Nemotron-VLM-Dataset v2, it was curated as a large-scale, multi-subdataset release where each subset ships a standardized conversation JSONL alongside a dataset card describing sources, licensing, and media layout.… See the full description on the dataset page: https://huggingface.co/datasets/liyah1616/Nemotron-Image-Training-v3.textvisual-question-answering1M<n<10M0 likes1.4k downloads5mo agoHugging Face08taesiri /imagenet_hard_review_datatabular1K<n<10K0 likes888 downloads3y agoHugging Face09TexasNotFound /Imagetabular1K<n<10K0 likes842 downloads18d agoHugging Face10AndyZijianZhang /webdsh-images webdsh-images Disk images for the emulated machines webdsh offers. Why this exists v86 can run about a hundred and twenty-five machines in a browser, and every one of them is the same emulator with a different disk. What copy.sh/v86 has that a fork does not is a CDN with the disks on it: its own host, i.copy.sh, refuses browser requests from anywhere else — deliberately, and it is their bandwidth to protect. So webdsh's catalog was complete and its machines were… See the full description on the dataset page: https://huggingface.co/datasets/AndyZijianZhang/webdsh-images.geospatialn<1K0 likes572 downloads26d agoHugging Face11lingamvamshikrishnareddy /ramanv-image-editinggated ramanv-image-editing Image editing dataset for training FLUX.1-Kontext / InstructPix2Pix style models. Size 592,141 total editing pairs Sources: ultraedit Schema Each shard tar contains {uid}_src.jpg, {uid}_edit.jpg, {uid}_mask.png (where available). Metadata per record: instruction, prompt, edit_type, caption_before/after, license, sha256. Licenses MagicBrush, InstructPix2Pix, Pico-Banana, HumanEdit: CC-BY-4.0 UltraEdit, AnyEdit… See the full description on the dataset page: https://huggingface.co/datasets/lingamvamshikrishnareddy/ramanv-image-editing.image1K<n<10K6 likes523 downloads20d agoHugging Face12NuTonic /sat-image-boundingbox-sft NU-TONIC raw SFT init Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning LFM-VL (leap-finetune vlm_sft format). JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters. Provenance Locations: GeoGuessr-style POIs (default HF source: stochastic/random_streetview_images_pano_v0.0.2) via download_geoguessr_poi_imagery.py. Optical: multispectral optical COGs from a public STAC catalog… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft.imageimage-text-to-text10K<n<100K0 likes408 downloads5mo agoHugging Face13lingamvamshikrishnareddy /ramanv-image-textrendergatedtextn<1K2 likes353 downloads1mo agoHugging Face14BennoKrojer /ImageCoDe Dataset Card for ImageCoDe To get started quickly, load descriptions via: from datasets import load_dataset examples = load_dataset('BennoKrojer/ImageCoDe') And download image_sets.zip for all images sets (each directory consisting of 10 images). Dataset Summary We introduce ImageCoDe, a vision-and-language benchmark that requires contextual language understanding in the form of pragmatics, temporality, long descriptions and visual nuances. The task: Given a detailed… See the full description on the dataset page: https://huggingface.co/datasets/BennoKrojer/ImageCoDe.image10K<n<100K5 likes292 downloads4y agoHugging Face15DiffSynth-Studio /ImagePulseV2-Edit-Change ImagePulseV2 Dataset - Foreground Editing The ImagePulseV2 dataset is a custom-built dataset we created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB. Open-source code: DiffSynth-Studio Technical report: arXiv Project homepage: GitHub Documentation: English Version, Chinese Version Online demo: ModelScope Studio… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Change.text10K<n<100K0 likes266 downloads5mo agoHugging Face16tigerking009 /zoya-image-1-experiments ZOYA IMAGE-1 — Reproducible GGUF Experiments Purpose This dataset stores reproducible ZOYA IMAGE-1 image-generation experiments together with the exact generation parameters, model identities, SHA256 fingerprints, and validation reports. The package is designed for controlled comparisons where the tested variable is changed explicitly and all other relevant variables remain fixed. Current baseline Experiment ID: ZOYA_PHASE0_BASELINE_00001… See the full description on the dataset page: https://huggingface.co/datasets/tigerking009/zoya-image-1-experiments.imageimage-to-imagen<1K0 likes250 downloads29d agoHugging Face17janellecai /imagenet_sketch_resized10K<n<100K0 likes245 downloads2y agoHugging Face18DiffSynth-Studio /ImagePulseV2-Edit-AddRemove ImagePulseV2 Dataset - Local Add/Delete The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB. Open-source code: DiffSynth-Studio Technical report: arXiv Project homepage: GitHub Documentation: English Version, Chinese Version Online demo: ModelScope Studio Models:… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-AddRemove.text10K<n<100K0 likes242 downloads5mo agoHugging Face19wanqi16 /med-imagestextn<1K0 likes230 downloads8mo agoHugging Face20hshjerry0315 /VideoEspresso_train_multi_image VideoEspresso This dataset is the multi-image version. Leaderboard Model Params Frames Overall Narrative Analysis Event Dynamic Preparation Steps Causal Analysis Theme Analysis Contextual Analysis Influence Analysis Role Analysis Interaction Analysis Behavior Analysis Emotion Analysis Cooking Process Traffic Analysis Situation Analysis LLaVA-Video 72B 64 66.3% 68.4% 66.2% 74.5% 62.7% 62.3% 71.6% 62.5% 63.5% 67.7% 63.2% 60.0% 75.5% 76.7% 74.0% LLaVA-OneVision… See the full description on the dataset page: https://huggingface.co/datasets/hshjerry0315/VideoEspresso_train_multi_image.text100K<n<1M0 likes223 downloads1y agoHugging Face21Agents-X /PyVision-Image-SFT-Data PyVision-Image-RL-Data Project Page | Paper | GitHub This repository contains the reinforcement learning (RL) data used to train PyVision-Image-RL, as presented in the paper "PyVision-RL: Forging Open Agentic Vision Models via RL". PyVision-RL is a reinforcement learning framework for open-weight multimodal models designed to stabilize training and sustain interaction, preventing interaction collapse and encouraging multi-turn tool use in agentic tasks. Citation… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/PyVision-Image-SFT-Data.textimage-text-to-text1K<n<10K3 likes198 downloads7mo agoHugging Face22nielsr /arxiv-chandra-ocr-2-include-images-first50-20260415 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415 Source paper IDs in input list: 27,584 Processed IDs recorded in state/processed_ids.txt: 50 Successes: 50 Partial successes: 0 Errors: 0 Next shard index: 10 Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.imagen<1K0 likes130 downloads5mo agoHugging Face23wittenator /imagenet-metric-refstabularn<1K0 likes124 downloads9d agoHugging Face24fizzlepoof /MMH3_Image_Edit_WorkflowThis is just an example of using MiniMax H3 as an image editor. The actual workflow that I use requires several custom nodes, some of which are not published, so this one is simply a bare bones demonstration. This uses the hybrid MiniMax H3 model from here: https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/tree/main It uses the custom VAE from here: https://huggingface.co/Mamad8/MiniMax-H3-Image-VAE/tree/main It uses the LoRA from here:… See the full description on the dataset page: https://huggingface.co/datasets/fizzlepoof/MMH3_Image_Edit_Workflow.tabularn<1K1 likes122 downloads1mo agoHugging Face25yuhuanstudio /wikipedia-image-zh-tw 中文維基百科圖文配對資料集 一圖一列,附圖說、替代文字、所在條目與章節。繁體(wiki_images_dataset.jsonl) 與簡體(wiki_images_dataset_CN.jsonl)各一份。 📅 目前版本:2609(維基百科 dump 日期:2026/9/1) HF 的預設 train split 會讀入繁簡兩個檔案。只需要其中一種字體時,請明確指定: from datasets import load_dataset tw = load_dataset("yuhuanstudio/wikipedia-image-zh-tw", data_files="wiki_images_dataset.jsonl", split="train") cn = load_dataset("yuhuanstudio/wikipedia-image-zh-tw", data_files="wiki_images_dataset_CN.jsonl", split="train") 欄位 欄位 說明… See the full description on the dataset page: https://huggingface.co/datasets/yuhuanstudio/wikipedia-image-zh-tw.imageimage-to-text1M<n<10M1 likes120 downloads19d agoHugging Face26DiffSynth-Studio /ImagePulseV2-Edit-Pose ImagePulseV2 Dataset - Pose Adjustment The ImagePulseV2 dataset is a custom-built dataset created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB. Open-source code: DiffSynth-Studio Technical report: arXiv Project homepage: GitHub Documentation: English Version, Chinese Version Online demo: ModelScope Studio Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Pose.text10K<n<100K0 likes118 downloads5mo agoHugging Face27DiffSynth-Studio /ImagePulseV2-Edit-Style ImagePulseV2 Dataset - Style Transfer The ImagePulseV2 dataset is a custom-built dataset we created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB. Open-source code: DiffSynth-Studio Technical report: arXiv Project homepage: GitHub Documentation: English Version, Chinese Version Online demo: ModelScope Studio Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Style.text10K<n<100K0 likes116 downloads5mo agoHugging Face28Sugita-daichi /LoRA-Merge-Imagesimageimage-to-image100K<n<1M0 likes115 downloads5mo agoHugging Face29Agents-X /PyVision-Image-RL-Data PyVision-Image-RL-Data Project Page | Paper | GitHub This repository contains the Reinforcement Learning (RL) training data used to train PyVision-Image-RL, as presented in the paper "PyVision-RL: Forging Open Agentic Vision Models via RL". Dataset Summary PyVision-RL is a reinforcement learning framework for open-weight multimodal models designed to stabilize training and sustain interaction in agentic tasks. This dataset specifically supports the training of… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/PyVision-Image-RL-Data.textimage-text-to-text10K<n<100K1 likes112 downloads7mo agoHugging Face30KilersBotz /Qwen-Image-Bench Qwen-Image-Bench A creator-centric benchmark for evaluating Text-to-Image models beyond semantic alignment. Links Resource Link 📑 Paper http://arxiv.org/abs/2605.28091 📊 Benchmark Dataset (HuggingFace) https://huggingface.co/datasets/Qwen/Qwen-Image-Bench 📊 Benchmark Dataset (ModelScope) https://www.modelscope.cn/datasets/Qwen/Qwen-Image-Bench 💻 GitHub https://github.com/QwenLM/Qwen-Image-Bench 🧑‍⚖️ Q-Judger Model… See the full description on the dataset page: https://huggingface.co/datasets/KilersBotz/Qwen-Image-Bench.imageimage-to-text1K<n<10K0 likes111 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.