datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
witYFCC100M_OpenAI_subsetThe YFCC100M is one of the largest publicly and freely useable multimedia collection, containing the metadata of around 99.2 million photos and 0.8 million videos from Flickr, all of which were shared under one of the various Creative Commons licenses.
This version is a subset defined in openai/CLIP.dalle-3-dataset
Dataset Card for LAION DALL·E 3 Discord Dataset
Description: This dataset consists of caption and image pairs scraped from the LAION share-dalle-3 discord channel. The purpose is to collect image-text pairs for research and exploration.
Source Code: The code used to generate this data can be found here.
Contributors
Zach Nagengast
Eduardo Pach
Seva Maltsev
Ben Egan
The LAION community
Data Attributes
caption: The text description or prompt associated with… See the full description on the dataset page: https://huggingface.co/datasets/OpenDatasets/dalle-3-dataset.open-imagessynthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.dallestreet
Citation Information
@misc{mukherjee2024crossroadscontinentsautomatedartifact,
title={Crossroads of Continents: Automated Artifact Extraction for Cultural Adaptation with Large Multimodal Models},
author={Anjishnu Mukherjee and Ziwei Zhu and Antonios Anastasopoulos},
year={2024},
eprint={2407.02067},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.02067},
}
700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset
Rapidata Image Generation Preference Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.vqgan-pairs
VQGAN Pairs
This dataset contains ~2.4 million image pairs intended for improvement of image quality in VQGAN predictions. Each pair consists of:
A 512x512 crop of an image taken from Open Images.
A 256x256 image encoded and decoded using VQGAN, corresponding to the same image crop as the original.
This is the VQGAN implementation that was used for encoding and decoding: https://github.com/patil-suraj/vqgan-jax
License
This dataset is created using Open Images… See the full description on the dataset page: https://huggingface.co/datasets/dalle-mini/vqgan-pairs.Flux_SD3_MJ_Dalle_Human_Alignment_Dataset
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.jobs-dalle-2
Dataset Card for "dataset-dalle"
More Information needed
dalle-3-datasetUse the Edit dataset card button to edit.
Dalle310,000 high-quality captions with image pairs produced by dalle3 with a raw.zip incase i uploaded it wrong.
dalle-3-images
🎨 DALL•E 3 Images Dataset
This is datase with images made by Dalle3.
Dataset parameters
Count of images: 3310
Zip file with dataset: True
Captions with images: False
License
License for this dataset: MIT
Use in datasets
pip install -q datasets
from datasets import load_dataset
dataset = load_dataset(
"ehristoforu/dalle-3-images",
revision="main"
)
Enjoy with this dataset!
Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.prof_images_blip__dalle-2
Dataset Card for "prof_images_blip__dalle-2"
More Information needed
animals_with_fruits_dalle3
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
A Hugging Face Datasets repository accompanying the paper "Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings".
Code: https://github.com/aziksh-ospanov/scendi-score
Dataset Information
This dataset consists of images depicting various animals eating different kinds of fruits, generated using DALL-E 3. It is released as a companion to the… See the full description on the dataset page: https://huggingface.co/datasets/aziksh/animals_with_fruits_dalle3.Dalle-3-1MThis dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically share their best results online, this dataset reflects a diverse and high quality compilation of human preferences and high quality creative works. Captions for the images were generated using 4-bit CogVLM with custom… See the full description on the dataset page: https://huggingface.co/datasets/bitmind/Dalle-3-1M.dalle2-facesmidjourney-dalle-sd-nanobananapro-dataset
Dataset Card: Midjourney, DALL-E, Stable Diffusion & Nano Banana Pro vs Real Images
Description
Dataset de classification binaire pour détecter les images générées par IA (Midjourney, DALL-E, Stable Diffusion et Nano Banana Pro) vs images réelles.
Dataset Structure
Train set: 10,000 images
Real: 5000 images
Fake (AI-generated): 5000 images
Test set: 2,000 images
Real: 1000 images
Fake (AI-generated): 1000 images
Features
{
"image": Image… See the full description on the dataset page: https://huggingface.co/datasets/julienlucas/midjourney-dalle-sd-nanobananapro-dataset.test-dalle-3This is a test database. Please ignore.
prof_report__dalle-2__multi__24
Dataset Card for "prof_report__dalle-2__multi__24"
More Information needed
dalle-3-dataset
Dataset Card for LAION DALL·E 3 Discord Dataset
Description: This dataset consists of caption and image pairs scraped from the LAION share-dalle-3 discord channel. The purpose is to collect image-text pairs for research and exploration.
Source Code: The code used to generate this data can be found here.
Contributors
Zach Nagengast
Eduardo Pach
Seva Maltsev
Ben Egan
The LAION community
Data Attributes
caption: The text description or prompt… See the full description on the dataset page: https://huggingface.co/datasets/ShishirB434/dalle-3-dataset.dalle-3-contrastive-captions-updated
Dataset Card for "dalle-3-contrastive-captions-updated"
More Information needed
midjourney-dalle-sd-dataset
Dataset Card: Midjourney, DALL-E, Stable Diffusion vs Real Images
Description
Dataset de classification binaire pour détecter les images générées par IA (Midjourney, DALL-E, Stable Diffusion) vs images réelles.
Dataset Structure
Train set: 5,000 images
Real: 2,500 images
Fake (AI-generated): 2,500 images
Test set: 1,000 images
Real: 500 images
Fake (AI-generated): 500 images
Features
{
"image": Image,
"label": "real" | "fake"
}… See the full description on the dataset page: https://huggingface.co/datasets/julienlucas/midjourney-dalle-sd-dataset.prof_report__dalle-2__multi__24
Dataset Card for "prof_report__dalle-2__multi__24"
More Information needed
identities-dalle-2
Dataset Card for "identities-dalle-2"
More Information needed
dalle-3-contrastive-captions
Dataset Card for "dalle-3-contrastive-captions"
More Information needed
DALL-E-Prompts-OpenAI-ChatGPT
Dataset Card for Dataset Name
Dataset Summary
This dataset has been generated using Prompt Generator for OpenAI's DALL-E.
Languages
English
Dataset Structure
1.000.000 Prompts
dalle3-eval-samples
DALL-E 3 Evaluation Samples
This repository contains text-to-image samples collected for the evaluations of DALL-E 3 in the whitepaper. We provide samples not only from DALL-E 3, but from the competitors we compare against in the paper.
The intent of this repository is to enable researchers in the text-to-image space to reproduce our results and foster forward progress of the text-to-image field as a whole. The samples from this repository are not meant to be demonstrations of the… See the full description on the dataset page: https://huggingface.co/datasets/johko/dalle3-eval-samples.dalle-3-palette
