datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midjourney-v6-recap
Midjourney v6 Recaptioned
~1.2M Midjourney v6 images with captions from three VLMs:
llava: Original LLaVA captions from the source dataset
gemini: Gemini Flash 1.5 captions
qwen3: Qwen3 VL 8B captions
Caption coverage
llava: available for all 1,235,432 images (from original dataset)
gemini and qwen3: available for 1,017,105 images (82.3%)
Source
Based on brivangl/midjourney-v6-llava.
i1-midjourneyv6-tfrecordi1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu
Princeton University
[arXiv][code][model][project page]
Overview
To prepare the dataset for training, we store the image-caption pairs as TFRecords.
This HuggingFace dataset contains the TFRecords corresponding to the midjourneyv6 dataset at 256×256 resolution.
It also serves as an example of what a dataset processed using… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/i1-midjourneyv6-tfrecord.i1-midjourneyv6-512-resolution-1m-tfrecordi1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu
Princeton University
[arXiv][code][model][project page]
Overview
To prepare the dataset for training, we store the image-caption pairs as TFRecords.
This HuggingFace dataset contains the TFRecords corresponding to the midjourneyv6 dataset at 512×512 resolution. Concretely, we only retain raw images with a shorter edge of at… See the full description on the dataset page: https://huggingface.co/datasets/i1-datasets/i1-midjourneyv6-512-resolution-1m-tfrecord.midjourney-v6-llavaThis dataset based on https://huggingface.co/datasets/CortexLM/midjourney-v6 dataset, captioned with LLava-1.6 model.
This dataset was released as is. By accessing and using this dataset, you acknowledge and agree that Cortex Foundation and the author of this repo are not responsible for any copyright violations or legal consequences that may arise from the use of these images.
i1-midjourneyv6-1024-resolution-1m-tfrecordi1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu
Princeton University
[arXiv][code][model][project page]
Overview
To prepare the dataset for training, we store the image-caption pairs as TFRecords.
This HuggingFace dataset contains the TFRecords corresponding to the midjourneyv6 dataset at 1024×1024 resolution. Concretely, we only retain raw images with a shorter edge of… See the full description on the dataset page: https://huggingface.co/datasets/i1-datasets/i1-midjourneyv6-1024-resolution-1m-tfrecord.midjourney-messages
midjourney-messages
Description
This dataset contains the raw messages from Midjourney.
Total messages: 55,082,563
midjourney-niji-1m-llavanext
Dataset Card for midjourney-niji-1m-llavanext
Dataset Summary
This is a dataset of 2,079,886 synthetic captions for 1,039,943 images from midjourney-v6-520k-raw and nijijourney-v6-520k-raw. The captions were produced using https://huggingface.co/lmms-lab/llama3-llava-next-8b inferenced in float16 after tags were generated with wd-swinv2-tagger-v3, followed by cleanup and shortening with Meta-Llama-3-8B.
All images with metadata are available as MozJPEG encoded JPEGs… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/midjourney-niji-1m-llavanext.midjourney-messages
midjourney-messages
Description
This dataset contains the raw messages from Midjourney.
Initial dataset is https://huggingface.co/datasets/vivym/midjourney-messages, but this one has the images attached.
AIGeneratedImages_Midjourney
AI Generated Image for Image Classification
This dataset contains AI generated images by Midjourney and Human images taken from Imagenet. The dataset is meant for Image Classification tasks.
Dataset Details
Dataset Description
Curated by: Deepankar Sharma
GenImage_MidJourneymidjourney-images
⛵ Midjourney Images Dataset
This is datase with images made by Midjourney V5/V6.
Dataset parameters
Count of images: ~10.000
Zip file with dataset: True
Captions with images: False
License
License for this dataset: MIT
Use in datasets
pip install -q datasets
from datasets import load_dataset
dataset = load_dataset(
"ehristoforu/midjourney-images",
revision="main"
)
Enjoy with this dataset!
midjourney-prompts-embeddings
Midjourney Prompt–Embedding Dataset
This dataset is derived from our COLM 2024 paper, Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images. The paper studies whether multimodal language models can infer prompts that generate images visually similar to target images produced by text-to-image systems or found in stock image collections, highlighting the relationship between real-world prompts and generated images as well as broader economic and security… See the full description on the dataset page: https://huggingface.co/datasets/AliN96/midjourney-prompts-embeddings.midjourney-promptsMidjourney is an independent research lab whose broad mission is to "explore new mediums of thought". In 2022, they launched a text-to-image service that, given a natural language prompt, produces visual depictions that are faithful to the description. Their service is accessible via a public Discord server: users issue a query in natural language, and the Midjourney bot returns AI-generated images that follow the given description. The raw dataset (with Discord messages) can be found on… See the full description on the dataset page: https://huggingface.co/datasets/succinctly/midjourney-prompts.Midjourney-23Mmidjourney-texttoimage
Dataset Card for Midjourney User Prompts & Generated Images (250k)
Dataset Summary
General Context
Midjourney is an independent research lab whose broad mission is to "explore new mediums of thought". In 2022, they launched a text-to-image service that, given a natural language prompt, produces visual depictions that are faithful to the description. Their service is accessible via a public Discord server, where users interact with a Midjourney bot. When issued… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/midjourney-texttoimage.midjourney-prompts
midjourney-prompts
Description
This dataset contains the cleaned midjourney prompts from Midjourney.
Total prompts: 9,085,397
Version
Count
5.2
2,272,465
5.1
2,060,106
5.0
3,530,770
4.0
1,204,384
3.0
14,991
2.0
791
1.0
1,239
Style
Count
default
8,874,181
raw
177,953
expressive
27,919
scenic
2,146
cute
2,036
original
511
midjourney-texttoimage-new
Dataset Card for Midjourney User Prompts & Generated Images (250k)
Dataset Summary
General Context
Midjourney is an independent research lab whose broad mission is to "explore new mediums of thought". In 2022, they launched a text-to-image service that, given a natural language prompt, produces visual depictions that are faithful to the description. Their service is accessible via a public Discord server, where users interact with a Midjourney bot. When issued… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/midjourney-texttoimage-new.midjourney-prompts-FLUXmidjourney-niji-1Mmidjourney-v5-2M
midjourney-v5-2M dataset in LMDB format
Combine and extract the tar files:
cat mjv5.tar.gz.part-* | tar -xzf -
Files look like this:
mjv5.lmdb/
data.mdb
lock.mdb
midjourney-threads
Dataset Card for Midjourney-Threads 🧵💬
This dataset contains users prompts from the Midjourney discord channel, organized into "threads of interaction".
Each thread contains a user’s trails to create one target image.
The dataset was introduced as part of the paper: Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney.
Dataset Sources
Repository: https://github.com/shachardon/Mid-Journey-to-alignment
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/shachardon/midjourney-threads.midjourney-dalle-sd-nanobananapro-dataset
Dataset Card: Midjourney, DALL-E, Stable Diffusion & Nano Banana Pro vs Real Images
Description
Dataset de classification binaire pour détecter les images générées par IA (Midjourney, DALL-E, Stable Diffusion et Nano Banana Pro) vs images réelles.
Dataset Structure
Train set: 10,000 images
Real: 5000 images
Fake (AI-generated): 5000 images
Test set: 2,000 images
Real: 1000 images
Fake (AI-generated): 1000 images
Features
{
"image": Image… See the full description on the dataset page: https://huggingface.co/datasets/julienlucas/midjourney-dalle-sd-nanobananapro-dataset.felix-midjourney-archive
Felix Midjourney Archive
A deduplicated, checksum-addressed preservation dataset of AI-generated images created by Felix / waffles13 with Midjourney. Images are stored in deterministic WebDataset TAR shards with searchable Parquet, JSONL, CSV, and SQLite catalogs.
Contents
Unique images: 32,477
Exact duplicate source copies excluded: 12,560
Images with full embedded prompts: 13,914
Unique Midjourney Job IDs represented: 22,859
Total image bytes before TAR… See the full description on the dataset page: https://huggingface.co/datasets/wafflefan/felix-midjourney-archive.midjourney-messages-cleaned
midjourney-messages-cleaned
This is vivym/midjourney-messages but with the following cleaning steps:
remove most columns (keep id columns for reference vs. original)
Apply clean-text to all rows (keep casing)
rename content to text (ffs)
remove intermediate ID/tag (???) in angle brackets at the end, remove double asterisks **
remove exact duplicate rows
dataset structure
overall:
DatasetDict({
train: Dataset({
features: ['id', 'channel_id', 'text']… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/midjourney-messages-cleaned.midjourney-detailed-prompts
Midjourney: Detailed Prompts
This dataset is my attempt in providing a high quality text-to-image dataset with detailed and several levels of prompting for images.
Hope it helps anyone in his research ^^
Thanks goes to ...
midjourney-images dataset
Qwen-VL-Max for descriping images in huge detail.
Command R for long & short prompt generation
midjourney-v5-202304-clean
midjourney-v5-202304-clean
简介 Brief Introduction
非官方的,爬取自midjourney v5的2023年4月的数据,一共1701420条。
Unofficial, crawled from midjourney v5 for April 2023, 1,701,420 pairs in total.
数据集信息 Dataset Information
原始项目地址:https://huggingface.co/datasets/tarungupta83/MidJourney_v5_Prompt_dataset
我做了一些清洗,清理出了两个文件:
ori_prompts_df.parquet (1,255,812对,midjourney的四格图)
upscaled_prompts_df.parquet (445,608对,使用了高清指令的图,这意味着这个图更受欢迎。)
Original project address:… See the full description on the dataset page: https://huggingface.co/datasets/wanng/midjourney-v5-202304-clean.MidJourney-generated-imagesmidjourney-dalle-sd-dataset
Dataset Card: Midjourney, DALL-E, Stable Diffusion vs Real Images
Description
Dataset de classification binaire pour détecter les images générées par IA (Midjourney, DALL-E, Stable Diffusion) vs images réelles.
Dataset Structure
Train set: 5,000 images
Real: 2,500 images
Fake (AI-generated): 2,500 images
Test set: 1,000 images
Real: 500 images
Fake (AI-generated): 500 images
Features
{
"image": Image,
"label": "real" | "fake"
}… See the full description on the dataset page: https://huggingface.co/datasets/julienlucas/midjourney-dalle-sd-dataset.Midjourney_v6_Classification_small_shuffledmidjourney-captioned-23m-images-1-0000
midjourney-captioned-23m-images-1-0000
Mirror of the exact bmcore v24 local holdout subset: 4998 image files.
Benchmark label: synthetic. This preserves the source benchmark label; it is not an independent label review.
Source reference: https://huggingface.co/datasets/deepghs/midjourney_captioned_23m_full.
Source-reported terms: other. No additional rights are granted by this mirror.
Source revision reviewed: e82031e89c813a2b0907dbecc1b54cf4025be63e.
Original upstream uses an… See the full description on the dataset page: https://huggingface.co/datasets/34data/midjourney-captioned-23m-images-1-0000.
