datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midjourney-v6-recap
Midjourney v6 Recaptioned
~1.2M Midjourney v6 images with captions from three VLMs:
llava: Original LLaVA captions from the source dataset
gemini: Gemini Flash 1.5 captions
qwen3: Qwen3 VL 8B captions
Caption coverage
llava: available for all 1,235,432 images (from original dataset)
gemini and qwen3: available for 1,017,105 images (82.3%)
Source
Based on brivangl/midjourney-v6-llava.
midjourney-v6-llavaThis dataset based on https://huggingface.co/datasets/CortexLM/midjourney-v6 dataset, captioned with LLava-1.6 model.
This dataset was released as is. By accessing and using this dataset, you acknowledge and agree that Cortex Foundation and the author of this repo are not responsible for any copyright violations or legal consequences that may arise from the use of these images.
midjourney-messages
midjourney-messages
Description
This dataset contains the raw messages from Midjourney.
Total messages: 55,082,563
midjourney-niji-1m-llavanext
Dataset Card for midjourney-niji-1m-llavanext
Dataset Summary
This is a dataset of 2,079,886 synthetic captions for 1,039,943 images from midjourney-v6-520k-raw and nijijourney-v6-520k-raw. The captions were produced using https://huggingface.co/lmms-lab/llama3-llava-next-8b inferenced in float16 after tags were generated with wd-swinv2-tagger-v3, followed by cleanup and shortening with Meta-Llama-3-8B.
All images with metadata are available as MozJPEG encoded JPEGs… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/midjourney-niji-1m-llavanext.AIGeneratedImages_Midjourney
AI Generated Image for Image Classification
This dataset contains AI generated images by Midjourney and Human images taken from Imagenet. The dataset is meant for Image Classification tasks.
Dataset Details
Dataset Description
Curated by: Deepankar Sharma
GenImage_MidJourneymidjourney-images
⛵ Midjourney Images Dataset
This is datase with images made by Midjourney V5/V6.
Dataset parameters
Count of images: ~10.000
Zip file with dataset: True
Captions with images: False
License
License for this dataset: MIT
Use in datasets
pip install -q datasets
from datasets import load_dataset
dataset = load_dataset(
"ehristoforu/midjourney-images",
revision="main"
)
Enjoy with this dataset!
Midjourney-23Mmidjourney-prompts
midjourney-prompts
Description
This dataset contains the cleaned midjourney prompts from Midjourney.
Total prompts: 9,085,397
Version
Count
5.2
2,272,465
5.1
2,060,106
5.0
3,530,770
4.0
1,204,384
3.0
14,991
2.0
791
1.0
1,239
Style
Count
default
8,874,181
raw
177,953
expressive
27,919
scenic
2,146
cute
2,036
original
511
midjourney-niji-1Mmidjourney-threads
Dataset Card for Midjourney-Threads 🧵💬
This dataset contains users prompts from the Midjourney discord channel, organized into "threads of interaction".
Each thread contains a user’s trails to create one target image.
The dataset was introduced as part of the paper: Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney.
Dataset Sources
Repository: https://github.com/shachardon/Mid-Journey-to-alignment
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/shachardon/midjourney-threads.midjourney-dalle-sd-nanobananapro-dataset
Dataset Card: Midjourney, DALL-E, Stable Diffusion & Nano Banana Pro vs Real Images
Description
Dataset de classification binaire pour détecter les images générées par IA (Midjourney, DALL-E, Stable Diffusion et Nano Banana Pro) vs images réelles.
Dataset Structure
Train set: 10,000 images
Real: 5000 images
Fake (AI-generated): 5000 images
Test set: 2,000 images
Real: 1000 images
Fake (AI-generated): 1000 images
Features
{
"image": Image… See the full description on the dataset page: https://huggingface.co/datasets/julienlucas/midjourney-dalle-sd-nanobananapro-dataset.midjourney-detailed-prompts
Midjourney: Detailed Prompts
This dataset is my attempt in providing a high quality text-to-image dataset with detailed and several levels of prompting for images.
Hope it helps anyone in his research ^^
Thanks goes to ...
midjourney-images dataset
Qwen-VL-Max for descriping images in huge detail.
Command R for long & short prompt generation
midjourney-v5-202304-clean
midjourney-v5-202304-clean
简介 Brief Introduction
非官方的,爬取自midjourney v5的2023年4月的数据,一共1701420条。
Unofficial, crawled from midjourney v5 for April 2023, 1,701,420 pairs in total.
数据集信息 Dataset Information
原始项目地址:https://huggingface.co/datasets/tarungupta83/MidJourney_v5_Prompt_dataset
我做了一些清洗,清理出了两个文件:
ori_prompts_df.parquet (1,255,812对,midjourney的四格图)
upscaled_prompts_df.parquet (445,608对,使用了高清指令的图,这意味着这个图更受欢迎。)
Original project address:… See the full description on the dataset page: https://huggingface.co/datasets/wanng/midjourney-v5-202304-clean.MidJourney-generated-imagesMidjourney_v6_Classification_small_shuffledSyn_MidJourneymidjourney-v5-202304
midjourney-v5-202304-clean
简介 Brief Introduction
转载自wanng/midjourney-v5-202304-clean
非官方的,爬取自midjourney v5的2023年4月的数据,一共1701420条。
Unofficial, crawled from midjourney v5 for April 2023, 1,701,420 pairs in total.
数据集信息 Dataset Information
原始项目地址:https://huggingface.co/datasets/tarungupta83/MidJourney_v5_Prompt_dataset
我做了一些清洗,清理出了两个文件:
ori_prompts_df.parquet (1,255,812对,midjourney的四格图)
upscaled_prompts_df.parquet (445,608对,使用了高清指令的图,这意味着这个图更受欢迎。)
Original… See the full description on the dataset page: https://huggingface.co/datasets/JohnTeddy3/midjourney-v5-202304.midjourney-v6
MidJourney v6 Dataset by Bittensor Network (NetUID 19)
Description : This dataset was generated by Subnetwork 19 (Bittensor), utilizing the capabilities of MidJourney v6.
Disclaimer: Image Attribution and Copyright Notice
The images included in this dataset have been sourced from an API. While every effort has been made to ensure compliance with copyright and intellectual property rights, Cortex Foundation cannot guarantee the absence of any copyright or intellectual property… See the full description on the dataset page: https://huggingface.co/datasets/yvdao/midjourney-v6.midjourney-prompts-highquality
Thank you to the Akash Network for sponsoring this project and providing A100s/H100s for compute!
About
A filtered version of the vivym/midjourney-prompts dataset
Filtering criteria
top 10% in length (assuming that longer prompts = more effort and higher quality)
used on an image to be upscaled (assuming that users are more likely to upscale an image that is aesthetically pleasing)
used on midjourney version 5.0+
deduplicated
Run yourself
filter.py script… See the full description on the dataset page: https://huggingface.co/datasets/gaodrew/midjourney-prompts-highquality.midjourney-v6.1
Dataset Card for Midjourney v6.1 Dataset
This dataset is trained on 638 publicly available images sourced from the Midjourney Discord Server. The prompts are cleaned with no parameters etc. Niji images have ", anime style" added to the end of the prompt.
Copyright Notice
This dataset was created using publicly available images and text sourced from the MidJourney Discord server. The images and text included in this dataset are the intellectual property of their respective… See the full description on the dataset page: https://huggingface.co/datasets/saq1b/midjourney-v6.1.Midjourney_gallerymidjourney-vs-real-full-analysismidjourney-kaggle-clean
midjourney-v5-202304-clean
简介 Brief Introduction
非官方的,对Kaggle (Midjourney User Prompts & Generated Images (250k))[https://www.kaggle.com/datasets/succinctlyai/midjourney-texttoimage?select=general-01_2022_06_20.json] 上的数据集进行了清理,一共有 248,167对。
Unofficially, a cleanup of the dataset on Kaggle (Midjourney User Prompts & Generated Images (250k))[https://www.kaggle.com/datasets/succinctlyai/midjourney-texttoimage?select=general-01_2022_06_20.json] yielded 248,167 pairs.… See the full description on the dataset page: https://huggingface.co/datasets/wanng/midjourney-kaggle-clean.midjourney-v6
MidJourney v6 Dataset by Bittensor Network (NetUID 19)
Description : This dataset was generated by Subnetwork 19 (Bittensor), utilizing the capabilities of MidJourney v6.
Disclaimer: Image Attribution and Copyright Notice
The images included in this dataset have been sourced from an API. While every effort has been made to ensure compliance with copyright and intellectual property rights, Cortex Foundation cannot guarantee the absence of any copyright or intellectual property… See the full description on the dataset page: https://huggingface.co/datasets/ruiandcromwell/midjourney-v6.genimage-midjourney-10kMidjourneySrefsFiltered subset of deepghs/midjourney_captioned_23m_full including only the prompts with the sref param passed in midjourney.
midjourney_images_sref_4_images_horizontal
midjourney_v625k images generated by midjourney v6
scrapped from discord
midjourney_captioned_23m_full
Midjourney Captioned Full Dataset
This is the full dataset of Midjourney Captioned 23M dataset. And all the original images are maintained here.
Thanks to the contribution of a certain third-party data provider who wishes to remain anonymous.
Information
Images
There are 23167456 images in total. The maximum ID of these images is 23167456. Last updated at 2024-12-01 12:11:43 UTC.
These are the information of recent 50 images:
id
width
height
filename… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/midjourney_captioned_23m_full.
