datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midjourney-v6-recap
Midjourney v6 Recaptioned
~1.2M Midjourney v6 images with captions from three VLMs:
llava: Original LLaVA captions from the source dataset
gemini: Gemini Flash 1.5 captions
qwen3: Qwen3 VL 8B captions
Caption coverage
llava: available for all 1,235,432 images (from original dataset)
gemini and qwen3: available for 1,017,105 images (82.3%)
Source
Based on brivangl/midjourney-v6-llava.
midjourney-v6-llavaThis dataset based on https://huggingface.co/datasets/CortexLM/midjourney-v6 dataset, captioned with LLava-1.6 model.
This dataset was released as is. By accessing and using this dataset, you acknowledge and agree that Cortex Foundation and the author of this repo are not responsible for any copyright violations or legal consequences that may arise from the use of these images.
midjourney-messages
midjourney-messages
Description
This dataset contains the raw messages from Midjourney.
Total messages: 55,082,563
midjourney-prompts-embeddings
Midjourney Prompt–Embedding Dataset
This dataset is derived from our COLM 2024 paper, Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images. The paper studies whether multimodal language models can infer prompts that generate images visually similar to target images produced by text-to-image systems or found in stock image collections, highlighting the relationship between real-world prompts and generated images as well as broader economic and security… See the full description on the dataset page: https://huggingface.co/datasets/AliN96/midjourney-prompts-embeddings.midjourney-promptsMidjourney is an independent research lab whose broad mission is to "explore new mediums of thought". In 2022, they launched a text-to-image service that, given a natural language prompt, produces visual depictions that are faithful to the description. Their service is accessible via a public Discord server: users issue a query in natural language, and the Midjourney bot returns AI-generated images that follow the given description. The raw dataset (with Discord messages) can be found on… See the full description on the dataset page: https://huggingface.co/datasets/succinctly/midjourney-prompts.Midjourney-23Mmidjourney-prompts
midjourney-prompts
Description
This dataset contains the cleaned midjourney prompts from Midjourney.
Total prompts: 9,085,397
Version
Count
5.2
2,272,465
5.1
2,060,106
5.0
3,530,770
4.0
1,204,384
3.0
14,991
2.0
791
1.0
1,239
Style
Count
default
8,874,181
raw
177,953
expressive
27,919
scenic
2,146
cute
2,036
original
511
midjourney-prompts-FLUXmidjourney-niji-1Mmidjourney-threads
Dataset Card for Midjourney-Threads 🧵💬
This dataset contains users prompts from the Midjourney discord channel, organized into "threads of interaction".
Each thread contains a user’s trails to create one target image.
The dataset was introduced as part of the paper: Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney.
Dataset Sources
Repository: https://github.com/shachardon/Mid-Journey-to-alignment
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/shachardon/midjourney-threads.midjourney-messages-cleaned
midjourney-messages-cleaned
This is vivym/midjourney-messages but with the following cleaning steps:
remove most columns (keep id columns for reference vs. original)
Apply clean-text to all rows (keep casing)
rename content to text (ffs)
remove intermediate ID/tag (???) in angle brackets at the end, remove double asterisks **
remove exact duplicate rows
dataset structure
overall:
DatasetDict({
train: Dataset({
features: ['id', 'channel_id', 'text']… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/midjourney-messages-cleaned.midjourney-detailed-prompts
Midjourney: Detailed Prompts
This dataset is my attempt in providing a high quality text-to-image dataset with detailed and several levels of prompting for images.
Hope it helps anyone in his research ^^
Thanks goes to ...
midjourney-images dataset
Qwen-VL-Max for descriping images in huge detail.
Command R for long & short prompt generation
midjourney-v5-202304-clean
midjourney-v5-202304-clean
简介 Brief Introduction
非官方的,爬取自midjourney v5的2023年4月的数据,一共1701420条。
Unofficial, crawled from midjourney v5 for April 2023, 1,701,420 pairs in total.
数据集信息 Dataset Information
原始项目地址:https://huggingface.co/datasets/tarungupta83/MidJourney_v5_Prompt_dataset
我做了一些清洗,清理出了两个文件:
ori_prompts_df.parquet (1,255,812对,midjourney的四格图)
upscaled_prompts_df.parquet (445,608对,使用了高清指令的图,这意味着这个图更受欢迎。)
Original project address:… See the full description on the dataset page: https://huggingface.co/datasets/wanng/midjourney-v5-202304-clean.midjourney-captioned-23m-images-1-0000
midjourney-captioned-23m-images-1-0000
Mirror of the exact bmcore v24 local holdout subset: 4998 image files.
Benchmark label: synthetic. This preserves the source benchmark label; it is not an independent label review.
Source reference: https://huggingface.co/datasets/deepghs/midjourney_captioned_23m_full.
Source-reported terms: other. No additional rights are granted by this mirror.
Source revision reviewed: e82031e89c813a2b0907dbecc1b54cf4025be63e.
Original upstream uses an… See the full description on the dataset page: https://huggingface.co/datasets/34data/midjourney-captioned-23m-images-1-0000.midjourney-prompts-only
Dataset Card for "midjourney-prompts-only"
More Information needed
MidJourney_v5_Prompt_datasetDataset contain raw prompts from Mid Journey v5
Total Records : 4245117
Sample Data
AuthorID
Author
Date
Content
Attachments
Reactions
936929561302675456
Midjourney Bot#9282
04/20/2023 12:00 AM
benjamin frankling with rayban sunglasses reflecting a usa flag walking on a side of penguin, whit...
Link
936929561302675456
Midjourney Bot#9282
04/20/2023 12:00 AM
Street vendor robot in 80's Poland, meat market, fruit stall, communist style, real photo, real ph...
Link… See the full description on the dataset page: https://huggingface.co/datasets/tarungupta83/MidJourney_v5_Prompt_dataset.midjourney-v5-202304
midjourney-v5-202304-clean
简介 Brief Introduction
转载自wanng/midjourney-v5-202304-clean
非官方的,爬取自midjourney v5的2023年4月的数据,一共1701420条。
Unofficial, crawled from midjourney v5 for April 2023, 1,701,420 pairs in total.
数据集信息 Dataset Information
原始项目地址:https://huggingface.co/datasets/tarungupta83/MidJourney_v5_Prompt_dataset
我做了一些清洗,清理出了两个文件:
ori_prompts_df.parquet (1,255,812对,midjourney的四格图)
upscaled_prompts_df.parquet (445,608对,使用了高清指令的图,这意味着这个图更受欢迎。)
Original… See the full description on the dataset page: https://huggingface.co/datasets/JohnTeddy3/midjourney-v5-202304.midjourney-captioned-23m-images-1-0001
midjourney-captioned-23m-images-1-0001
Mirror of the exact bmcore v24 local holdout subset: 4998 image files.
Benchmark label: synthetic. This preserves the source benchmark label; it is not an independent label review.
Source reference: https://huggingface.co/datasets/deepghs/midjourney_captioned_23m_full.
Source-reported terms: other. No additional rights are granted by this mirror.
Source revision reviewed: e82031e89c813a2b0907dbecc1b54cf4025be63e.
Original upstream uses an… See the full description on the dataset page: https://huggingface.co/datasets/34data/midjourney-captioned-23m-images-1-0001.midjourney-v6
MidJourney v6 Dataset by Bittensor Network (NetUID 19)
Description : This dataset was generated by Subnetwork 19 (Bittensor), utilizing the capabilities of MidJourney v6.
Disclaimer: Image Attribution and Copyright Notice
The images included in this dataset have been sourced from an API. While every effort has been made to ensure compliance with copyright and intellectual property rights, Cortex Foundation cannot guarantee the absence of any copyright or intellectual property… See the full description on the dataset page: https://huggingface.co/datasets/yvdao/midjourney-v6.midjourney-prompts-highquality
Thank you to the Akash Network for sponsoring this project and providing A100s/H100s for compute!
About
A filtered version of the vivym/midjourney-prompts dataset
Filtering criteria
top 10% in length (assuming that longer prompts = more effort and higher quality)
used on an image to be upscaled (assuming that users are more likely to upscale an image that is aesthetically pleasing)
used on midjourney version 5.0+
deduplicated
Run yourself
filter.py script… See the full description on the dataset page: https://huggingface.co/datasets/gaodrew/midjourney-prompts-highquality.midjourney-prompts
Midjourney Prompts
Extracted text prompts from Midjourney Discord messages.
Splits
train: 219,209 prompts (90%)
validation: 12,178 prompts (5%)
test: 12,179 prompts (5%)
Usage
from datasets import load_dataset
ds = load_dataset("text", data_files={
"train": "https://huggingface.co/datasets/suriyagunasekar/midjourney-prompts/resolve/main/train.txt",
"validation":… See the full description on the dataset page: https://huggingface.co/datasets/suriyagunasekar/midjourney-prompts.Midjourneyv5-5Kmidjourney-v6.1
Dataset Card for Midjourney v6.1 Dataset
This dataset is trained on 638 publicly available images sourced from the Midjourney Discord Server. The prompts are cleaned with no parameters etc. Niji images have ", anime style" added to the end of the prompt.
Copyright Notice
This dataset was created using publicly available images and text sourced from the MidJourney Discord server. The images and text included in this dataset are the intellectual property of their respective… See the full description on the dataset page: https://huggingface.co/datasets/saq1b/midjourney-v6.1.midjourney-prompts-dataset-3
Midjourney Text-to-Image Prompts Dataset
This dataset contains extracted and cleaned Midjourney text prompts for training language models.
Dataset Description
This dataset is extracted from Discord messages containing Midjourney bot interactions. It contains text prompts that were used to generate images using Midjourney.
Data Source
Source: Midjourney Discord server messages (Discord JSON export)
Message Types: INITIAL_OR_VARIATION and UPSCALE messages only… See the full description on the dataset page: https://huggingface.co/datasets/suriyagunasekar/midjourney-prompts-dataset-3.midjourney-prompts
Midjourney Prompts
Extracted text prompts from Midjourney Discord messages.
Splits
train: 219,209 prompts (90%)
validation: 12,178 prompts (5%)
test: 12,179 prompts (5%)
Usage
from datasets import load_dataset
ds = load_dataset("text", data_files={
"train": "https://huggingface.co/datasets/suriyagunasekar/midjourney-prompts/resolve/main/train.txt",
"validation":… See the full description on the dataset page: https://huggingface.co/datasets/bwdyaks778/midjourney-prompts.midjourney-leaks
Midjourney prompts leaks
About
This dataset contains 5000 raw Midjourney prompts leaked from their Discord.
How to analyze the dataset with phospho?
phospho is a platform to do text analytics, even with raw, uncleaned data. Here's how to do it:
Create an account @https://phospho.ai.
Load the CSV file.
In Clusters, go to Configure clusters detection. Change the instruction to type of image generated. Select 10 clusters.
Run the clustering and enjoy the… See the full description on the dataset page: https://huggingface.co/datasets/phospho-ai/midjourney-leaks.midjourney-vs-real-full-analysismidjourney-kaggle-clean
midjourney-v5-202304-clean
简介 Brief Introduction
非官方的,对Kaggle (Midjourney User Prompts & Generated Images (250k))[https://www.kaggle.com/datasets/succinctlyai/midjourney-texttoimage?select=general-01_2022_06_20.json] 上的数据集进行了清理,一共有 248,167对。
Unofficially, a cleanup of the dataset on Kaggle (Midjourney User Prompts & Generated Images (250k))[https://www.kaggle.com/datasets/succinctlyai/midjourney-texttoimage?select=general-01_2022_06_20.json] yielded 248,167 pairs.… See the full description on the dataset page: https://huggingface.co/datasets/wanng/midjourney-kaggle-clean.midjourney-v6
MidJourney v6 Dataset by Bittensor Network (NetUID 19)
Description : This dataset was generated by Subnetwork 19 (Bittensor), utilizing the capabilities of MidJourney v6.
Disclaimer: Image Attribution and Copyright Notice
The images included in this dataset have been sourced from an API. While every effort has been made to ensure compliance with copyright and intellectual property rights, Cortex Foundation cannot guarantee the absence of any copyright or intellectual property… See the full description on the dataset page: https://huggingface.co/datasets/ruiandcromwell/midjourney-v6.genimage-midjourney-10k
