datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midjourney-threads
Dataset Card for Midjourney-Threads 🧵💬
This dataset contains users prompts from the Midjourney discord channel, organized into "threads of interaction".
Each thread contains a user’s trails to create one target image.
The dataset was introduced as part of the paper: Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney.
Dataset Sources
Repository: https://github.com/shachardon/Mid-Journey-to-alignment
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/shachardon/midjourney-threads.MidJourney_v5_Prompt_datasetDataset contain raw prompts from Mid Journey v5
Total Records : 4245117
Sample Data
AuthorID
Author
Date
Content
Attachments
Reactions
936929561302675456
Midjourney Bot#9282
04/20/2023 12:00 AM
benjamin frankling with rayban sunglasses reflecting a usa flag walking on a side of penguin, whit...
Link
936929561302675456
Midjourney Bot#9282
04/20/2023 12:00 AM
Street vendor robot in 80's Poland, meat market, fruit stall, communist style, real photo, real ph...
Link… See the full description on the dataset page: https://huggingface.co/datasets/tarungupta83/MidJourney_v5_Prompt_dataset.midjourney-prompts-highquality
Thank you to the Akash Network for sponsoring this project and providing A100s/H100s for compute!
About
A filtered version of the vivym/midjourney-prompts dataset
Filtering criteria
top 10% in length (assuming that longer prompts = more effort and higher quality)
used on an image to be upscaled (assuming that users are more likely to upscale an image that is aesthetically pleasing)
used on midjourney version 5.0+
deduplicated
Run yourself
filter.py script… See the full description on the dataset page: https://huggingface.co/datasets/gaodrew/midjourney-prompts-highquality.midjourney-leaks
Midjourney prompts leaks
About
This dataset contains 5000 raw Midjourney prompts leaked from their Discord.
How to analyze the dataset with phospho?
phospho is a platform to do text analytics, even with raw, uncleaned data. Here's how to do it:
Create an account @https://phospho.ai.
Load the CSV file.
In Clusters, go to Configure clusters detection. Change the instruction to type of image generated. Select 10 clusters.
Run the clustering and enjoy the… See the full description on the dataset page: https://huggingface.co/datasets/phospho-ai/midjourney-leaks.mid_journey_promptsmidjourney-prompt-enhancementmidjourney_prompty_datasetmidjourney-sentiment130k midjourney prompts and their evaluated sentiment using NLTK library and the "Opinion Mining" positive/negative words library.
https://www.cs.uic.edu/~liub/FBS/sentiment-analysis.html
midjourney_smallmidjourney_harsh_10k
