datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.midjourney_captioned_23m_full
Midjourney Captioned Full Dataset
This is the full dataset of Midjourney Captioned 23M dataset. And all the original images are maintained here.
Thanks to the contribution of a certain third-party data provider who wishes to remain anonymous.
Information
Images
There are 23167456 images in total. The maximum ID of these images is 23167456. Last updated at 2024-12-01 12:11:43 UTC.
These are the information of recent 50 images:
id
width
height
filename… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/midjourney_captioned_23m_full.e621_2024-captions-1ktar
E621 2024 captions only in 1k tar
Raw captions jointed from lodestones/e621-captions
It doesn't align to any dataset yet.
meta_cap.json has been provided in compressed format if you want to train with kohyas triner. Currently I'm trying to merge this with my 2024 version.
Core logic
The script building this 1ktar
There is not much choice, I don't have GPU to run for 1M captions with VLM so I just "take it or leave it".
rearranged_tags = [row.regular_summary… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/e621_2024-captions-1ktar.
