datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laion2b-en-a65_cogvlm2-4bit_captions
Abstract
This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8).
The synthetic images are best viewed locally by cloning this repo with:
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.Minecraft-Skins-Captioned-1M
Dataset Card for Minecraft Skins
Dataset Summary
This dataset contains 981,079 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier.
Dataset Structure
Data Fields
This dataset includes the following fields:
hash: A data dependent hash. These hashes are generated from raw bytes and will be same if the skin is identical.
image: The skin image encoded in base64 format.… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/Minecraft-Skins-Captioned-1M.IntraOral_Gingivitis_Image_Captioning
A DENTAL INTRAORAL IMAGE DATASET OF GINGIVITIS FOR IMAGE CAPTIONING
Dataset Description
This dataset is a copy of A Dental IntraOral Image Dataset of Gingivitis for Image Captioning which is shared with the license CC BY 4.0.
This dataset contains 1,096 samples organized across multiple splits.
The dataset includes image data.
Splits
train: 732 samples
test: 182 samples
validation: 182 samples
Dataset Creation
This dataset was created using… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/IntraOral_Gingivitis_Image_Captioning.french-lot-department-captioned-photos
Lot Department, France Image Dataset
A collection of high-resolution scenic photographs from the Lot region of France with AI-generated descriptive captions.
Dataset Summary
This dataset contains scenic photographs from three notable locations in France's Lot department: Rocamadour, Autoire, and Padirac. All images were captured using a Sony A6600 camera and are paired with detailed English captions generated by Mistral AI's Pixtral-Large model.
Key Features:… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/french-lot-department-captioned-photos.Textual-Image-Caption-Dataset
Update: OCT-2023
Add v2 with recent SoTA model swinV2 classifier for both soft/hard-label visual_caption_cosine_score_v2 with person label (0.2, 0.3 and 0.4)
Introduction
Modern image captaining relies heavily on extracting knowledge, from images such as objects,
to capture the concept of static story in the image. In this paper, we propose a textual visual context dataset
for captioning, where the publicly available dataset COCO caption (Lin et al., 2014) has been… See the full description on the dataset page: https://huggingface.co/datasets/AhmedSSabir/Textual-Image-Caption-Dataset.TreeOfLife-10M-Captions
Dataset Card for TreeOfLife-10M Captions
This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model.
Dataset Details
This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M-Captions.albi-captioned-photos
Albi, France Image Dataset
A collection of high-resolution scenic photographs from Albi, France with AI-generated descriptive captions.
Dataset Summary
This dataset contains scenic photographs from Albi, France, including the city center, the Toulouse Lautrec museum, and the Sainte-Cécile Cathedral. All images were captured using a Sony A6600 camera and are paired with detailed English captions generated by Mistral AI's Pixtral-Large model.
Key Features:
High-resolution… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/albi-captioned-photos.danbooru2024-captions-1ktar
Danbooru 2024 captions only in 1k tar
Raw captions jointed by 7.62M unpublished extended dataset from KBlueLeaf/danbooru2023-metadata-database and 0.48M generated dataset via Minthy/ToriiGate-v0.4-7B in exl2-8bpw mode. There are 8.13M in total.
python convert_meta_to_tar.py
Reading source JSON
Keys count: 8136011
max id: 8360499
100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1000/1000… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/danbooru2024-captions-1ktar.vintage-photography-450k-high-quality-captionsThis is a 450k image datastet focused on photography from the 20th century, and their analog aspect. Many of the images are in high resolution. This dataset currently has 20k images captioned with InternVL2 26B, and is a work in progress (I plan to caption the entire dataset and also have short captions for all of the images, compute is an issue for now).
vintage-artworks-60k-captionedThis is a dataset consisting of 60k vintage artworks from the 20th century, consisting of vintage pulp, sci-fi and pinup artworks from that era.
The dataset has short and long captions for each image, as well as resolution information. The large captions (large_caption column) were made with florence-2-large-ft, and then shortened with llama 3 8b (see short_caption column).
PokeFA-pokemon-fanart-captioned
PokeFA — Pokémon fan-art metadata with relevance/aesthetic scores and hybrid captions
PokeFA is a large-scale Pokémon fan-art dataset released as metadata + URLs only (no image bytes).~30,000 candidate images are collected across 1,025 Pokémon using a popularity-banded budget with following curation pipeline:
NSFW filtering → OCR localization & inpainting → resizing → relevance & aesthetic scoring (GPT-5-mini vision) → near-duplicate removal → quality filtering to the top ~16… See the full description on the dataset page: https://huggingface.co/datasets/Kev0208/PokeFA-pokemon-fanart-captioned.synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.midjourney_captioned_23m_full
Midjourney Captioned Full Dataset
This is the full dataset of Midjourney Captioned 23M dataset. And all the original images are maintained here.
Thanks to the contribution of a certain third-party data provider who wishes to remain anonymous.
Information
Images
There are 23167456 images in total. The maximum ID of these images is 23167456. Last updated at 2024-12-01 12:11:43 UTC.
These are the information of recent 50 images:
id
width
height
filename… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/midjourney_captioned_23m_full.e621_2024-captions-1ktar
E621 2024 captions only in 1k tar
Raw captions jointed from lodestones/e621-captions
It doesn't align to any dataset yet.
meta_cap.json has been provided in compressed format if you want to train with kohyas triner. Currently I'm trying to merge this with my 2024 version.
Core logic
The script building this 1ktar
There is not much choice, I don't have GPU to run for 1M captions with VLM so I just "take it or leave it".
rearranged_tags = [row.regular_summary… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/e621_2024-captions-1ktar.deviantart_gt2likes_captioned_8m_full
DeviantArt Like>=2k Captioned Full Dataset
This is the full dataset of DeviantArt dataset, which like count no less than 2k. And all the original images are maintained here.
Thanks to the contribution of a certain third-party data provider who wishes to remain anonymous.
Information
Images
There are 7930631 images in total. The maximum ID of these images is 7932093. Last updated at 2024-11-03 14:16:47 UTC.
These are the information of recent 50 images:
id… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/deviantart_gt2likes_captioned_8m_full.vintage-photography-450k-high-quality-captionsThis is a 450k image datastet focused on photography from the 20th century, and their analog aspect. Many of the images are in high resolution. This dataset currently has 20k images captioned with InternVL2 26B, and is a work in progress (I plan to caption the entire dataset and also have short captions for all of the images, compute is an issue for now).
danbooru2023-captions-1ktar
Danbooru 2023 captions only in 1k tar
Raw captions jointed by unpublished extended dataset from KBlueLeaf/danbooru2023-metadata-database
It aligns to nyanko7/danbooru2023. There are around 200k missing for the 2024 version, I'll try to use Minthy/ToriiGate-v0.4-7B to fill in the rest.
meta_cap.json has been provided in compressed format if you want to train with kohyas triner. Currently I'm trying to merge this with my 2024 version.
Core logic
The script building… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/danbooru2023-captions-1ktar.
