datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conceptual-captions-12m-webdatasetmajestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.wds_mscoco_captionswds_mscoco_captions2017freesound-commercially-permissive-subset-with-captionsobjaverse_processed_renders_and_captionsContains rendered views and captions from Objaverse XL objects. the objects are from the alignment and TRELLIS500K (over 1 Millionen processed objects) dataset. We downloaded and rendered 4 views of each object. We added TRELLIS and CAP3D Captions where available. If there were no captions we generated new captions with the large version of Florence 2. This is the base dataset we used to generate MeshFleet which is described in MeshFleet: Filtered and Annotated 3D Vehicle Dataset for Domain… See the full description on the dataset page: https://huggingface.co/datasets/DamianBoborzi/objaverse_processed_renders_and_captions.audioset-with-grounded-captionsSAM-LLaVA-Captions10Mwikiart_with_BLIP_captionsmajestrino-unified-detailed-captions-temporal
Majestrino Unified Detailed Captions with Temporal Aspects
Filtered subset of laion/majestrino-data containing only samples with unified_detailed_caption_with_temporal_aspects.
Stats
4,128,665 samples
826 tar files (~1.1 GB each)
~878 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption with temporal aspects
caption_type — always unified_detailed_caption_with_temporal_aspects
transcription — speech… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions-temporal.audioset-with-captions2025-captionssynthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.SRF_captionse621_2024-captions-1ktar
E621 2024 captions only in 1k tar
Raw captions jointed from lodestones/e621-captions
It doesn't align to any dataset yet.
meta_cap.json has been provided in compressed format if you want to train with kohyas triner. Currently I'm trying to merge this with my 2024 version.
Core logic
The script building this 1ktar
There is not much choice, I don't have GPU to run for 1M captions with VLM so I just "take it or leave it".
rearranged_tags = [row.regular_summary… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/e621_2024-captions-1ktar.pokemon-blip-captions-wdsWebdataset version of: lambdalabs/pokemon-blip-captions
WikiArt-81K-BLIP_2-captions
WikiArt Enhanced Dataset
Description
This dataset contains 81,444 artistic images from WikiArt, organized into different artistic genres. It has undergone several improvements and corrections to optimize its use in machine learning tasks and computational art analysis. Credits to the original author of daset go to: WikiArt
Enhancements
1. Encoding Issues Correction
Fixed encoding issues in filenames and artist information.
All filenames were renamed… See the full description on the dataset page: https://huggingface.co/datasets/Dant33/WikiArt-81K-BLIP_2-captions.gpt4o_captions_1k5samples
PACO
WebDataset export for PACO-style localized caption data.
Summary
Samples: 1500
Shards: 2
Payload format inside each shard: pickle
Split: train
Config: PACO
Layout
Media files are stored in WebDataset tar shards.
Each sample key is stable and becomes __key__ in the dataset viewer.
Hugging Face will infer columns such as jpg, pickle, json, __key__, and __url__ from the shard contents.
Manifest file: PACO/annotations.json
