datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conceptual-captions-12m-webdatasetBLIP3o-Pretrain-Long-Caption
BLIP3o Pretrain Long-Caption Dataset
This collection contains 27 million images, each paired with a long (~120 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct.
Download
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="BLIP3o/BLIP3o-Pretrain-Long-Caption",
repo_type="dataset"
)
Load Dataset without Extracting
You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Long-Caption.BLIP3o-Pretrain-Short-Caption
BLIP3o Pretrain Short-Caption Dataset
This collection contains 5 million images, each paired with a short (~20 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct.
Download
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="BLIP3o/BLIP3o-Pretrain-Short-Caption",
repo_type="dataset"
)
Load Dataset without Extracting
You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Short-Caption.InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption
wds_mscoco_captionsMECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.wds_mscoco_captions2017CaptionedSynthTextThis dataset has been created by Stability AI and LAION.
SynthText is a popular OCR dataset, where random texts are rendered into random locations in images based on depth maps.
In this dataset, we additionally computed image captions using BLIP2.
Caption: "a close up of a leopard's face with a blurry background"
objaverse_processed_renders_and_captionsContains rendered views and captions from Objaverse XL objects. the objects are from the alignment and TRELLIS500K (over 1 Millionen processed objects) dataset. We downloaded and rendered 4 views of each object. We added TRELLIS and CAP3D Captions where available. If there were no captions we generated new captions with the large version of Florence 2. This is the base dataset we used to generate MeshFleet which is described in MeshFleet: Filtered and Annotated 3D Vehicle Dataset for Domain… See the full description on the dataset page: https://huggingface.co/datasets/DamianBoborzi/objaverse_processed_renders_and_captions.ffhq_captioned_1024
ffhq_captioned_1024
A captioned bucketed-shards export of gaunernst/ffhq-1024-wds.
This export contains 70,000 square face and portrait images from FFHQ, stored as
JPEG TAR shards in a single 1024 x 1024 bucket. The source images are decoded
from the original dataset, deterministically converted to RGB, and re-encoded as
high-quality JPEG (quality=95, adaptive subsampling). Captions were generated
with a Gemini 2.5 Flash Lite primary pass and a Mistral Medium 3.1 fallback.
Intended… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/ffhq_captioned_1024.SynRIS-captionednsfw-video-still-caption-grid-onlyInternVL-SA-1B-Caption-512synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.midjourney_captioned_23m_full
Midjourney Captioned Full Dataset
This is the full dataset of Midjourney Captioned 23M dataset. And all the original images are maintained here.
Thanks to the contribution of a certain third-party data provider who wishes to remain anonymous.
Information
Images
There are 23167456 images in total. The maximum ID of these images is 23167456. Last updated at 2024-12-01 12:11:43 UTC.
These are the information of recent 50 images:
id
width
height
filename… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/midjourney_captioned_23m_full.NWPU-Captionpokemon-blip-captions-wdsWebdataset version of: lambdalabs/pokemon-blip-captions
nsfw-video-still-caption-testWikiArt-81K-BLIP_2-captions
WikiArt Enhanced Dataset
Description
This dataset contains 81,444 artistic images from WikiArt, organized into different artistic genres. It has undergone several improvements and corrections to optimize its use in machine learning tasks and computational art analysis. Credits to the original author of daset go to: WikiArt
Enhancements
1. Encoding Issues Correction
Fixed encoding issues in filenames and artist information.
All filenames were renamed… See the full description on the dataset page: https://huggingface.co/datasets/Dant33/WikiArt-81K-BLIP_2-captions.aesthetic-cleaned-captionedsynthetic-image-caption-pairssoundnet-flux-captiongpt4o_captions_1k5samples
PACO
WebDataset export for PACO-style localized caption data.
Summary
Samples: 1500
Shards: 2
Payload format inside each shard: pickle
Split: train
Config: PACO
Layout
Media files are stored in WebDataset tar shards.
Each sample key is stable and becomes __key__ in the dataset viewer.
Hugging Face will infer columns such as jpg, pickle, json, __key__, and __url__ from the shard contents.
Manifest file: PACO/annotations.json
SD-2.1_with_coco_captionimagenet_val_caption1CaptionedSynthTextThis dataset has been created by Stability AI and LAION.
SynthText is a popular OCR dataset, where random texts are rendered into random locations in images based on depth maps.
In this dataset, we additionally computed image captions using BLIP2.
Caption: "a close up of a leopard's face with a blurry background"
medonethinker-source-captioning
