Tran1312/Image-Caption
BLIP3o Long-Caption 100K Image-Text Subset This dataset is a locally reorganized subset of BLIP3o/BLIP3o-Pretrain-Long-Caption, containing approximately 100,000 image-text pairs selected from the original BLIP3o long-caption pretraining dataset [1]. The original BLIP3o long-caption collection contains approximately 27 million images, each paired with a long caption of roughly 120 tokens generated using Qwen2.5-VL-7B-Instruct [1]. The BLIP3-o project was introduced as part of a… See the full description on the dataset page: https://huggingface.co/datasets/Tran1312/Image-Caption.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face