datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blip3o-caption-mini-arrow
blip3o-caption-mini-arrow
blip3o-caption-mini-arrow is a high-quality, curated image-caption dataset derived and optimized from the original BLIP3o/BLIP3o-Pretrain-Long-Caption. This dataset is specifically filtered and processed for tasks involving long-form image captioning and vision-language understanding.
Overview
Total Samples: 91,600
Modality: Image ↔ Text
Format: Arrow (auto-converted to Parquet)
License: Apache 2.0
Language: English
Size: ~4.5 GB… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/blip3o-caption-mini-arrow.BLIP3o-60k_wan_i2v_14b_480p_512_512_81_rewrite
