datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blip3-kale
🥬 BLIP3-KALE:Knowledge Augmented Large-scale Dense Captions
BLIP3-KALE is an open-source dataset of 218 million image-text pairs, featuring knowledge-augmented dense captions combining web-scale knowledge with detailed image descriptions.
Paper: [To be added]
Uses
BLIP3-KALE is designed to facilitate research in multimodal pretraining. The dataset can be used for training large multimodal models that require factually grounded, dense image captions. It has already been an… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-kale.comicstrips-gpt4o-blip3
Comic Strips
Dataset Details
Dataset Description
This dataset contains indie comics from Reddit, then captioned with GPT4o and BLIP3.
Currently, only the GPT4o captions are available in this repository. The BLIP3 captions will be uploaded soon.
Roughly 1400 images were captioned at a cost of ~$11 using GPT4o (25 May 2024 version).
Curated by: @pseudoterminalx
Funded by @pseudoterminalx
License: MIT
Dataset Sources
Unlike other free-to-use… See the full description on the dataset page: https://huggingface.co/datasets/bghira/comicstrips-gpt4o-blip3.flickr30k-blip2max-densemscoco-blip-densemscoco-blip2max-densemscoco-blip2avg-denseBLIP_VQA_Vietnameseblip-kd-results-fitnets-fixed-fashion200ktextvqa_mini_validation_Salesforce_blip2-flan-t5-xxl_ns_100textvqa_valid_Salesforce_blip2-flan-t5-xxl_ns_5000flickr30k-blip2avg-denseblip-kd-student-attn-distill-fashion200k-15kwikiart_kaggle_blip_captionsVibration_datasetwikiart-blip-captions
