CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Salesforce /blip3-kale 🥬 BLIP3-KALE:Knowledge Augmented Large-scale Dense Captions BLIP3-KALE is an open-source dataset of 218 million image-text pairs, featuring knowledge-augmented dense captions combining web-scale knowledge with detailed image descriptions. Paper: [To be added] Uses BLIP3-KALE is designed to facilitate research in multimodal pretraining. The dataset can be used for training large multimodal models that require factually grounded, dense image captions. It has already been an… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-kale.imageimage-to-text100M<n<1B46 likes6.6k downloads2y agoHugging Face02BLIP3o /BLIP3o-Pretrain-Long-Caption BLIP3o Pretrain Long-Caption Dataset This collection contains 27 million images, each paired with a long (~120 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Long-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Long-Caption.image10M<n<100M74 likes6k downloads1y agoHugging Face03lambda /pokemon-blip-captionsgated Notice of DMCA Takedown Action We have received a DMCA takedown notice from The Pokémon Company International, Inc. In response to this action, we have taken down the dataset. We appreciate your understanding. imagetext-to-imagen<1K314 likes4.9k downloads3y agoHugging Face04BLIP3o /BLIP3o-Pretrain-Short-Caption BLIP3o Pretrain Short-Caption Dataset This collection contains 5 million images, each paired with a short (~20 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Short-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Short-Caption.image1M<n<10M10 likes4.5k downloads1y agoHugging Face05sasha /prof_images_blip__22h-vintedois-diffusion-v0-1 Dataset Card for "prof_images_blip__22h-vintedois-diffusion-v0-1" More Information needed image10K<n<100K0 likes3.3k downloads3y agoHugging Face06sasha /prof_images_blip__stabilityai-stable-diffusion-2 Dataset Card for "prof_images_blip__stabilityai-stable-diffusion-2" More Information needed image10K<n<100K0 likes2.7k downloads3y agoHugging Face07sasha /prof_images_blip__SG161222-Realistic_Vision_V1.4 Dataset Card for "prof_images_blip__SG161222-Realistic_Vision_V1.4" More Information needed image10K<n<100K0 likes2.3k downloads3y agoHugging Face08BLIP3o /BLIP3o-Pretrain-JourneyDB BLIP3o Pretrain JourneyDB Dataset This collection contains 4 million JourneyDB images. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-JourneyDB", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import load_dataset import glob data_files = glob.glob("/your/data/path/*.tar")… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-JourneyDB.image1M<n<10M7 likes2k downloads1y agoHugging Face09Salesforce /blip3-ocr-200m BLIP3-OCR-200M Dataset Overview The BLIP3-OCR-200M dataset is designed to address the limitations of current Vision-Language Models (VLMs) in processing and interpreting text-rich images, such as documents and charts. Traditional image-text datasets often struggle to capture nuanced textual information, which is crucial for tasks requiring complex text comprehension and reasoning. Key Features OCR Integration: The dataset incorporates Optical Character… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-ocr-200m.image10M<n<100M45 likes2k downloads2y agoHugging Face10Salesforce /BLIP3o-NEXT-EDIT-ENSEMBLE-DATASETS1 likes1.8k downloads11mo agoHugging Face11pufanyi /BLIP3o-60kimage10K<n<100K1 likes1.8k downloads1y agoHugging Face12lambda /naruto-blip-captions Dataset Card for Naruto BLIP captions Dataset used to train TBD. The original images were obtained from narutopedia.com and captioned with the pre-trained BLIP model. For each row the dataset contains image and text keys. image is a varying size PIL jpeg, and text is the accompanying text caption. Only a train split is provided. Example stable diffusion outputs "Bill Gates with a hoodie", "John Oliver with Naruto style", "Hello Kitty with Naruto style", "Lebron… See the full description on the dataset page: https://huggingface.co/datasets/lambda/naruto-blip-captions.56 likes1.7k downloads4y agoHugging Face13sasha /prof_images_blip__andite-anything-v4.0 Dataset Card for "prof_images_blip__andite-anything-v4.0" More Information needed image1K<n<10K0 likes1.3k downloads3y agoHugging Face14diffusion-bench /blip3o-256image1K<n<10K1 likes1.2k downloads6mo agoHugging Face15BLIP3o /BLIP3o-60kThis is BLIP3o-60k Text-to-Image instruction tuning dataset distilled from GPT-4o, including the following categories: JourneyDB Human (including MSCOCO with human caption, human gestures, occupations) Dalle3 Geneval (no overlap with test set) Common objects Simple text Here we provide the code guidance to download tar file: from huggingface_hub import snapshot_download snapshot_download(repo_id='BLIP3o/BLIP3o-60k', repo_type=‘dataset’) And you can use huggingface datasets to read the tar… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-60k.text1K<n<10K40 likes1.1k downloads1y agoHugging Face16MakiPan /hagrid250k-blip2 Dataset Card for "hagrid250k-blip2" More Information needed image100K<n<1M5 likes1.1k downloads3y agoHugging Face17lmms-lab /blip3o-60kimage1K<n<10K0 likes1.1k downloads1y agoHugging Face18sasha /prof_images_blip__andite-pastel-mix Dataset Card for "prof_images_blip__andite-pastel-mix" More Information needed image10K<n<100K0 likes1k downloads3y agoHugging Face19henryscheible /coco_val2014_blip2_processed Dataset Card for "coco_val2014_blip2_processed" More Information needed text10K<n<100K0 likes939 downloads3y agoHugging Face20sasha /prof_images_blip__CompVis-stable-diffusion-v1-4 Dataset Card for "prof_images_blip__CompVis-stable-diffusion-v1-4" More Information needed image10K<n<100K0 likes892 downloads3y agoHugging Face21reach-vb /pokemon-blip-captions Dataset Card for Pokémon BLIP captions Dataset used to train Pokémon text to image model BLIP generated captions for Pokémon images from Few Shot Pokémon dataset introduced by Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image Synthesis (FastGAN). Original images were obtained from FastGAN-pytorch and captioned with the pre-trained BLIP model. For each row the dataset contains image and text keys. image is a varying size PIL jpeg, and text is the… See the full description on the dataset page: https://huggingface.co/datasets/reach-vb/pokemon-blip-captions.imagetext-to-imagen<1K26 likes575 downloads3y agoHugging Face22Ashenone3 /BLIP3o-JourneyDBimage1M<n<10M1 likes402 downloads1y agoHugging Face23Salesforce /blip3-grounding-50m BLIP3-GROUNDING-50M Dataset Overview The BLIP3-GROUNDING-50M dataset is designed to enhance the ability of Vision-Language Models (VLMs) to ground semantic concepts in visual features, which is crucial for tasks like object detection, semantic segmentation, and understanding referring expressions (e.g., "the object to the left of the dog"). Traditional datasets often lack the necessary granularity for such tasks, making it challenging for models to accurately localize and… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-grounding-50m.image10M<n<100M30 likes391 downloads2y agoHugging Face24rbeauchamp /blip_50k_train Dataset Card for "blip_50k_train" More Information needed image10K<n<100K0 likes341 downloads3y agoHugging Face25stable-bias /prof_images_blip__SD_v1.4_random_seeds Dataset Card for "prof_images_blip__SD_v1.4_random_seeds" More Information needed image10K<n<100K0 likes339 downloads3y agoHugging Face26svjack /pokemon-blip-captions-en-zh Dataset Card for Pokémon BLIP captions with English and Chinese. Dataset used to train Pokémon text to image model, add a Chinese Column of Pokémon BLIP captions BLIP generated captions for Pokémon images from Few Shot Pokémon dataset introduced by Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image Synthesis (FastGAN). Original images were obtained from FastGAN-pytorch and captioned with the pre-trained BLIP model. For each row the dataset contains image… See the full description on the dataset page: https://huggingface.co/datasets/svjack/pokemon-blip-captions-en-zh.imagetext-to-imagen<1K54 likes330 downloads4y agoHugging Face27hahminlew /kream-product-blip-captions KREAM Product Blip Captions Dataset Information KREAM Product Blip Captions Dataset is a dataset card for finetuning a text-to-image generative model collected from KREAM, one of the best online-resell market in Korea. This dataset consists of 'image' and 'text' key pairs. The format of 'text' is 'category (e.g. outer), product original name (e.g. The North Face 1996 Eco Nuptse Jacket Black), blip captions (e.g. a photography of the north face black down jacket)'. You can easily… See the full description on the dataset page: https://huggingface.co/datasets/hahminlew/kream-product-blip-captions.imagetext-to-image10K<n<100K10 likes328 downloads3y agoHugging Face28sasha /prof_images_blip__wavymulder-Analog-Diffusion Dataset Card for "prof_images_blip__wavymulder-Analog-Diffusion" More Information needed image10K<n<100K0 likes326 downloads3y agoHugging Face29stable-bias /prof_images_blip__dalle-2 Dataset Card for "prof_images_blip__dalle-2" More Information needed image10K<n<100K0 likes295 downloads3y agoHugging Face30thaottn /DataComp_large_pool_BLIP2_captions Dataset Card for DataComp_large_pool_BLIP2_captions Dataset Summary Supported Tasks and Leaderboards We have used this dataset for pre-training CLIP models and found that it rivals or outperforms models trained on raw web captions on average across the 38 evaluation tasks proposed by DataComp. Refer to the DataComp leaderboard (https://www.datacomp.ai/leaderboard.html) for the top baselines uncovered in our work. Languages Primarily English.… See the full description on the dataset page: https://huggingface.co/datasets/thaottn/DataComp_large_pool_BLIP2_captions.textimage-to-text10M<n<100M1 likes278 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.