CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TommyBsk /Embodied-Captioning Embodied Image Captioning – Manually Annotated Test Set Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning 📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.tabularimage-to-text1K<n<10K0 likes8.8k downloads1y agoHugging Face02ibm-esa-geospatial /Llama3-SSL4EO-S12-v1.1-captions Llama3-SSL4EO-S12-Captions The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model. Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper. Code: https://github.com/IBM/MS-CLIP Data Structure We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.tabularzero-shot-image-classification100K<n<1M5 likes1.3k downloads1y agoHugging Face03flax-community /conceptual-captions-12This file contains English captions from Conceptual 12M dataset by Google. Since we don't own the images, we have provided the link to images, name of downloaded file, and caption for that image in the TSV file. We would like to thank Luke Melas for helping us get the cleaned CC-12M data on our TPU-VMs. image10M<n<100M5 likes368 downloads3y agoHugging Face04AbstractPhil /human-templated-captions-1bcsv delimiter is = ".,|,." apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading. This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon. Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.texttext-generation100M<n<1B1 likes328 downloads1y agoHugging Face05thaottn /DataComp_large_pool_BLIP2_captions Dataset Card for DataComp_large_pool_BLIP2_captions Dataset Summary Supported Tasks and Leaderboards We have used this dataset for pre-training CLIP models and found that it rivals or outperforms models trained on raw web captions on average across the 38 evaluation tasks proposed by DataComp. Refer to the DataComp leaderboard (https://www.datacomp.ai/leaderboard.html) for the top baselines uncovered in our work. Languages Primarily English.… See the full description on the dataset page: https://huggingface.co/datasets/thaottn/DataComp_large_pool_BLIP2_captions.textimage-to-text10M<n<100M1 likes265 downloads3y agoHugging Face06qidouxiong619 /dreamlip_long_captions Dataset Card for DreamLIP-30M Dataset Summary DreamLIP-Long-Captions is a dataset consisting of ~30M image annotations, i.e. detailed long captions. In contrast with the curated style of other synthetic image caption annotations, DreamLIP-30M utilizes pre-trained Multi-modality Large Language Model to obtain detailed descriptions with an average length of 247. More precisely, the detailed descriptions are generated by asking the ShareGPT4V/InstructBLIP/LLava1.5 the… See the full description on the dataset page: https://huggingface.co/datasets/qidouxiong619/dreamlip_long_captions.imagetext-to-image10M<n<100M19 likes105 downloads2y agoHugging Face07thaottn /DataComp_medium_pool_BLIP2_captions Dataset Card for DataComp_medium_pool_BLIP2_captions Dataset Summary Supported Tasks and Leaderboards We have used this dataset for pre-training CLIP models and found that it rivals or outperforms models trained on raw web captions on average across the 38 evaluation tasks proposed by DataComp. Refer to the DataComp leaderboard (https://www.datacomp.ai/leaderboard.html) for the top baselines uncovered in our work. Languages Primarily English.… See the full description on the dataset page: https://huggingface.co/datasets/thaottn/DataComp_medium_pool_BLIP2_captions.textimage-to-text10M<n<100M5 likes98 downloads3y agoHugging Face08svjack /Nagi_no_Asukara_Videos_Captioned Reorganized version of Wild-Heart/Disney-VideoGeneration-Dataset. This is needed for Mochi-1 fine-tuning. text1K<n<10K0 likes82 downloads2y agoHugging Face09kubernetes-bad /character-captions-opusDeduplicated set of character portraits that have been described by Anthropic Claude Opus as characters with stories and visual attributes. Images obtained from CivitAI by filtering for SD XL-derived models only. Original Stable Diffusion prompt and metadata is also included. Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided. Character-like description for each image is given by Claude Opus. Here is an example: {… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/character-captions-opus.image10K<n<100K5 likes31 downloads2y agoHugging Face10dragoon49 /florence2-ofa-captions-500 OFA Florence-2 Dataset (500 Samples) This dataset was generated using microsoft/Florence-2-large on a subset of COCO 2017 Validation images. It is pre-formatted for OFA Stage-1 fine-tuning (Headerless TSV, URL-safe base64, max 512x512 resolution). tabularn<1K0 likes30 downloads2mo agoHugging Face11uygarkurt /simple-image-captionsimagen<1K5 likes28 downloads1y agoHugging Face12Anish /nepali-meme-captions NeMeme-CAP: Nepali Meme Captions Dataset Summary English-language captions generated by Google Gemini for the CHiPSAL 2026 SubtaskA Nepali Meme Datset. The context-aware captions was generated accross the training, validation, and test splits. Supported Tasks Hateful Meme Classification: Predict whether the meme is non-hateful (label=0) and hateful (label=1). Multimodal Meme Understanding: Useful as auxiliary text features or as ground-truth explanations for… See the full description on the dataset page: https://huggingface.co/datasets/Anish/nepali-meme-captions.text1K<n<10K0 likes22 downloads6mo agoHugging Face13Boni98 /PixLore-Rich-CaptionsRich image captioning dataset used for training PixLore model: https://arxiv.org/abs/2312.05349 "image_path" contains the path to the COCO dataset image (change the path accordingly), "rich_caption" contains the rich caption created using the technique described in the paper. The rest of the columns are used for debugging or improving the prompt. textimage-to-text10K<n<100K0 likes19 downloads3y agoHugging Face14tennant /imagenet-partially-captionedtext1M<n<10M0 likes18 downloads1y agoHugging Face15Arabic-Image-Captioning-latest /testimage1M<n<10M2 likes17 downloads3y agoHugging Face16wufeim /imagenet1k_captions_minigpt4 ImageNet1k Captions Generated with MiniGPT-4 MiniGPT-4 captions generated for ImageNet1k images. Can be used for training/finetuning diffusion models for image generation. ImageNet1k: link MiniGPT-4: link text1M<n<10M1 likes17 downloads2y agoHugging Face17baobab-trees /baobab_coco_evaluate_caption_24 Dataset Summary This dataset contains Japanese captions for COCO images and their English translations.The format is CSV. Dataset Structure Data Fields The data fields are the same among all lines. filename(str): The name of the COCO image file chatgpt text(str): The text generated by gpt-4o gemini text(str): The text generated by gemini-1.5-pro claude text(str): The text generated by claude-3.5-sonnet-20240620 llama text(str): The text generated by… See the full description on the dataset page: https://huggingface.co/datasets/baobab-trees/baobab_coco_evaluate_caption_24.textn<1K0 likes16 downloads2y agoHugging Face18Eunju2834 /img_captioning_oilcanvas_styletexttext-generation1K<n<10K0 likes15 downloads3y agoHugging Face19pbevan11 /GPT4V-captions-from-LVIS-typography GPT4V-captions-from-LVIS-typography by: Peter Bevan, 21 March 2023 This dataset is a typography subset of 220k-GPT4Vision-captions-from-LIVIS. This dataset comprises a subset of 8,857 captioned images from the LVIS dataset. This subset was creating by selecting only image-caption pairs which contain typography that is accurately reflected in the caption. The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/GPT4V-captions-from-LVIS-typography.image1K<n<10K1 likes15 downloads3y agoHugging Face20AIML-TUDA /t2i-diversity-gender-neutral-captionsThis dataset contains different synthetic captions for our image samples. We have selected the best-performing caption set from our experiments, the random-length captions. Then, we have used Gemma-2-9b-it and instructed it to remove different genders from the captions. We obtained three sets from the original set, namly (i) all genders neutralized, (ii) only female gender neutralized, and (iii) only male gender neutralized. To this end, we have removed all gender indicative words such as… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/t2i-diversity-gender-neutral-captions.texttext-to-image1M<n<10M1 likes14 downloads1y agoHugging Face21ccvl /3D-DST-captionsgated 3D-DST-captions As part of our data release in 3D-DST, we present MiniGPT4-generated captions for all 1000 classes in ImageNet-1k. See wufeim/DST3D for synthetic data generation with 3D annotations using the captions here. These captions can be used to produce other synthetic datasets for fair comparisons between different data generation procedures. text1M<n<10M0 likes12 downloads2y agoHugging Face22Ekaterina2002 /wikiart-captions-5k WikiArt Captions — 5k subset image_row — индекс в 5k-сабсете caption — автоописание (nlpconnect/vit-gpt2-image-captioning) Usage from datasets import load_dataset ds = load_dataset("USERNAME/wikiart-captions-5k", data_files="data/train/captions.csv") text1K<n<10K0 likes12 downloads11mo agoHugging Face23capstone-pad3 /PAD3-dataset-video-captiontextvideo-classification1K<n<10K0 likes12 downloads5mo agoHugging Face24ayush-vatsal /description_to_caption Dataset Card for Dataset Name Dataset Summary This is a tiny dataset made with the help of flicker dataset and ChatGPT. 121 image descriptions were taken from the flickr dataset and captions were AI generated, prompted to generate social media like captions. Data Fields Description, and Caption Data Splits No splits Source Data A portion of the dataset was taken from the Flickr dataset linked here:… See the full description on the dataset page: https://huggingface.co/datasets/ayush-vatsal/description_to_caption.textn<1K0 likes10 downloads3y agoHugging Face25JourneyBench /JourneyBench_Captioningimagetext-generation1K<n<10K0 likes10 downloads2y agoHugging Face26Rohan3 /ShapeNetCore_CaptionsgatedThis ShapeNetCore_Captions dataset contains text captions of all the 52,472 models of the ShapeNetCore v2 dataset. As these were generated by LLMs, there is a chance that these captions will comprise slightly incorrect phrases or information. If something incorrect is discovered during manual inspection, kindly make an effort to share your findings with me. texttext-to-3d10K<n<100K2 likes10 downloads2y agoHugging Face27Lenkashell /unsafe-safe-captions-3.0 Safety Caption Pairs Deduplicated Filtered using n-gram and word-diff rules. Contains only pairs with 1–2 word changes and minimal n-gram overlap. tabularn<1K0 likes10 downloads1y agoHugging Face28eduvedras /CaptionTemplatesDatasettextn<1K0 likes7 downloads3y agoHugging Face29dhruvindia /print_ad_Captiontextn<1K0 likes7 downloads2y agoHugging Face30selinax10010 /newyorker_caption_contest_testimage1K<n<10K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.