CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jmhessel /newyorker_caption_contest Dataset Card for New Yorker Caption Contest Benchmarks Dataset Summary See capcon.dev for more! Data from: Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest @inproceedings{hessel2023androids, title={Do Androids Laugh at Electric Sheep? {Humor} ``Understanding'' Benchmarks from {The New Yorker Caption Contest}}, author={Hessel, Jack and Marasovi{\'c}, Ana and Hwang, Jena D. and Lee, Lillian and… See the full description on the dataset page: https://huggingface.co/datasets/jmhessel/newyorker_caption_contest.imageimage-to-text100K<n<1M76 likes27k downloads3y agoHugging Face02google-research-datasets /conceptual_captions Dataset Card for Conceptual Captions Dataset Summary Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.imageimage-to-text1M<n<10M111 likes12k downloads2y agoHugging Face03lambda /pokemon-blip-captionsgated Notice of DMCA Takedown Action We have received a DMCA takedown notice from The Pokémon Company International, Inc. In response to this action, we have taken down the dataset. We appreciate your understanding. imagetext-to-imagen<1K314 likes5.4k downloads3y agoHugging Face04lmms-lab-encoder /COCO-Caption Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2014-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption.image10K<n<100K15 likes4.2k downloads3y agoHugging Face05jxie /coco_captions Dataset Card for "coco_captions" More Information needed image100K<n<1M18 likes4k downloads3y agoHugging Face06lmms-lab-encoder /COCO-Caption2017 Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2017-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption2017.image10K<n<100K24 likes3.1k downloads3y agoHugging Face07Ryan-sjtu /ffhq512-captionimage10K<n<100K5 likes3.1k downloads3y agoHugging Face08Borise /CaptionQA 📌 CaptionQA Benchmark A high-density, taxonomy-grounded benchmark for evaluating image caption quality and the alignment between image information and generated captions 📄 Paper: CaptionQA: Is Your Caption as Useful as the Image Itself? 📦 Evaluation Code: GitHub Repository Sample Usage You can load the dataset using the Hugging Face datasets library: from datasets import load_dataset # Load the entire dataset dataset = load_dataset("Borise/CaptionQA") # Load a… See the full description on the dataset page: https://huggingface.co/datasets/Borise/CaptionQA.imageimage-text-to-textn<1K9 likes1.8k downloads10mo agoHugging Face09limingcv /Captioned_COCOStuffimage100K<n<1M2 likes1.5k downloads3y agoHugging Face10Multimodal-Fatima /COCO_captions_train Dataset Card for "COCO_captions_train" More Information needed image100K<n<1M7 likes1.3k downloads4y agoHugging Face11diffusers /pokemon-gpt4-captions Dataset Card for "pokemon-gpt4-captions" This dataset is just lambdalabs/pokemon-blip-captions but the captions come from GPT-4 (Turbo). Code used to generate the captions: import base64 from io import BytesIO import requests from PIL import Image def encode_image(image): buffered = BytesIO() image.save(buffered, format="JPEG") img_str = base64.b64encode(buffered.getvalue()) returnimg_str.decode("utf-8") def create_payload(image_string): payload = {… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/pokemon-gpt4-captions.imagetext-to-imagen<1K42 likes1k downloads3y agoHugging Face12unography /movie-scenes-captionedimage100K<n<1M0 likes1k downloads2y agoHugging Face13shunk031 /STAIR-Captions Dataset Card for STAIR-Captions Dataset Summary STAIR Captions is a large-scale dataset containing 820,310 Japanese captions. This dataset can be used for caption generation, multimodal retrieval, and image generation. Supported Tasks and Leaderboards [More Information Needed] Languages The language data in JDocQA is in Japanese (BCP-47 ja-JP). Dataset Structure Data Instances [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/shunk031/STAIR-Captions.imageimage-to-text100K<n<1M5 likes984 downloads2y agoHugging Face14laion /220k-GPT4Vision-captions-from-LIVIS 220k-GPT4Vision-captions-from-LVIS by: Christoph Schuhmann, Peter Bevan, 21 Nov, 2023 This dataset comprises 220,000 captioned images from the LVIS dataset. The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted into captions using Mistral-7B-OpenOrca. PROMPT """<<SYS>> You are a highly intelligent, empathic, helpful, respectful, and honest assistant with high emotional intelligence. Always… See the full description on the dataset page: https://huggingface.co/datasets/laion/220k-GPT4Vision-captions-from-LIVIS.image100K<n<1M64 likes955 downloads3y agoHugging Face15graph-based-captions /GBC10M Graph-based captioning (GBC) is a new image annotation paradigm that combines the strengths of long captions, region captions, and scene graphs GBC interconnects region captions to create a unified description akin to a long caption, while also providing structural information similar to scene graphs. ** The associated data point can be found at demo/water_tower.json Description and data format The GBC10M dataset, derived from the original images in CC12M, is… See the full description on the dataset page: https://huggingface.co/datasets/graph-based-captions/GBC10M.imageimage-to-text10M<n<100M35 likes947 downloads2y agoHugging Face16alexandrainst /nordjylland-news-image-captioning Dataset Card for "nordjylland-news-image-captioning" Dataset Summary This dataset is a collection of image-caption pairs from the Danish newspaper TV2 Nord. Supported Tasks and Leaderboards Image captioning is the intended task for this dataset. No leaderboard is active at this point. Languages The dataset is available in Danish (da). Dataset Structure An example from the dataset looks as follows. { "file_name": "1.jpg", "caption":… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nordjylland-news-image-captioning.imageimage-to-text10K<n<100K4 likes946 downloads3y agoHugging Face17Multimodal-Fatima /COCO_captions_validation Dataset Card for "COCO_captions_validation" More Information needed image1K<n<10K0 likes843 downloads4y agoHugging Face18Obscure-Entropy /CONCEPTUAL_CAPTIONS_HU_FILTEREDimage1M<n<10M0 likes743 downloads2y agoHugging Face19Joctor /bokete_oogiri_captionimage100K<n<1M3 likes725 downloads2y agoHugging Face20fusing /wikiart_captionsimage1K<n<10K12 likes664 downloads4y agoHugging Face21BootsofLagrangian /danbooru-multitier-captions-202606 Danbooru — multi-tier captions (202606) Per-post Danbooru data for the 202606 crawl: native tags, the raw API metadata, model-generated multi-tier natural-language captions (long / refined long / medium / short), and post flags. One row per Danbooru post_id. Images are not included — each post is referenced by post_id, danbooru_url, md5, and the Danbooru CDN URLs. (The two example previews below are downscaled for illustration.) Based on:… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/danbooru-multitier-captions-202606.imagetext-to-image10M<n<100M2 likes649 downloads1mo agoHugging Face22none-yet /anime-captionsimage100K<n<1M30 likes643 downloads10mo agoHugging Face23laicsiifes /coco-captions-pt-br 🎉 COCO Captions Dataset Translation for Portuguese Image Captioning 💾 Dataset Summary COCO Captions Portuguese Translation, a multimodal dataset for Portuguese image captioning with 123,287 images, each accompanied by five descriptive captions that have been generated by human annotators for every individual image. The original English captions were rendered into Portuguese through the utilization of the Google Translator API. 🧑‍💻 Hot to Get… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/coco-captions-pt-br.imagetext-to-image100K<n<1M6 likes597 downloads4mo agoHugging Face24irodkin /ffhq_with_llava_shorter_captions Dataset Card for "ffhq_with_llava_shorter_captions" More Information needed image10K<n<100K2 likes551 downloads3y agoHugging Face25ekacare /IntraOral_Gingivitis_Image_Captioning A DENTAL INTRAORAL IMAGE DATASET OF GINGIVITIS FOR IMAGE CAPTIONING Dataset Description This dataset is a copy of A Dental IntraOral Image Dataset of Gingivitis for Image Captioning which is shared with the license CC BY 4.0. This dataset contains 1,096 samples organized across multiple splits. The dataset includes image data. Splits train: 732 samples test: 182 samples validation: 182 samples Dataset Creation This dataset was created using… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/IntraOral_Gingivitis_Image_Captioning.imageimage-classification1K<n<10K0 likes536 downloads1y agoHugging Face26summykai /minecraft-skins-captioned-900k 🎮 minecraft-skins-captioned-900k 854,116 high-quality, captioned Minecraft player skins — deduplicated, Steve-model only, ready for text-to-image training. 📋 Dataset Summary A rigorously filtered and quality-controlled version of neurlang/Minecraft-Skins-Captioned-1M specifically curated for training high-performance generative models that require precise UV topology constraints. This dataset is optimized for models like ST-DiT (Sparse Template-Aware… See the full description on the dataset page: https://huggingface.co/datasets/summykai/minecraft-skins-captioned-900k.imagetext-to-image100K<n<1M4 likes519 downloads10mo agoHugging Face27ituperceptron /image-captioning-turkish Türkçe Image Captioning Veri Seti Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz. Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.imageimage-to-text1M<n<10M7 likes506 downloads8mo agoHugging Face28PursuitOfDataScience /llama4-maverick-coco-captionsimage100K<n<1M0 likes500 downloads1y agoHugging Face29takara-ai /image_captions From the Frontier Research Team at takara.ai we present over 1 million curated captioned images for multimodal text and image tasks. Usage from datasets import load_dataset ds = load_dataset("takara-ai/image_captions") print(ds) Example 10,000 images from the dataset. Methodology We consolidated multiple open source datasets through an intensive 96-hour computational process across three nodes. This involved standardizing and validating the… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/image_captions.imagetext-to-image1M<n<10M24 likes487 downloads2y agoHugging Face30PerRing /coco_captioning_complete_formatimage100K<n<1M0 likes471 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.