CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jmhessel /newyorker_caption_contest Dataset Card for New Yorker Caption Contest Benchmarks Dataset Summary See capcon.dev for more! Data from: Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest @inproceedings{hessel2023androids, title={Do Androids Laugh at Electric Sheep? {Humor} ``Understanding'' Benchmarks from {The New Yorker Caption Contest}}, author={Hessel, Jack and Marasovi{\'c}, Ana and Hwang, Jena D. and Lee, Lillian and… See the full description on the dataset page: https://huggingface.co/datasets/jmhessel/newyorker_caption_contest.imageimage-to-text100K<n<1M76 likes25k downloads3y agoHugging Face02quarterturn /danbooru-1024-eq-captioned Danbooru 1024 e/q Captioned Dataset 59,495 high-resolution (1024px) anime-style images from Danbooru's explicit and questionable rated pools. Each image includes comprehensive JSON captions generated via MiniMax-M3 with structured per-character state-of-dress inventories, camera notes, mood palettes, and post-processing detections. Directory Structure danbooru-1024-eq-captioned.parquet <- consolidated metadata manifest originals/ <-… See the full description on the dataset page: https://huggingface.co/datasets/quarterturn/danbooru-1024-eq-captioned.image10K<n<100K6 likes14k downloads1mo agoHugging Face03google-research-datasets /conceptual_captions Dataset Card for Conceptual Captions Dataset Summary Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.imageimage-to-text1M<n<10M111 likes12k downloads2y agoHugging Face04laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes6.5k downloads5y agoHugging Face05BLIP3o /BLIP3o-Pretrain-Long-Caption BLIP3o Pretrain Long-Caption Dataset This collection contains 27 million images, each paired with a long (~120 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Long-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Long-Caption.image10M<n<100M74 likes6.1k downloads1y agoHugging Face06BLIP3o /BLIP3o-Pretrain-Short-Caption BLIP3o Pretrain Short-Caption Dataset This collection contains 5 million images, each paired with a short (~20 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Short-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Short-Caption.image1M<n<10M10 likes5.6k downloads1y agoHugging Face07lambda /pokemon-blip-captionsgated Notice of DMCA Takedown Action We have received a DMCA takedown notice from The Pokémon Company International, Inc. In response to this action, we have taken down the dataset. We appreciate your understanding. imagetext-to-imagen<1K314 likes5k downloads3y agoHugging Face08lmms-lab-encoder /COCO-Caption Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2014-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption.image10K<n<100K15 likes4.3k downloads3y agoHugging Face09jxie /coco_captions Dataset Card for "coco_captions" More Information needed image100K<n<1M18 likes4k downloads3y agoHugging Face10Ryan-sjtu /ffhq512-captionimage10K<n<100K5 likes3.1k downloads3y agoHugging Face11lmms-lab-encoder /COCO-Caption2017 Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2017-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption2017.image10K<n<100K24 likes3.1k downloads3y agoHugging Face12GeroldMeisinger /laion2b-en-a65_cogvlm2-4bit_captions Abstract This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8). The synthetic images are best viewed locally by cloning this repo with: git lfs install git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.imageimage-classification1K<n<10K6 likes3k downloads2y agoHugging Face13inclusionAI /VenusBench-CAPTCHA VenusBench-CAPTCHA: A Real-World CAPTCHA Screenshot–Action Benchmark for GUI Agents Evaluation Code: https://github.com/inclusionAI/UI-Venus/tree/VenusBench-CAPTCHA Introduction CAPTCHA solving is a practical challenge for multimodal GUI agents because it requires more than isolated visual recognition. An agent must understand the challenge instruction, identify the relevant interface region, recognize or reason about the visual target, ground the result… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/VenusBench-CAPTCHA.imageimage-text-to-textn<1K6 likes2.2k downloads24d agoHugging Face14Borise /CaptionQA 📌 CaptionQA Benchmark A high-density, taxonomy-grounded benchmark for evaluating image caption quality and the alignment between image information and generated captions 📄 Paper: CaptionQA: Is Your Caption as Useful as the Image Itself? 📦 Evaluation Code: GitHub Repository Sample Usage You can load the dataset using the Hugging Face datasets library: from datasets import load_dataset # Load the entire dataset dataset = load_dataset("Borise/CaptionQA") # Load a… See the full description on the dataset page: https://huggingface.co/datasets/Borise/CaptionQA.imageimage-text-to-textn<1K9 likes2k downloads10mo agoHugging Face15ProGamerGov /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M154 likes1.9k downloads2y agoHugging Face16Felldude /Gradients_Gradients_and_Text_Full_Logic_Captionsimage1K<n<10K2 likes1.5k downloads15d agoHugging Face17limingcv /Captioned_COCOStuffimage100K<n<1M2 likes1.5k downloads3y agoHugging Face18hanlincs /InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption image10M<n<100M1 likes1.4k downloads1y agoHugging Face19yguooo /newyorker_caption_ranking New Yorker Caption Ranking Dataset Dataset Descriptions Homepage: https://nextml.github.io/caption-contest-data/ Repository: https://github.com/yguooo/cartoon-caption-generation Paper: Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning Point of Contact: yguo@cs.wisc.edu Dataset Summary We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/newyorker_caption_ranking.imagetext-generation1M<n<10M6 likes1.4k downloads2y agoHugging Face20MelanieCo /Capillary-Dataset Capillary dataset Paper: Capillary Dataset: A dataset of nail-fold capillaries captured by microscopy for diabetes detection Github: https://github.com/urgonguyen/Capillarydataset.git The dataset are structured as follows: Capillary dataset ├── Classification ├── data_1x1_224 ├── data_concat_1x9_224 ├── data_concat_2x2_224 ├── data_concat_3x3_224 ├── data_concat_4x1_224 └── data_concat_4x4_224 ├── Morphology_detection… See the full description on the dataset page: https://huggingface.co/datasets/MelanieCo/Capillary-Dataset.imageimage-classification10K<n<100K0 likes1.4k downloads3mo agoHugging Face21Multimodal-Fatima /COCO_captions_train Dataset Card for "COCO_captions_train" More Information needed image100K<n<1M7 likes1.3k downloads4y agoHugging Face22internlm /CapRL-QA-75K CapRL 75K QA Training Dataset This dataset is the carefully filtered 75K QA training set used by CapRL to train CapRL-3B, a lightweight image captioning model initialized from Qwen2.5-VL-3B. It contains 75,285 samples, where each image is paired with multiple multiple-choice QA items. The dataset is designed for the two-stage CapRL training objective, where caption quality is evaluated through answerability of visual questions. The QA construction pipeline is fully open-sourced in… See the full description on the dataset page: https://huggingface.co/datasets/internlm/CapRL-QA-75K.imageimage-text-to-text10K<n<100K5 likes1.3k downloads5mo agoHugging Face23nakasyou /captcha-like-suicaimagen<1K1 likes1.2k downloads3mo agoHugging Face24lukman48 /captcha-dataimage10K<n<100K0 likes1k downloads2mo agoHugging Face25diffusers /pokemon-gpt4-captions Dataset Card for "pokemon-gpt4-captions" This dataset is just lambdalabs/pokemon-blip-captions but the captions come from GPT-4 (Turbo). Code used to generate the captions: import base64 from io import BytesIO import requests from PIL import Image def encode_image(image): buffered = BytesIO() image.save(buffered, format="JPEG") img_str = base64.b64encode(buffered.getvalue()) returnimg_str.decode("utf-8") def create_payload(image_string): payload = {… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/pokemon-gpt4-captions.imagetext-to-imagen<1K42 likes1k downloads3y agoHugging Face26OpenCaptchaWorld /Open_CaptchaWorld Open CaptchaWorld Dataset This dataset accompanies the paper Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents. It contains 20 distinct CAPTCHA types, each testing different visual reasoning capabilities. The dataset is designed for evaluating the visual reasoning and interaction capabilities of Multimodal Large Language Model (MLLM)-powered agents. Project Page | Github The dataset includes: 20 CAPTCHA Types: A diverse set of… See the full description on the dataset page: https://huggingface.co/datasets/OpenCaptchaWorld/Open_CaptchaWorld.imagevisual-document-retrievaln<1K0 likes1k downloads1y agoHugging Face27allenai /pixmo-cap PixMo-Cap PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions. It can be used to pre-train and fine-tune vision-language models. PixMo-Cap was created by recording annotators speaking about an image for 60-90 seconds and then using the Claude large language model to turn the audio transcripts(s) into a long caption. The audio transcripts are also included. PixMo-Cap is part of the PixMo dataset collection and was used to train the Molmo family of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-cap.imageimage-to-text100K<n<1M48 likes987 downloads2y agoHugging Face28remiai3 /synthetic-captchas-library 🌍 Synthetic Multilingual CAPTCHA Library Repository: remiai3/synthetic-captchas-libraryA multilingual dataset of synthetic 4-character CAPTCHA images designed for OCR, multilingual vision models, and script recognition research. This dataset spans 44 world writing systems and is especially useful for low-resource script OCR training. 📌 Dataset Summary Each script includes 100,000 unique CAPTCHA images.The dataset is provided in two parallel formats: CSV version… See the full description on the dataset page: https://huggingface.co/datasets/remiai3/synthetic-captchas-library.image1M<n<10M1 likes985 downloads8mo agoHugging Face29LiuzhipengUCAS /CAPEval CAPEval CAPEval (Coverage And Precision Evaluation) is a checklist-based caption evaluation benchmark. It decouples caption quality into Coverage (C) and Precision (P) (0–100), and studies how each profile transfers to VLM understanding and T2I generation. Code / docs: liuzhipenggg/CAPEval Project page: liuzhipenggg.github.io/CAPEval Paper: arXiv:2608.02589 Leaderboard: leaderboard Dataset contents Path Description image/ 300 high-resolution images… See the full description on the dataset page: https://huggingface.co/datasets/LiuzhipengUCAS/CAPEval.imageimage-to-textn<1K2 likes975 downloads1mo agoHugging Face30graph-based-captions /GBC10M Graph-based captioning (GBC) is a new image annotation paradigm that combines the strengths of long captions, region captions, and scene graphs GBC interconnects region captions to create a unified description akin to a long caption, while also providing structural information similar to scene graphs. ** The associated data point can be found at demo/water_tower.json Description and data format The GBC10M dataset, derived from the original images in CC12M, is… See the full description on the dataset page: https://huggingface.co/datasets/graph-based-captions/GBC10M.imageimage-to-text10M<n<100M35 likes968 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.