CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jmhessel /newyorker_caption_contest Dataset Card for New Yorker Caption Contest Benchmarks Dataset Summary See capcon.dev for more! Data from: Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest @inproceedings{hessel2023androids, title={Do Androids Laugh at Electric Sheep? {Humor} ``Understanding'' Benchmarks from {The New Yorker Caption Contest}}, author={Hessel, Jack and Marasovi{\'c}, Ana and Hwang, Jena D. and Lee, Lillian and… See the full description on the dataset page: https://huggingface.co/datasets/jmhessel/newyorker_caption_contest.imageimage-to-text100K<n<1M76 likes25k downloads3y agoHugging Face02argilla /distilabel-capybara-dpo-7k-binarized Capybara-DPO 7K binarized A DPO dataset built with distilabel atop the awesome LDJnr/Capybara This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models. Why? Multi-turn dialogue data is key to fine-tune capable chat models. Multi-turn preference data has been used by the most relevant RLHF works (Anthropic, Meta Llama2, etc.). Unfortunately, there are very few… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-dpo-7k-binarized.tabularquestion-answering1K<n<10K184 likes23k downloads2y agoHugging Face03Hemabhushan /capstone_sakuga_preproc_optical_flowtabular100K<n<1M0 likes14k downloads2y agoHugging Face04quarterturn /danbooru-1024-eq-captioned Danbooru 1024 e/q Captioned Dataset 59,495 high-resolution (1024px) anime-style images from Danbooru's explicit and questionable rated pools. Each image includes comprehensive JSON captions generated via MiniMax-M3 with structured per-character state-of-dress inventories, camera notes, mood palettes, and post-processing detections. Directory Structure danbooru-1024-eq-captioned.parquet <- consolidated metadata manifest originals/ <-… See the full description on the dataset page: https://huggingface.co/datasets/quarterturn/danbooru-1024-eq-captioned.image10K<n<100K6 likes14k downloads1mo agoHugging Face05google-research-datasets /conceptual_captions Dataset Card for Conceptual Captions Dataset Summary Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.imageimage-to-text1M<n<10M111 likes12k downloads2y agoHugging Face06tiange /Cap3DThis repository hosts data for Scalable 3D Captioning with Pretrained Models and View Selection for 3D Captioning via Diffusion Ranking, including descriptive captions for 3D objects in Objaverse, Objaverse-XL, ABO, and ShapeNet. This repo also includes point clouds and rendered images with camera, depth, and MatAlpha information of Objaverse objects, as well as their Shap-E latent codes. All the captions and data provided by our papers are released under ODC-By 1.0 license. Important… See the full description on the dataset page: https://huggingface.co/datasets/tiange/Cap3D.text-to-3d127 likes11k downloads8mo agoHugging Face07TommyBsk /Embodied-Captioning Embodied Image Captioning – Manually Annotated Test Set Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning 📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.tabularimage-to-text1K<n<10K0 likes8.8k downloads1y agoHugging Face08laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes6.5k downloads5y agoHugging Face09BLIP3o /BLIP3o-Pretrain-Long-Caption BLIP3o Pretrain Long-Caption Dataset This collection contains 27 million images, each paired with a long (~120 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Long-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Long-Caption.image10M<n<100M74 likes6.1k downloads1y agoHugging Face10Hemabhushan /capstone_mlm_hidden_statestabular100K<n<1M0 likes5.7k downloads2y agoHugging Face11BLIP3o /BLIP3o-Pretrain-Short-Caption BLIP3o Pretrain Short-Caption Dataset This collection contains 5 million images, each paired with a short (~20 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Short-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Short-Caption.image1M<n<10M10 likes5.6k downloads1y agoHugging Face12laion /laions_got_talent_enhanced_flash_annotations_and_long_captions18 likes5.4k downloads2y agoHugging Face13lambda /pokemon-blip-captionsgated Notice of DMCA Takedown Action We have received a DMCA takedown notice from The Pokémon Company International, Inc. In response to this action, we have taken down the dataset. We appreciate your understanding. imagetext-to-imagen<1K314 likes5k downloads3y agoHugging Face14lingamvamshikrishnareddy /ramanv-image-captions-realtext100K<n<1M0 likes5k downloads27d agoHugging Face15webshart /conceptual-captions-12m-webdataset-metadata Conceptual Captions 12M — Webshart metadata indices Per-shard webshart metadata indices for laion/conceptual-captions-12m-webdataset: 1,100 JSON files under data/, one per source tar shard, mirroring the source's shard layout. Each index records every tar member's byte offset and length (enabling ranged reads without downloading whole shards), image geometry (width/height for aspect bucketing), and — as of August 2026 — embedded captions for all 10,994,853 samples, coalesced… See the full description on the dataset page: https://huggingface.co/datasets/webshart/conceptual-captions-12m-webdataset-metadata.1 likes4.8k downloads1mo agoHugging Face16trl-lib /Capybaratext10K<n<100K26 likes4.5k downloads2y agoHugging Face17lmms-lab-encoder /COCO-Caption Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2014-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption.image10K<n<100K15 likes4.3k downloads3y agoHugging Face18jxie /coco_captions Dataset Card for "coco_captions" More Information needed image100K<n<1M18 likes4k downloads3y agoHugging Face19RoboCOIN /AgiBot-g1_left_capture_partgated AgiBot-g1_left_capture_part 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: ruantong_a2d | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: factory 🤖 Atomic Actions This dataset includes the following atomic actions: grasp 📊 Dataset Statistics Metric Value Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_left_capture_part.tabularrobotics100K<n<1M0 likes4k downloads9mo agoHugging Face20AbstractPhil /conceptual-captions-12m-webdataset-bertstext10M<n<100M1 likes3.9k downloads2mo agoHugging Face21architect-ubc-capstone /rtl-augmented-v3 RTL Bug Fix — Augmented Dataset Auto-generated dashboard snapshot (2026-04-14T10:53:43). Overview Metric Value Total problems 718 Repos with data 57 / 81 Modules augmented 408 Bug types 11/11 Augmentation success 48.9% Coverage Distribution Augmentation Health Topic Coverage Warnings lucky-wfw_IC_System_Design: 0 problems from 48 attempts — likely systematic sim issue meiniKi_FazyRV:… See the full description on the dataset page: https://huggingface.co/datasets/architect-ubc-capstone/rtl-augmented-v3.text-generation1K<n<10K0 likes3.7k downloads5mo agoHugging Face22nvidia /BridgeData2-Subset-Synthetic-Captions BridgeData2 Subset Synthetic Captions Dataset Summary nvidia/BridgeData2-Subset-Synthetic-Captions is a subset of BridgeData V2 packaged with short robot-manipulation video clips and synthetic video captions. It is intended for supervised fine-tuning (SFT), prompt generation, and evaluation workflows involving text-to-video, image-to-video, and video-to-video generation of robot manipulation scenes. The source data is derived from BridgeData V2, a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/BridgeData2-Subset-Synthetic-Captions.10 likes3.7k downloads4mo agoHugging Face23ming030890 /youtube_caption_yue YouTube ASR Caption Dataset (Cantonese) This dataset was built from YouTube videos with manually provided captions in Cantonese. We used SenseVoice to re-transcribe the audio and filtered segments to build a high-quality collection of audio-caption pairs. What’s included Segments where the ASR output is identical to the original caption — likely clean. Segments where differences are only homophones (同音字) or English words — likely ASR mistakes. This combination supports… See the full description on the dataset page: https://huggingface.co/datasets/ming030890/youtube_caption_yue.audio10K<n<100K2 likes3.5k downloads1y agoHugging Face24kdexd /red_capsRedCaps is a large-scale dataset of 12M image-text pairs collected from Reddit. Images and captions from Reddit depict and describe a wide variety of objects and scenes. The data is collected from a manually curated set of subreddits (350 total), which give coarse image labels and allow steering of the dataset composition without labeling individual instances.image-to-text10M<n<100M60 likes3.3k downloads3y agoHugging Face25Ryan-sjtu /ffhq512-captionimage10K<n<100K5 likes3.1k downloads3y agoHugging Face26lmms-lab-encoder /COCO-Caption2017 Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2017-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption2017.image10K<n<100K24 likes3.1k downloads3y agoHugging Face27GeroldMeisinger /laion2b-en-a65_cogvlm2-4bit_captions Abstract This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8). The synthetic images are best viewed locally by cloning this repo with: git lfs install git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.imageimage-classification1K<n<10K6 likes3k downloads2y agoHugging Face28hf-internal-testing /fixtures-captioning\\n0 likes2.7k downloads1y agoHugging Face29capleaf /viVoicegated Important Note ⚠️ This dataset is only to be used for research purposes. Access requests must be made via your school, institution, or work email. Requests from common email services will be rejected. We apologize for any inconvenience. viVoice: Enabling Vietnamese Multi-Speaker Speech Synthesis For a comprehensive description, please visit https://github.com/thinhlpg/viVoice This dataset is licensed under CC-BY-NC-SA-4.0 and is intended for research purposes only.… See the full description on the dataset page: https://huggingface.co/datasets/capleaf/viVoice.audiotext-to-speech100K<n<1M93 likes2.5k downloads2y agoHugging Face30friedrichor /ActivityNet_Captions About ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark. We adopt the official split: Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions) Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.texttext-to-video10K<n<100K15 likes2.4k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.