CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TommyBsk /Embodied-Captioning Embodied Image Captioning – Manually Annotated Test Set Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning 📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.tabularimage-to-text1K<n<10K0 likes8.8k downloads1y agoHugging Face02hf-internal-testing /fixtures-captioning\\n0 likes2.8k downloads1y agoHugging Face03zhiqiulin /video_captioningvideo1K<n<10K0 likes1.5k downloads1y agoHugging Face04laion /Segmentation-Captioning-Assistant-Tuning-Data1 likes1.4k downloads10mo agoHugging Face05alexandrainst /nordjylland-news-image-captioning Dataset Card for "nordjylland-news-image-captioning" Dataset Summary This dataset is a collection of image-caption pairs from the Danish newspaper TV2 Nord. Supported Tasks and Leaderboards Image captioning is the intended task for this dataset. No leaderboard is active at this point. Languages The dataset is available in Danish (da). Dataset Structure An example from the dataset looks as follows. { "file_name": "1.jpg", "caption":… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nordjylland-news-image-captioning.imageimage-to-text10K<n<100K4 likes745 downloads3y agoHugging Face06svjack /Chinese_Children_Image_Captioning_Dataset_Split0 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.image1K<n<10K0 likes533 downloads1y agoHugging Face07ekacare /IntraOral_Gingivitis_Image_Captioning A DENTAL INTRAORAL IMAGE DATASET OF GINGIVITIS FOR IMAGE CAPTIONING Dataset Description This dataset is a copy of A Dental IntraOral Image Dataset of Gingivitis for Image Captioning which is shared with the license CC BY 4.0. This dataset contains 1,096 samples organized across multiple splits. The dataset includes image data. Splits train: 732 samples test: 182 samples validation: 182 samples Dataset Creation This dataset was created using… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/IntraOral_Gingivitis_Image_Captioning.imageimage-classification1K<n<10K0 likes531 downloads1y agoHugging Face08ituperceptron /image-captioning-turkish Türkçe Image Captioning Veri Seti Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz. Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.imageimage-to-text1M<n<10M7 likes490 downloads8mo agoHugging Face09PerRing /coco_captioning_complete_formatimage100K<n<1M0 likes414 downloads11mo agoHugging Face10AKCIT /coco2017-captioningimage100K<n<1M0 likes412 downloads6mo agoHugging Face11svjack /Chinese_Children_Image_Captioning_Dataset_Split1 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split1.image1K<n<10K0 likes368 downloads1y agoHugging Face12TMICCProj /Closed_Captioning_Lecture_Datasetaudio10K<n<100K0 likes364 downloads6mo agoHugging Face13DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes323 downloads2y agoHugging Face14TeeA /Pokemon-Captioning-Classification 2000+ download monthly. Really appreciate for all of you guys: Buy me a coffee: https://buymeacoffee.com/tridoan Disclaimer: This model is provided "as-is" without any warranties. The authors are not responsible for any misuse or damages arising from its use. image1K<n<10K2 likes197 downloads21d agoHugging Face15MagiBoss /COCO-Image-Captioningimage100K<n<1M0 likes172 downloads2y agoHugging Face16gijs /tacos-captioningaudio10K<n<100K1 likes167 downloads1y agoHugging Face17fancyfeast /joy-captioning-20250408aThis is the dataset used to do the initial training for JoyCaption Beta One (https://huggingface.co/fancyfeast/llama-joycaption-beta-one-hf-llava), before post-training. Contents Most of the dataset focusses on descriptions and captions for images, with a smaller subset covering general VQA tasks. Some of the questions and answers are human written, some are automated, some are machine written. The is_human column is True when the answer text is human written. WARNING… See the full description on the dataset page: https://huggingface.co/datasets/fancyfeast/joy-captioning-20250408a.textvisual-question-answering100K<n<1M10 likes142 downloads7mo agoHugging Face18gorovuha /ru_image_captioningimage1K<n<10K0 likes140 downloads2y agoHugging Face19fancyfeast /joy-captioning-20250328b Work In Progress I'm still going back through my data to add in the URLs. textvisual-question-answering1M<n<10M24 likes136 downloads1y agoHugging Face20tavish-mishra /my_image_captioning_datasetimage10K<n<100K1 likes133 downloads1y agoHugging Face21alinasdkey /graph-captioning-train-onlyimagen<1K0 likes133 downloads1y agoHugging Face22NTQAI /MSVD-Video-Captioning-Vi MSVD-Video-Captioning-Vi 📌 Overview MSVD-Video-Captioning-Vi is a Vietnamese video captioning dataset derived from the MSVD dataset originally hosted by the user friedrichor on Hugging Face. This dataset provides Vietnamese captions for short video clips and is intended for: Video captioning research Vision–Language model training Multimodal instruction tuning Video-to-text generation 🔁 Dataset Origin This dataset is a translated and derived version… See the full description on the dataset page: https://huggingface.co/datasets/NTQAI/MSVD-Video-Captioning-Vi.textvisual-question-answering1K<n<10K4 likes132 downloads8mo agoHugging Face23orzhan /minecraft-captioningimageimage-to-textn<1K0 likes126 downloads3y agoHugging Face24ChristophSchuhmann /Mega-Fast-KNN-Captioningtext1B<n<10B2 likes114 downloads3y agoHugging Face25sylvan54 /Bean_Captioning_Datasettext10K<n<100K0 likes109 downloads2y agoHugging Face26Image-Captioning-ML /ucf101-captioned-mappedtext1K<n<10K0 likes105 downloads1y agoHugging Face27MykMaks /nordjylland-news-image-captioningimagezero-shot-classification10K<n<100K1 likes103 downloads2y agoHugging Face28MSEarth /MSEarth_Captioningimage1K<n<10K2 likes97 downloads1y agoHugging Face29roshbeed /ai-residency-multimodal-captioning-dataimagen<1K0 likes87 downloads1mo agoHugging Face30SEACrowd /bloom_captioningThis is a Bloom Library dataset developed for the image captioning task. It covers 74 languages indigenous to SEA overall, amounting to total data of 21K. This dataset belongs to a CC license, where its datapoints has specific license attached to it. Before using this dataloader, please accept the acknowledgement at https://huggingface.co/datasets/sil-ai/bloom-captioning and use huggingface-cli login for authentication. Languages abc, ahk, bfn, bjn, bkx, brb, brv, bya, bzi, ceb… See the full description on the dataset page: https://huggingface.co/datasets/SEACrowd/bloom_captioning.0 likes83 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.