CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexandrainst /nordjylland-news-image-captioning Dataset Card for "nordjylland-news-image-captioning" Dataset Summary This dataset is a collection of image-caption pairs from the Danish newspaper TV2 Nord. Supported Tasks and Leaderboards Image captioning is the intended task for this dataset. No leaderboard is active at this point. Languages The dataset is available in Danish (da). Dataset Structure An example from the dataset looks as follows. { "file_name": "1.jpg", "caption":… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nordjylland-news-image-captioning.imageimage-to-text10K<n<100K4 likes750 downloads3y agoHugging Face02svjack /Chinese_Children_Image_Captioning_Dataset_Split0 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.image1K<n<10K0 likes610 downloads1y agoHugging Face03ekacare /IntraOral_Gingivitis_Image_Captioning A DENTAL INTRAORAL IMAGE DATASET OF GINGIVITIS FOR IMAGE CAPTIONING Dataset Description This dataset is a copy of A Dental IntraOral Image Dataset of Gingivitis for Image Captioning which is shared with the license CC BY 4.0. This dataset contains 1,096 samples organized across multiple splits. The dataset includes image data. Splits train: 732 samples test: 182 samples validation: 182 samples Dataset Creation This dataset was created using… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/IntraOral_Gingivitis_Image_Captioning.imageimage-classification1K<n<10K0 likes531 downloads1y agoHugging Face04takara-ai /image_captions From the Frontier Research Team at takara.ai we present over 1 million curated captioned images for multimodal text and image tasks. Usage from datasets import load_dataset ds = load_dataset("takara-ai/image_captions") print(ds) Example 10,000 images from the dataset. Methodology We consolidated multiple open source datasets through an intensive 96-hour computational process across three nodes. This involved standardizing and validating the… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/image_captions.imagetext-to-image1M<n<10M24 likes494 downloads2y agoHugging Face05ituperceptron /image-captioning-turkish Türkçe Image Captioning Veri Seti Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz. Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.imageimage-to-text1M<n<10M7 likes487 downloads8mo agoHugging Face06Obscure-Entropy /ImageCaptioning_EN-HUimage10M<n<100M1 likes480 downloads1y agoHugging Face07wangherr /coco2017_train_512x_image_caption_cannyhttps://github.com/wangherr/coco2017_for_huggingface image100K<n<1M2 likes423 downloads1y agoHugging Face08Obscure-Entropy /ImageCaptioning_SmallParquetsimage1M<n<10M0 likes402 downloads1y agoHugging Face09svjack /Chinese_Children_Image_Captioning_Dataset_Split1 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split1.image1K<n<10K0 likes367 downloads1y agoHugging Face10arikatokachi /indic-multilingual-image-captions Indic Multilingual Image Caption Dataset This dataset contains 4,500 unique images with captions in: English Hindi Bengali Tamil Source composition 3,000 images from COCO Caption 2017 1,500 images from TextCaps Each image is stored once and paired with four multilingual caption variants. Hindi, Bengali and Tamil captions were generated from the selected English captions using the NLLB-200 distilled translation model. Intended use The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/arikatokachi/indic-multilingual-image-captions.imageimage-to-text1K<n<10K0 likes313 downloads2mo agoHugging Face11wangherr /coco2017_train_1024x_image_caption_cannyimage100K<n<1M0 likes292 downloads1y agoHugging Face12Kanrawee /my-image-caption-datasetimage100K<n<1M1 likes236 downloads2y agoHugging Face13wangherr /coco2017_train_512x_image_caption_depthimage100K<n<1M7 likes216 downloads2y agoHugging Face14wangherr /coco2017_train_image_captionhttps://github.com/wangherr/coco2017_for_huggingface image100K<n<1M0 likes207 downloads1y agoHugging Face15MagiBoss /COCO-Image-Captioningimage100K<n<1M0 likes183 downloads2y agoHugging Face16theojiang /image-text-dataset-subset-300k-captions_onlyimage100K<n<1M1 likes177 downloads3y agoHugging Face17gorovuha /ru-image-captions Image Caprioning for Russian language This dataset is a Russian part of dinhanhx/crossmodal-3600 Dataset Details 3.11k rows. Two description for each picture. Cracked pictures were deleted from the original source. The main feature is that all the descriptions are written by the native russian speakers. Paper [https://google.github.io/crossmodal-3600/] Uses It is intended to be used for fine-tuning image captioning models. imageimage-to-text1K<n<10K4 likes171 downloads2y agoHugging Face18gapatron /astrobridge-image-captions AstroBridge Legacy Survey Captions 3,487 imaging cutouts from the Legacy Survey (DR10 South + North), crossmatched against published literature mentions and captioned in four independent stages by Gemini (gemini-3.7-flash), following the AstroLLaVA data-generation approach (Zaman et al. 2025, arXiv:2504.08583): no caption is ever told the object's real name or catalog designation, and no caption states a fact that isn't derivable from the pixels or the (redacted-at-the-model… See the full description on the dataset page: https://huggingface.co/datasets/gapatron/astrobridge-image-captions.imageimage-to-text1K<n<10K0 likes162 downloads5d agoHugging Face19gorovuha /ru_image_captioningimage1K<n<10K0 likes140 downloads2y agoHugging Face20tavish-mishra /my_image_captioning_datasetimage10K<n<100K1 likes134 downloads1y agoHugging Face21fwd4xl /enhanced_image_captionsimage100K<n<1M0 likes115 downloads1y agoHugging Face22MykMaks /nordjylland-news-image-captioningimagezero-shot-classification10K<n<100K1 likes103 downloads2y agoHugging Face23kamruzzaman-asif /image_captions_x Dataset Card for image_captions_x (URL + Caption) This dataset provides a lightweight, web-scale resource of image-caption pairs in the form of URLs and their associated textual descriptions (captions). It is designed for training and evaluating vision-language models where users retrieve images independently from the provided links. This dataset card is based on the Hugging Face dataset card template. Dataset Details Dataset Description This… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/image_captions_x.image100M<n<1B0 likes95 downloads1y agoHugging Face24svjack /Prince_Xiang_qwen_image_2509_F2P_inpaint_repose_zh_caption reference face grid image imagen<1K0 likes83 downloads9mo agoHugging Face25aryananand19 /imageCaptioningimage10K<n<100K0 likes79 downloads2y agoHugging Face26Arabic-Clip-Archive /ImageCaptions-7M-Translations-Arabicimage100K<n<1M1 likes74 downloads3y agoHugging Face27Khanalnishan /nepali-image-caption-datasetimage10K<n<100K0 likes73 downloads1y agoHugging Face28indrad123 /image-captioning-idimage1K<n<10K0 likes69 downloads2y agoHugging Face29SeyedAli /Persian-Image-Captioning Dataset Card for "Persian-Image-Captioning" More Information needed image10K<n<100K2 likes62 downloads3y agoHugging Face30svjack /Peng_UNO_qwen_image_outpainting_Keye_ZH_Captionedimagen<1K0 likes62 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.