datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nordjylland-news-image-captioning
Dataset Card for "nordjylland-news-image-captioning"
Dataset Summary
This dataset is a collection of image-caption pairs from the Danish newspaper TV2 Nord.
Supported Tasks and Leaderboards
Image captioning is the intended task for this dataset. No leaderboard is active at this point.
Languages
The dataset is available in Danish (da).
Dataset Structure
An example from the dataset looks as follows.
{
"file_name": "1.jpg",
"caption":… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nordjylland-news-image-captioning.Chinese_Children_Image_Captioning_Dataset_Split0
CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions)
CODP-1200: An AIGC based benchmark for assisting in child language acquisition
数据集介绍
目前已知最大的儿童图像描述数据集,children image captioning
共有1200张图片
每张图片对应五个中文描述,每两张图片为一组
描述文字600*5=3000
如果使用CODP-1200数据集,请引用以下文章
@article{LENG2024102627,
title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition},
journal = {Displays},
volume = {82},
pages = {102627},
year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.IntraOral_Gingivitis_Image_Captioning
A DENTAL INTRAORAL IMAGE DATASET OF GINGIVITIS FOR IMAGE CAPTIONING
Dataset Description
This dataset is a copy of A Dental IntraOral Image Dataset of Gingivitis for Image Captioning which is shared with the license CC BY 4.0.
This dataset contains 1,096 samples organized across multiple splits.
The dataset includes image data.
Splits
train: 732 samples
test: 182 samples
validation: 182 samples
Dataset Creation
This dataset was created using… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/IntraOral_Gingivitis_Image_Captioning.image_captions
From the Frontier Research Team at takara.ai we present over 1 million curated captioned images for multimodal text and image tasks.
Usage
from datasets import load_dataset
ds = load_dataset("takara-ai/image_captions")
print(ds)
Example
10,000 images from the dataset.
Methodology
We consolidated multiple open source datasets through an intensive 96-hour computational process across three nodes. This involved standardizing and validating the… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/image_captions.image-captioning-turkish
Türkçe Image Captioning Veri Seti
Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz.
Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.ImageCaptioning_EN-HUcoco2017_train_512x_image_caption_cannyhttps://github.com/wangherr/coco2017_for_huggingface
ImageCaptioning_SmallParquetsChinese_Children_Image_Captioning_Dataset_Split1
CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions)
CODP-1200: An AIGC based benchmark for assisting in child language acquisition
数据集介绍
目前已知最大的儿童图像描述数据集,children image captioning
共有1200张图片
每张图片对应五个中文描述,每两张图片为一组
描述文字600*5=3000
如果使用CODP-1200数据集,请引用以下文章
@article{LENG2024102627,
title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition},
journal = {Displays},
volume = {82},
pages = {102627},
year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split1.indic-multilingual-image-captions
Indic Multilingual Image Caption Dataset
This dataset contains 4,500 unique images with captions in:
English
Hindi
Bengali
Tamil
Source composition
3,000 images from COCO Caption 2017
1,500 images from TextCaps
Each image is stored once and paired with four multilingual caption
variants. Hindi, Bengali and Tamil captions were generated from the
selected English captions using the NLLB-200 distilled translation model.
Intended use
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/arikatokachi/indic-multilingual-image-captions.coco2017_train_1024x_image_caption_cannymy-image-caption-datasetcoco2017_train_512x_image_caption_depthcoco2017_train_image_captionhttps://github.com/wangherr/coco2017_for_huggingface
COCO-Image-Captioningimage-text-dataset-subset-300k-captions_onlyru-image-captions
Image Caprioning for Russian language
This dataset is a Russian part of dinhanhx/crossmodal-3600
Dataset Details
3.11k rows.
Two description for each picture. Cracked pictures were deleted from the original source.
The main feature is that all the descriptions are written by the native russian speakers.
Paper [https://google.github.io/crossmodal-3600/]
Uses
It is intended to be used for fine-tuning image captioning models.
astrobridge-image-captions
AstroBridge Legacy Survey Captions
3,487 imaging cutouts from the Legacy Survey (DR10 South + North), crossmatched against
published literature mentions and captioned in four independent stages by Gemini
(gemini-3.7-flash), following the AstroLLaVA data-generation approach (Zaman et al. 2025,
arXiv:2504.08583): no caption is ever told the object's
real name or catalog designation, and no caption states a fact that isn't derivable from the
pixels or the (redacted-at-the-model… See the full description on the dataset page: https://huggingface.co/datasets/gapatron/astrobridge-image-captions.ru_image_captioningmy_image_captioning_datasetenhanced_image_captionsnordjylland-news-image-captioningimage_captions_x
Dataset Card for image_captions_x (URL + Caption)
This dataset provides a lightweight, web-scale resource of image-caption pairs in the form of URLs and their associated textual descriptions (captions). It is designed for training and evaluating vision-language models where users retrieve images independently from the provided links.
This dataset card is based on the Hugging Face dataset card template.
Dataset Details
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/image_captions_x.Prince_Xiang_qwen_image_2509_F2P_inpaint_repose_zh_caption
reference face
grid image
imageCaptioningImageCaptions-7M-Translations-Arabicnepali-image-caption-datasetimage-captioning-idPersian-Image-Captioning
Dataset Card for "Persian-Image-Captioning"
More Information needed
Peng_UNO_qwen_image_outpainting_Keye_ZH_Captioned
