datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image_captions
From the Frontier Research Team at takara.ai we present over 1 million curated captioned images for multimodal text and image tasks.
Usage
from datasets import load_dataset
ds = load_dataset("takara-ai/image_captions")
print(ds)
Example
10,000 images from the dataset.
Methodology
We consolidated multiple open source datasets through an intensive 96-hour computational process across three nodes. This involved standardizing and validating the… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/image_captions.image-text-dataset-subset-300k-captions_onlyru-image-captions
Image Caprioning for Russian language
This dataset is a Russian part of dinhanhx/crossmodal-3600
Dataset Details
3.11k rows.
Two description for each picture. Cracked pictures were deleted from the original source.
The main feature is that all the descriptions are written by the native russian speakers.
Paper [https://google.github.io/crossmodal-3600/]
Uses
It is intended to be used for fine-tuning image captioning models.
astrobridge-image-captions
AstroBridge Legacy Survey Captions
3,487 imaging cutouts from the Legacy Survey (DR10 South + North), crossmatched against
published literature mentions and captioned in four independent stages by Gemini
(gemini-3.7-flash), following the AstroLLaVA data-generation approach (Zaman et al. 2025,
arXiv:2504.08583): no caption is ever told the object's
real name or catalog designation, and no caption states a fact that isn't derivable from the
pixels or the (redacted-at-the-model… See the full description on the dataset page: https://huggingface.co/datasets/gapatron/astrobridge-image-captions.indic-multilingual-image-captions
Indic Multilingual Image Caption Dataset
This dataset contains 4,500 unique images with captions in:
English
Hindi
Bengali
Tamil
Source composition
3,000 images from COCO Caption 2017
1,500 images from TextCaps
Each image is stored once and paired with four multilingual caption
variants. Hindi, Bengali and Tamil captions were generated from the
selected English captions using the NLLB-200 distilled translation model.
Intended use
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/arikatokachi/indic-multilingual-image-captions.enhanced_image_captionsimage_captions_x
Dataset Card for image_captions_x (URL + Caption)
This dataset provides a lightweight, web-scale resource of image-caption pairs in the form of URLs and their associated textual descriptions (captions). It is designed for training and evaluating vision-language models where users retrieve images independently from the provided links.
This dataset card is based on the Hugging Face dataset card template.
Dataset Details
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/image_captions_x.ImageCaptions-7M-Translations-Arabicimage-text-dataset-subset-300k-captions_only_with_latentsvg-captions-graphs-processed-image-graphssimple-image-captionsImageCaptions-7M-Translationsversion https://git-lfs.github.com/spec/v1
oid sha256:835f3f7d88a86e05a882c6a6b6333da6ab874776385f85473798769d767c2fca
size 27
image-captions-datasetdottrmstr-qwen-image-short-captions-no-style-datasetPokemon-Card-Plus-Pokemon-Actual-Image-And-Captions-13000
How this Dataset was made
Combined these 3 datasets together, dropping columns and downloading images to format them into image formats and just keeping the images and descriptions.
https://huggingface.co/datasets/TheFusion21/PokemonCards
https://huggingface.co/datasets/wanghaofan/pokemon-wiki-captions?row=3
https://huggingface.co/datasets/JoaoFassina/pokemon_anotated
sentinel-image-captionssmall-stock-image-captionsadaption-multilingual-image-captions
This dataset is a remastered version of
Reubencf/multilingual-image-annotations
prepared using Adaption's Adaptive Data platform.
Multilingual Image Captions (Adaption)
462 image rows with English + 6-language captions, 21 VQA pairs per image
(3 per language), conditional object detections with normalized bounding
boxes, and Adaption-sharpened enhanced_prompt / enhanced_completion
columns. Images and bounding-box visualizations are included as the first
two columns.… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-multilingual-image-captions.ru-image-captions-test1000image-text-dataset-subset-300k-captions_textmap-image-captionsunsafe_shocking_image_captions
Dataset Card for "unsafe_shocking_image_captions"
More Information needed
unsafe_violence_image_captions
Dataset Card for "unsafe_violence_image_captions"
More Information needed
Pokemon_Card_Actual_Image_And_Captions_5000Pokemon_Card_Actual_Image_And_Captions_13000unsafe_harassment_image_captions
Dataset Card for "unsafe_harassment_image_captions"
More Information needed
imageCaptionsDatasafe_harassment_image_captions
Dataset Card for "safe_harassment_image_captions"
More Information needed
Pokemon_Card_Actual_Image_And_Captions_11000image_captions
