datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ramanv-image-captions-realramanv-image-captions-11learn_hf_food_not_food_image_captions
Food/Not Food Image Caption Dataset
Small dataset of synthetic food and not food image captions.
Text generated using Mistral Chat/Mixtral.
Can be used to train a text classifier on food/not_food image captions as a demo before scaling up to a larger dataset.
See Colab notebook on how dataset was created.
Example usage
import random
from datasets import load_dataset
# Load dataset
loaded_dataset = load_dataset("mrdbourke/learn_hf_food_not_food_image_captions")
# Get… See the full description on the dataset page: https://huggingface.co/datasets/mrdbourke/learn_hf_food_not_food_image_captions.ramanv-image-captions-11image_captions
From the Frontier Research Team at takara.ai we present over 1 million curated captioned images for multimodal text and image tasks.
Usage
from datasets import load_dataset
ds = load_dataset("takara-ai/image_captions")
print(ds)
Example
10,000 images from the dataset.
Methodology
We consolidated multiple open source datasets through an intensive 96-hour computational process across three nodes. This involved standardizing and validating the… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/image_captions.image-text-dataset-subset-300k-captions_onlyru-image-captions
Image Caprioning for Russian language
This dataset is a Russian part of dinhanhx/crossmodal-3600
Dataset Details
3.11k rows.
Two description for each picture. Cracked pictures were deleted from the original source.
The main feature is that all the descriptions are written by the native russian speakers.
Paper [https://google.github.io/crossmodal-3600/]
Uses
It is intended to be used for fine-tuning image captioning models.
astrobridge-image-captions
AstroBridge Legacy Survey Captions
3,487 imaging cutouts from the Legacy Survey (DR10 South + North), crossmatched against
published literature mentions and captioned in four independent stages by Gemini
(gemini-3.7-flash), following the AstroLLaVA data-generation approach (Zaman et al. 2025,
arXiv:2504.08583): no caption is ever told the object's
real name or catalog designation, and no caption states a fact that isn't derivable from the
pixels or the (redacted-at-the-model… See the full description on the dataset page: https://huggingface.co/datasets/gapatron/astrobridge-image-captions.enhanced_image_captionsimage_captions_x
Dataset Card for image_captions_x (URL + Caption)
This dataset provides a lightweight, web-scale resource of image-caption pairs in the form of URLs and their associated textual descriptions (captions). It is designed for training and evaluating vision-language models where users retrieve images independently from the provided links.
This dataset card is based on the Hugging Face dataset card template.
Dataset Details
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/image_captions_x.ImageCaptions-7M-Translations-Arabicimage-text-dataset-subset-300k-captions_only_with_latentsfood_tracker_image_captionsvg-captions-graphs-processed-image-graphsramanv-image-captions-2-pqramanv-image-captions-10-pqramanv-image-captions-3-pqramanv-image-captions-4-pqramanv-image-captions-6ramanv-image-captions-pqramanv-image-captions-7-pqramanv-image-captions-5-pqramanv-image-captions-8-pqramanv-image-captions-9-pqImageCaptions-7M-Translationsversion https://git-lfs.github.com/spec/v1
oid sha256:835f3f7d88a86e05a882c6a6b6333da6ab874776385f85473798769d767c2fca
size 27
simple-image-captionsramanv-image-captions-6-pqimage-captions-datasetdottrmstr-qwen-image-short-captions-no-style-datasethf_food_not_food_image_captions
