datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
COCO-Mini-DeepCaption-10K
COCO-Mini-DeepCaption-10K
COCO-Mini-DeepCaption-10K is a dense image captioning dataset built from a 10,000-image subset of the COCO dataset, paired with long-form synthetic captions generated using the Qwen3.5 multimodal model. Each caption is produced through a dedicated Qwen3.5 captioning pipeline designed to yield detailed, high-fidelity descriptions of scene composition, subject attributes, and visual context rather than short, generic labels. The dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/COCO-Mini-DeepCaption-10K.DeepCaption-1K-Qwen35
DeepCaption-1K-Qwen35
A curated image captioning dataset containing 1,000 image and caption pairs. Each sample consists of an image paired with a detailed natural language description, making it suitable for training and evaluating image captioning and vision-language models.
Features
1,000 image-caption pairs
Detailed natural language captions
High-quality descriptive annotations
Suitable for image captioning and multimodal learning
Data Format
{… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/DeepCaption-1K-Qwen35.
