datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
newyorker_caption_contest
Dataset Card for New Yorker Caption Contest Benchmarks
Dataset Summary
See capcon.dev for more!
Data from:
Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest
@inproceedings{hessel2023androids,
title={Do Androids Laugh at Electric Sheep? {Humor} ``Understanding''
Benchmarks from {The New Yorker Caption Contest}},
author={Hessel, Jack and Marasovi{\'c}, Ana and Hwang, Jena D. and Lee, Lillian
and… See the full description on the dataset page: https://huggingface.co/datasets/jmhessel/newyorker_caption_contest.newyorker_caption_ranking
New Yorker Caption Ranking Dataset
Dataset Descriptions
Homepage: https://nextml.github.io/caption-contest-data/
Repository: https://github.com/yguooo/cartoon-caption-generation
Paper: Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
Point of Contact: yguo@cs.wisc.edu
Dataset Summary
We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/newyorker_caption_ranking.coco-captions-pt-br
🎉 COCO Captions Dataset Translation for Portuguese Image Captioning
💾 Dataset Summary
COCO Captions Portuguese Translation, a multimodal dataset for Portuguese image captioning with 123,287 images, each accompanied by five descriptive captions that have been
generated by human annotators for every individual image. The original English captions were rendered into Portuguese
through the utilization of the Google Translator API.
🧑💻 Hot to Get… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/coco-captions-pt-br.astrobridge-image-captions
AstroBridge Legacy Survey Captions
3,487 imaging cutouts from the Legacy Survey (DR10 South + North), crossmatched against
published literature mentions and captioned in four independent stages by Gemini
(gemini-3.7-flash), following the AstroLLaVA data-generation approach (Zaman et al. 2025,
arXiv:2504.08583): no caption is ever told the object's
real name or catalog designation, and no caption states a fact that isn't derivable from the
pixels or the (redacted-at-the-model… See the full description on the dataset page: https://huggingface.co/datasets/gapatron/astrobridge-image-captions.pokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions
You can use it in LLaMA Factory by specifying dataset: pokemon_cap.
leetcode_with_youtube_captionscoco-captions-pt-br
🎉 COCO Captions Dataset Translation for Portuguese Image Captioning
💾 Dataset Summary
COCO Captions Portuguese Translation, a multimodal dataset for Portuguese image captioning with 123,287 images, each accompanied by five descriptive captions that have been
generated by human annotators for every individual image. The original English captions were rendered into Portuguese
through the utilization of the Google Translator API.
🧑💻 Hot to Get… See the full description on the dataset page: https://huggingface.co/datasets/EliMC/coco-captions-pt-br.JourneyBench_Captioning
