datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
newyorker_caption_contest
Dataset Card for New Yorker Caption Contest Benchmarks
Dataset Summary
See capcon.dev for more!
Data from:
Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest
@inproceedings{hessel2023androids,
title={Do Androids Laugh at Electric Sheep? {Humor} ``Understanding''
Benchmarks from {The New Yorker Caption Contest}},
author={Hessel, Jack and Marasovi{\'c}, Ana and Hwang, Jena D. and Lee, Lillian
and… See the full description on the dataset page: https://huggingface.co/datasets/jmhessel/newyorker_caption_contest.microvqaMicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research (CVPR 2025)
🌐 Homepage / blog •
📝 arXiv •
🤗 HF Dataset •
💻 Code •
🏛 CC-BY-SA-4.0
MicroVQA is expert-curated benchmark for multimodal reasoning for microscopy-based scientific research, proposed in the paper MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research.
Paper abstract
Scientific research demands sophisticated reasoning over multimodal… See the full description on the dataset page: https://huggingface.co/datasets/jmhb/microvqa.VidDiffBench
Dataset card for "VidDiffBench"
This is the dataset / benchmark for Video Action Differencing (ICLR 2025), a new task that compares how an action is performed between two videos. This page introduces the task, the dataset structure, and how to access the data. See the paper for details on dataset construction.
Blog / project page / leaderboards: https://jmhb0.github.io/viddiff/
Eval code and benchmarking popular LMMs: https://github.com/jmhb0/viddiff
Paper: arXiv:2503.07860… See the full description on the dataset page: https://huggingface.co/datasets/jmhb/VidDiffBench.seoul_subwayKorean Subway images, filmed from Seoul Station to Seoul National University.
