jmh
Datasets
All datasets matching “jmh”newyorker_caption_contest
Dataset Card for New Yorker Caption Contest Benchmarks
Dataset Summary
See capcon.dev for more!
Data from:
Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest
@inproceedings{hessel2023androids,
title={Do Androids Laugh at Electric Sheep? {Humor} ``Understanding''
Benchmarks from {The New Yorker Caption Contest}},
author={Hessel, Jack and Marasovi{\'c}, Ana and Hwang, Jena D. and Lee, Lillian
and… See the full description on the dataset page: https://huggingface.co/datasets/jmhessel/newyorker_caption_contest.pubmed_bioasq_2022
PubMed BioASQ 2022 Corpus
This dataset contains the PubMed abstracts corpus from the BioASQ 2022 challenge, comprising approximately 23 million biomedical documents with MeSH (Medical Subject Headings) annotations.
Purpose
This is a convenience collection of the PubMed corpus from the BioASQ 2022 Challenge, reformatted for easier use in retrieval and QA systems. The original source is the BioASQ challenge data. We created this processed version with multiple formats (JSON… See the full description on the dataset page: https://huggingface.co/datasets/jmhb/pubmed_bioasq_2022.mmc4-core-ffThe core fewer faces subset of mmc4
# images
# docs
# tokens
Multimodal-C4 core fewer-faces (mmc4-core-ff)
22.4M
5.5M
1.8B
PaperSearchQA
PaperSearchQA Dataset
A large-scale biomedical question-answering dataset for training and evaluating search agents that reason over scientific literature.
Dataset Description
PaperSearchQA contains 60,000 question-answer pairs generated from PubMed abstracts, designed for training retrieval-augmented language models on biomedical question answering tasks.
Dataset Statistics
Training set: 54,907 Q&A pairs
Test set: 5,000 Q&A pairs
Retrieval corpus: 16 million… See the full description on the dataset page: https://huggingface.co/datasets/jmhb/PaperSearchQA.microvqaMicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research (CVPR 2025)
🌐 Homepage / blog •
📝 arXiv •
🤗 HF Dataset •
💻 Code •
🏛 CC-BY-SA-4.0
MicroVQA is expert-curated benchmark for multimodal reasoning for microscopy-based scientific research, proposed in the paper MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research.
Paper abstract
Scientific research demands sophisticated reasoning over multimodal… See the full description on the dataset page: https://huggingface.co/datasets/jmhb/microvqa.bioasq_yesno_trainv0_n1464_test100
