datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image_action_only_ballimage-text-dataset-subset-300k-captions_onlyimage-text-dataset-subset-300k-captions_only_with_latentstest_image_onlyrelaion_with_image_make_onlyvoc2012-image-onlycrisismmd_inf_image_onlyCrisismmd Informative Task Labels (image-only):
0: not_informative
1: informative
Category
Train
Dev
Test
Total
Informative
6,345
1,056
1,030
8,431
Not-informative
3,256
517
504
4,277
Total
9,601
1,573
1,534
12,708
JSWEMU_add_image_onlyMMLongBench_image_question_only_deepeyes_prompt_revised_wo_imageMMLongBench_image_question_only_deepeyes_multi_viz_2_1120_newcrisismmd_hum_image_onlyCrisismmd Humanitarian Task Labels (image-only):
0: affected_individuals
1: rescue_volunteering_or_donation_effort
2: infrastructure_and_utility_damage
3: other_relevant_information
4: not_humanitarian
Category
Train
Dev
Test
Total
Affected individuals
71
9
9
89
Rescue/Volunteering
912
149
126
1,187
Infrastructure damage
612
80
81
773
Other relevant
1,279
239
235
1,753
Not-humanitarian
3,252
521
504
4,277
Total
6,126
998
955
8,079
only_image_gpt-5-nano_wtMMLongBench_image_question_only_deepeyes_prompt_revisedgpt_gen_desc_image_only_logosMMLongBench_image_question_only_deepeyes_concat_viz_concatclevr1000_hf_image_text_rephrased_onlyimage_only_test-datasetthought_action_bid_only_version_3_no_imagedermoscopy-releasev0-image-report-onlyonly_image_qwen3-vl8b-instructonly_image_gpt-5-nanoimage_only_train-datasetonly_image_gpt-5-nano_simpleonly_image_qwen3-vl8bMMLongBench_image_question_only_deepeyes_concatonly_image_gpt5-nanoonly_image_qwen3-vl8b-instruct_simpleonly_image_qwen3-vl8b_simpleMMLongBench_image_question_only_deepeyes_multi_viz_2MMLongBench_image_question_only_deepeyes_multi_seed5
MMLongBench – 2025-12-12 09:08 UTC
Average accuracy: 39.35% (1052 samples with scores)
Subset metrics by evidence source:
Figure: samples=296, accuracy=35.14%
Pure-text (Plain-text): samples=296, accuracy=44.26%
Table: samples=217, accuracy=41.94%
Chart: samples=173, accuracy=38.73%
Generalized-text (Layout): samples=117, accuracy=27.35%
Subset metrics by evidence pages length:
no_pages: samples=214, accuracy=27.57%
single_page: samples=485, accuracy=53.20%
multiple_pages:… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/MMLongBench_image_question_only_deepeyes_multi_seed5.
