datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KMMMU
KMMMU (Korean MMMU)
technical report https://arxiv.org/abs/2604.13058
link to evaluation tutorial! https://github.com/HAE-RAE/KMMMU
KMMMU is a Korean version of MMMU: a multimodal benchmark designed to evaluate college-/exam-level reasoning that requires combining images + Korean text.
This dataset contains 3,466 questions collected from Korean exam sources including:
Civil service recruitment exams
National Technical Qualifications
National Competency Standard (NCS) exams… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMMU.HAERAE-VISION
HAERAE-VISION
A Korean visual QA benchmark featuring real-world, under-specified questions.
Dataset Description
This dataset includes two question types:
original: Under-specified, authentic user queries
explicit: Clarified queries with full context
Both share the same images and reference answers, allowing controlled evaluation of query under-specification.
Evaluation Code
See our GitHub repository for evaluation scripts.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAERAE-VISION.butterflies_and_moths_vqa
Butterflies and Moths VQA
Dataset Summary
butterflies_and_moths_vqa is a visual question answering (VQA) dataset focused on butterflies and moths. It features tasks such as fine-grained species classification and ecological reasoning. The dataset is designed to benchmark Vision-Language Models (VLMs) for both image-based and text-only training approaches.
Key Features
Fine-Grained Classification (Type1): Questions requiring detailed species identification.… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/butterflies_and_moths_vqa.
