datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya-mm-exams-spanish-medicalMedical Spanish Exams for the Multimodal Aya Exams Projects.
Questions available in file: data.json
Images stored in: /images
Original data and file available here: link
taiwan-examsMachine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-exams.EXAMS-V
EXAMS-V: ImageCLEF 2025 – Multimodal Reasoning
Dimitar Iliyanov Dimitrov, Hee Ming Shan, Zhuohan Xie, Rocktim Jyoti Das , Momina Ahsan, Sarfraz Ahmad, Nikolay Paev, Ali Mekky, Omar El Herraoui, Rania Hossam, Nurdaulet Mukhituly, Akhmed Sakip, Ivan Koychev, Preslav Nakov
INTRODUCTION
EXAMS-V is a multilingual, multimodal dataset created to evaluate and benchmark the visual reasoning abilities of AI systems, especially Vision-Language Models (VLMs). The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/EXAMS-V.aya-mm-exams-spanish-nursingNursing Spanish Exams for the Multimodal Aya Exams Projects.
Questions available in file: data.json
Images stored in: /images
Original data and file available here: link
taiwan-professional-exams-115-2Machine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/skyhong2002/taiwan-professional-exams-115-2.EXAMS-V
EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models
Rocktim Jyoti Das, Simeon Emilov Hristov, Haonan Li, Dimitar Iliyanov Dimitrov, Ivan Koychev, Preslav Nakov
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi & Sofia University
This is arxiv link for the EXAMS-V paper can be found here.
Introduction
We introduce EXAMS-V, a new challenging multi-discipline multimodal multilingual exam benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/Rocktim/EXAMS-V.Nigerian-Language-Exams-VQA
Nigerian Language Exams VQA Dataset
Dataset Summary
The Nigerian Language Exams VQA Dataset is a curated benchmark designed to evaluate the Document Visual Question Answering (VQA) capabilities of Multimodal Large Language Models (MLLMs) in indigenous West African languages. It consists of 837 manually cropped images of multiple-choice questions from standardized Nigerian examinations (WAEC and NECO) spanning the years 2008 to 2024.
This dataset addresses the critical… See the full description on the dataset page: https://huggingface.co/datasets/LyngualLabs/Nigerian-Language-Exams-VQA.taiwan-national-examsMachine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), assets/ (figures), and, when present, resources/ (the public
solver corpus). The entire resource… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-national-exams.flemish_multimodal_exams_physicianindian-entrance-exams-benchmarkEXAMS-Vchemistry-multimodal-exams-hindiEXAMS-V_IT
Dataset Card for EXAMS-V_IT
Dataset description
This is a formatted version of EXAMS-V including only the Italian data.
Citation
If you use this dataset in your research you should cite the original publication:
@misc{das2024examsv,
title={EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models},
author={Rocktim Jyoti Das and Simeon Emilov Hristov and Haonan Li and Dimitar Iliyanov Dimitrov and Ivan… See the full description on the dataset page: https://huggingface.co/datasets/swap-uniba/EXAMS-V_IT.taiwan-examsMachine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/TakalaWang/taiwan-exams.dots-ocr-test-exams
Document OCR using dots.ocr
This dataset contains OCR results from images in NationalLibraryOfScotland/Scottish-School-Exam-Papers using DoTS.ocr, a compact 1.7B multilingual model.
Processing Details
Source Dataset: NationalLibraryOfScotland/Scottish-School-Exam-Papers
Model: rednote-hilab/dots.ocr
Number of Samples: 10
Processing Time: 1.6 min
Processing Date: 2025-10-07 14:23 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/dots-ocr-test-exams.Brazil-Medicine-Schools-Entrance-Exams-FAMERP-SANTA-CASAEXAMS-V
EXAMS-V: ImageCLEF 2025 – Multimodal Reasoning
Dimitar Iliyanov Dimitrov, Hee Ming Shan, Zhuohan Xie, Rocktim Jyoti Das , Momina Ahsan, Sarfraz Ahmad, Nikolay Paev, Ali Mekky, Omar El Herraoui, Rania Hossam, Nurdaulet Mukhituly, Akhmed Sakip, Ivan Koychev, Preslav Nakov
INTRODUCTION
EXAMS-V is a multilingual, multimodal dataset created to evaluate and benchmark the visual reasoning abilities of AI systems, especially Vision-Language Models (VLMs). The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/EXAMS-V.exams-ocr-optimized-testchemistry-multimodal-exams-teluguEXAMS-V-cnenexams-ocrmultimodal-PL-examsEXAMS-V
EXAMS-V: ImageCLEF 2025 – Multimodal Reasoning
Dimitar Iliyanov Dimitrov, Hee Ming Shan, Zhuohan Xie, Rocktim Jyoti Das , Momina Ahsan, Sarfraz Ahmad, Nikolay Paev, Ali Mekky, Omar El Herraoui, Rania Hossam, Nurdaulet Mukhituly, Akhmed Sakip, Ivan Koychev, Preslav Nakov
INTRODUCTION
EXAMS-V is a multilingual, multimodal dataset created to evaluate and benchmark the visual reasoning abilities of AI systems, especially Vision-Language Models (VLMs). The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/dheenab/EXAMS-V.arabic_examsvmultimodal-SK-examsexams-ocr-gpu-testEXAMS-V
EXAMS-V: ImageCLEF 2025 – Multimodal Reasoning
Dimitar Iliyanov Dimitrov, Hee Ming Shan, Zhuohan Xie, Rocktim Jyoti Das , Momina Ahsan, Sarfraz Ahmad, Nikolay Paev, Ali Mekky, Omar El Herraoui, Rania Hossam, Nurdaulet Mukhituly, Akhmed Sakip, Ivan Koychev, Preslav Nakov
INTRODUCTION
EXAMS-V is a multilingual, multimodal dataset created to evaluate and benchmark the visual reasoning abilities of AI systems, especially Vision-Language Models (VLMs). The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AycanSon/EXAMS-V.multimodal-CS-examsexams-hunyuan-ocr
Document OCR using HunyuanOCR
This dataset contains OCR results from images in NationalLibraryOfScotland/Scottish-School-Exam-Papers using HunyuanOCR, a lightweight 1B VLM from Tencent.
Processing Details
Source Dataset: NationalLibraryOfScotland/Scottish-School-Exam-Papers
Model: tencent/HunyuanOCR
Number of Samples: 100
Processing Time: 9.8 min
Processing Date: 2025-11-25 16:15 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/exams-hunyuan-ocr.bengali-exams-public
