datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
exams
Dataset Card for [Dataset Name]
Dataset Summary
EXAMS is a benchmark dataset for multilingual and cross-lingual question answering from high school examinations. It consists of more than 24,000 high-quality high school exam questions in 16 languages, covering 8 language families and 24 school subjects from Natural Sciences and Social Sciences, among others.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The languages in the… See the full description on the dataset page: https://huggingface.co/datasets/mhardalov/exams.taiwan-examsMachine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-exams.Arabic_EXAMSoab_examsEXAMS-V
EXAMS-V: ImageCLEF 2025 – Multimodal Reasoning
Dimitar Iliyanov Dimitrov, Hee Ming Shan, Zhuohan Xie, Rocktim Jyoti Das , Momina Ahsan, Sarfraz Ahmad, Nikolay Paev, Ali Mekky, Omar El Herraoui, Rania Hossam, Nurdaulet Mukhituly, Akhmed Sakip, Ivan Koychev, Preslav Nakov
INTRODUCTION
EXAMS-V is a multilingual, multimodal dataset created to evaluate and benchmark the visual reasoning abilities of AI systems, especially Vision-Language Models (VLMs). The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/EXAMS-V.exams-basic-and-quantum-cryptography-and-security-latex
Open Problem Exams: Cryptography and Security (LaTeX)
A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions.
Dataset Overview
Institution
Files
Topics
Questions
Caltech & TU Delft
8
38
145
EPFL
6
19
86
ETH Zurich
1
14
37
MIT
3
33
79
Total
18
104
347
Difficulty Distribution
Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.medical-exams-LDEK-EN-2013-2024
Dataset Card for medical-exams-LDEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LDEK-EN-2013-2024.swedish-medical-exams-mcq-1006-json
Dataset Card for Swedish Medical Exam MCQs
Dataset Description
This dataset contains multiple-choice questions from Swedish medical exams.
Languages
The dataset is in Swedish (sv).
Dataset Structure
Each entry in the dataset contains the following fields:
question: The question
options: An array of possible answers
answer: The correct answer
language: The language of the question (always "sv" for Swedish)
country: The country of origin (always… See the full description on the dataset page: https://huggingface.co/datasets/sarafuyu/swedish-medical-exams-mcq-1006-json.Arabic_EXAMS-Redux
Arabic_EXAMS-Redux
A corrected and text-repaired version of OALL/Arabic_EXAMS, the Arabic subset of the EXAMS multilingual high-school examinations benchmark.
What was fixed
Repaired corrupted Arabic text. The upstream benchmark contains widespread PDF-extraction damage to question stems and answer choices: split diacritics, fragmented words, and non-Arabic glyphs replacing standard characters. We restored these to readable Modern Standard Arabic.
Corrected the answer… See the full description on the dataset page: https://huggingface.co/datasets/inceptlabs/Arabic_EXAMS-Redux.aya-global-exams-catalanCatalan exams for the Aya Global Exams.
Original data and file available here: link
Github Repo: link
estonian_language_exams
Dataset Card for Estonian Language Proficiency Exam Samples
This is part of the initiative from Cohere For AI @CohereForAI to gather exams from around the world to build a new multilingual benchmark.
The web Scrapping code can be found at the source_scripts_data_aya repository.
The source data can be visually checked at
Sõeltestid.pdf and
Diagnoostestid.pdf.
Dataset Details
Dataset Description
This dataset contains sample questions from the Estonian language… See the full description on the dataset page: https://huggingface.co/datasets/Gabrui/estonian_language_exams.taiwan-professional-exams-115-2Machine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/skyhong2002/taiwan-professional-exams-115-2.taiwan-exams-resultsBenchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-<entry>/entry.json — that configuration's per-paper scores
results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-exams-results.EXAMS-V
EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models
Rocktim Jyoti Das, Simeon Emilov Hristov, Haonan Li, Dimitar Iliyanov Dimitrov, Ivan Koychev, Preslav Nakov
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi & Sofia University
This is arxiv link for the EXAMS-V paper can be found here.
Introduction
We introduce EXAMS-V, a new challenging multi-discipline multimodal multilingual exam benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/Rocktim/EXAMS-V.taiwan-national-exams-resultsBenchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-<entry>/entry.json — that configuration's per-paper scores
results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-national-exams-results.pt_exams
PHEB - Portuguese High School Exams MCQ
MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum.
For more details, see the PHEB paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.thema-panhellenic-exams
Thema
Thema is an open benchmark based on the Greek Panhellenic university entrance examinations (Πανελλαδικές Εξετάσεις ΓΕΛ). Models answer authentic exam questions in Greek. Responses are graded on the national 0 to 20 scale and can be converted to the admission points used by Greek university departments.
Website · Code · Leaderboard · Method
Dataset summary
Current release
Exam years
2023 to 2026
Published subjects
10 of 10
Complete exam… See the full description on the dataset page: https://huggingface.co/datasets/mpvasilis/thema-panhellenic-exams.examsmedical-exams-LEK-EN-2013-2024
Dataset Card for medical-exams-LEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LEK-EN-2013-2024.taiwan-professional-exams-115-2-resultsBenchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-<entry>/entry.json — that configuration's per-paper scores
results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/skyhong2002/taiwan-professional-exams-115-2-results.cambridge_exams_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://ilexir.co.uk/datasets/index.html
Original Dataset Paper:
Menglin Xia, Ekaterina Kochmar, and Ted Briscoe. 2016. Text Readability… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cambridge_exams_en.medical-exams-LEK-PL-2008-2024
Dataset Card for medical-exams-LEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LEK-PL-2008-2024.thai-investment-consultant-licensing-exams
Thai Public Investment Consultant (IC) Exams Dataset
Overview
This dataset comprises a collection of exam questions and answers from the Thai Public Investment Consultant (IC) Examinations. It's a valuable resource for developing and evaluating question-answering systems in the finance sector.
Dataset Source
The Stock Exchange of Thailand (SET)
Maintainer
Dr. Kobkrit Viriyayudhakorn
Email: kobkrit@iapp.co.th
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-investment-consultant-licensing-exams.taiwan-national-examsMachine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), assets/ (figures), and, when present, resources/ (the public
solver corpus). The entire resource… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-national-exams.examsReassess-Polish-Medical-Examstaiwan-exams-resultsBenchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-<entry>/entry.json — that configuration's per-paper scores
results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/TakalaWang/taiwan-exams-results.indian-entrance-exams-benchmarkUP_CET_Hindi_exams920_chemistry_exams_dataset_seed_46
