datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
exams-basic-and-quantum-cryptography-and-security-latex
Open Problem Exams: Cryptography and Security (LaTeX)
A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions.
Dataset Overview
Institution
Files
Topics
Questions
Caltech & TU Delft
8
38
145
EPFL
6
19
86
ETH Zurich
1
14
37
MIT
3
33
79
Total
18
104
347
Difficulty Distribution
Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.medical-exams-LDEK-EN-2013-2024
Dataset Card for medical-exams-LDEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LDEK-EN-2013-2024.swedish-medical-exams-mcq-1006-json
Dataset Card for Swedish Medical Exam MCQs
Dataset Description
This dataset contains multiple-choice questions from Swedish medical exams.
Languages
The dataset is in Swedish (sv).
Dataset Structure
Each entry in the dataset contains the following fields:
question: The question
options: An array of possible answers
answer: The correct answer
language: The language of the question (always "sv" for Swedish)
country: The country of origin (always… See the full description on the dataset page: https://huggingface.co/datasets/sarafuyu/swedish-medical-exams-mcq-1006-json.estonian_language_exams
Dataset Card for Estonian Language Proficiency Exam Samples
This is part of the initiative from Cohere For AI @CohereForAI to gather exams from around the world to build a new multilingual benchmark.
The web Scrapping code can be found at the source_scripts_data_aya repository.
The source data can be visually checked at
Sõeltestid.pdf and
Diagnoostestid.pdf.
Dataset Details
Dataset Description
This dataset contains sample questions from the Estonian language… See the full description on the dataset page: https://huggingface.co/datasets/Gabrui/estonian_language_exams.pt_exams
PHEB - Portuguese High School Exams MCQ
MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum.
For more details, see the PHEB paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.thema-panhellenic-exams
Thema
Thema is an open benchmark based on the Greek Panhellenic university entrance examinations (Πανελλαδικές Εξετάσεις ΓΕΛ). Models answer authentic exam questions in Greek. Responses are graded on the national 0 to 20 scale and can be converted to the admission points used by Greek university departments.
Website · Code · Leaderboard · Method
Dataset summary
Current release
Exam years
2023 to 2026
Published subjects
10 of 10
Complete exam… See the full description on the dataset page: https://huggingface.co/datasets/mpvasilis/thema-panhellenic-exams.medical-exams-LEK-EN-2013-2024
Dataset Card for medical-exams-LEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LEK-EN-2013-2024.cambridge_exams_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://ilexir.co.uk/datasets/index.html
Original Dataset Paper:
Menglin Xia, Ekaterina Kochmar, and Ted Briscoe. 2016. Text Readability… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cambridge_exams_en.medical-exams-LEK-PL-2008-2024
Dataset Card for medical-exams-LEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LEK-PL-2008-2024.UP_CET_Hindi_examsmedical-exams-PES-PL-2007-2024
Dataset Card for medical-exams-PES-PL-2007-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-PES-PL-2007-2024.filipino_elementary_periodical_examsuniv_exams_finnishgpt-exams
GPT-exams
Dataset summary
The dataset contains 8131 multi-domain question-answer pairs. It was created semi-automatically using the gpt-3.5-turbo-0613 model available in the OpenAI API. The process of building the dataset was as follows:
We manually prepared a list of 409 university-level courses from various fields. For each course, we instructed the model with the prompt: "Wygeneruj 20 przykładowych pytań na egzamin z [nazwa przedmiotu]" (Generate 20 sample questions… See the full description on the dataset page: https://huggingface.co/datasets/sdadas/gpt-exams.JEE_Main_Hindi_Examsexam-slovak-mathbioPart of INCLUDE
see research: https://arxiv.org/abs/2411.19799
full dataset: https://huggingface.co/datasets/CohereForAI/include-base-44
medical-exams-LDEK-PL-2008-2024
Dataset Card for medical-exams-LDEK-PL-2008-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LDEK-PL-2008-2024.EAMCET-Urdu-Examspolsci-exams-mcqfilipino_morals_culture_elementary_periodical_examsamalia-iave-exams-2024-2025
AMALIA IAVE national exams — verified MCQ pairs (2024-2025)
Verified multiple-choice question+answer pairs extracted from Portugal's
2024 and 2025 national secondary-school exams (Ensino Secundário, 12th
grade), built for specializing
AMALIA-9B-0626-DPO
toward a K-12 tutor use case. Ground truth by construction: every answer is
read directly off the official IAVE marking scheme (critérios de correção), never inferred by a model. Full pipeline, methodology, and the
rest of the… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-iave-exams-2024-2025.philippine_history_elementary_periodical_examsExams-MCQ-seed-Vie
Đây là dataset Exams MCQ Vietnames.
Được dùng để train / fine-tune model dành cho benchmark vlsp-2023-vllm/exams_vi
Dataset được tổng hợp lại từ 2 dataset roshansk23/Vietnam_HighSchool_Exam_Dataset và hllj/vi_grade_school_math_mcq
korean_land_mgmt_law_examsjapanese-criminal-law-exams
Japanese Criminal Law Multiple-Choice Exam Dataset
Overview
This dataset contains multiple-choice questions from Japanese criminal law (刑法) exams published by the Japanese Ministry of Justice (法務省). The dataset includes 100 questions from 5 different years of exams.
Data Source
All questions are collected from the official Japanese Ministry of Justice website:
https://www.moj.go.jp/jinji/shihoushiken/
Dataset Structure
The dataset is provided in JSON… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/japanese-criminal-law-exams.bengali-exams-publicexamsAlpaca-Medical-Reports-Exams-Pt-Datasetcie_examsamharic_primary_school_exams
