datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sounio-code-examples
Sounio Curated Code Examples
Curated compile-clean .sio examples for training and evaluating code models on
Sounio, a self-hosted systems and scientific programming language for epistemic
computing, uncertainty propagation, and algebraic effects.
This directory is the Cx-1 expansion lane for
chiuratto-AIgourakis/sounio-code-examples.
Current batch
Examples: 5,000
Metadata files: 5,000
Compiler gate: bin/souc check pass rate 5,000/5,000
Utility layer: 5,000… See the full description on the dataset page: https://huggingface.co/datasets/chiuratto-AIgourakis/sounio-code-examples.doc-formats-jsonl-1
[doc] formats - jsonl - 1
This dataset contains one jsonl file at the root.
chat_formatted_examplesthai_exam
Dataset Card for Thai_Exam
ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows:
ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.MedXpertQA-Exam
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/MedXpertQA-Exam.opengloss-v1.3-query-examples-flat
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.chinese_exam_train_dataexams-basic-and-quantum-cryptography-and-security-latex
Open Problem Exams: Cryptography and Security (LaTeX)
A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions.
Dataset Overview
Institution
Files
Topics
Questions
Caltech & TU Delft
8
38
145
EPFL
6
19
86
ETH Zurich
1
14
37
MIT
3
33
79
Total
18
104
347
Difficulty Distribution
Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.thai_exam-reformattedReformatted version of scb10x/thai_exam
Additional Changes:
Fix math incorrect answer
ถ้า \log_{\frac{1}{4}} 256 + \frac{2\log 625}{\log 5} = 3^a เมื่อ a เป็นจำนวนจริง แล้วคำตอบของ a เท่ากับเท่าใด?
a. \log_{3} 2
b. \log_{3} 4
c. \log_{3} \frac{33}{4}
d. \log_{3} 10
e. \log_{3} 12
# Original answer: d (\log_{3} 10)
# Correction : b (\log_{3} 4)
thai-exam-seacrowdmedical-exams-LDEK-EN-2013-2024
Dataset Card for medical-exams-LDEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LDEK-EN-2013-2024.swedish-medical-exams-mcq-1006-json
Dataset Card for Swedish Medical Exam MCQs
Dataset Description
This dataset contains multiple-choice questions from Swedish medical exams.
Languages
The dataset is in Swedish (sv).
Dataset Structure
Each entry in the dataset contains the following fields:
question: The question
options: An array of possible answers
answer: The correct answer
language: The language of the question (always "sv" for Swedish)
country: The country of origin (always… See the full description on the dataset page: https://huggingface.co/datasets/sarafuyu/swedish-medical-exams-mcq-1006-json.ENADE_Brazilian_national_university_examination_MCQ_483estonian_language_exams
Dataset Card for Estonian Language Proficiency Exam Samples
This is part of the initiative from Cohere For AI @CohereForAI to gather exams from around the world to build a new multilingual benchmark.
The web Scrapping code can be found at the source_scripts_data_aya repository.
The source data can be visually checked at
Sõeltestid.pdf and
Diagnoostestid.pdf.
Dataset Details
Dataset Description
This dataset contains sample questions from the Estonian language… See the full description on the dataset page: https://huggingface.co/datasets/Gabrui/estonian_language_exams.IPA_Exam_AM
独立行政法人 情報処理推進機構(IPA) 情報処理技術者試験 試験問題データセット(午前)
概要
本データセットは、独立行政法人 情報処理推進機構(以下、IPA)の情報処理技術者試験の午前問題及び公式解答をセットにした、非公式のデータセットです。
以下、公開されている過去問題pdfから抽出しています。
https://www.ipa.go.jp/shiken/mondai-kaiotu/index.html
人間の学習目的での利用の他、LLMのベンチマークやファインチューニング等、生成AIの研究開発用途での利用を想定しています。
データの詳細
現状、過去5年分(2020〜2024)の以下試験区分について、問題及び回答を収録しています。
応用情報技術者試験(ap)
高度共通午前1(koudo)
エンベデッドシステムスペシャリスト(es)
情報処理安全確保支援士(sc)
プロジェクトマネージャ(pm)
データベーススペシャリスト(db)
システム監査技術者(au)
ITストラテジスト(st)… See the full description on the dataset page: https://huggingface.co/datasets/Dimeiza/IPA_Exam_AM.OpenSeek-Synthetic-Reasoning-Data-Examples
OpenSeek-Reasoning-Data
OpenSeek [Github|Blog]
Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process.
News
🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.exambench
ExamBench
The ExamBench Dataset (~600M tokens, 405k examples) is one of the largest open-source corpora designed for competitive exam preparation and reasoning AI. Generated using advanced distillation techniques, it combines structured chain-of-thought reasoning with comprehensive coverage of over 25 Indian and international examinations. From JEE and NEET to UPSC, Banking, GRE, and IELTS, the dataset spans multiple domains like STEM, humanities, current affairs, language, and… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/exambench.example-space-to-dataset-jsonDemo to save data from a Space to a Dataset. Goal is to provide reusable snippets of code.
Documentation: https://huggingface.co/docs/huggingface_hub/main/en/guides/upload#scheduled-uploads
Space: https://huggingface.co/spaces/Wauplin/space_to_dataset_saver/
JSON dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json
Image dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image
Image (zipped) dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json.sycophancy_examples
Sycophancy Examples
Two sycophancy evaluation datasets from Kei et al., "Reward hacking can generalise across settings".
Original source: GeodesicResearch/Obfuscation_Generalization
Files
File
Examples
Description
sycophancy_opinion_political.jsonl
5,000
Political opinion questions with persona-aligned "sycophantic" answers
sycophancy_fact.jsonl
401
Factual questions where the persona holds a misconception; sycophantic answer agrees with the misconception… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/sycophancy_examples.multimodal-example
Multimodal Example Dataset
Small example dataset for testing multimodal (vision-language) fine-tuning with ms-swift.
Structure
├── train.jsonl # 10 training samples
├── test.jsonl # 2 validation samples
├── images/ # All referenced images (400x300 JPEG)
│ ├── dog_portrait.jpg
│ ├── forest_river.jpg
│ ├── laptop_desk.jpg
│ ├── mountain_lake.jpg
│ ├── ocean_rocks.jpg
│ ├── coffee_cup.jpg
│ ├── bookshelf.jpg
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/f13rnd/multimodal-example.2023_Pharmacist_Licensure_Examination-TCM_trackThe 2023 Chinese National Pharmacist Licensure Examination is divided into two distinct tracks: the Pharmacy track and the Traditional Chinese Medicine (TCM) Pharmacy track. The data provided here pertains to the Traditional Chinese Medicine (TCM) Pharmacy track examination. It is important to note that this dataset was collected from online sources, and there may be some discrepancies between this data and the actual examination.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/2023_Pharmacist_Licensure_Examination-TCM_track.pt_exams
PHEB - Portuguese High School Exams MCQ
MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum.
For more details, see the PHEB paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.ua-judge-exam
UA-JudgeExam
11,990 four-option multiple-choice items with official answer keys, from the question
bank the Higher Qualification Commission of Judges of Ukraine
(Вища кваліфікаційна комісія суддів України, VKKS) publishes for the anonymous written
testing of candidates for appellate-court judgeships.
Source: Commission decision of 15 July 2024, No. 221/зп-24 — five documents, 1,672 PDF pages,
covering general legal knowledge plus the administrative, commercial, criminal and… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ua-judge-exam.hydro_cali_agent_exampleflutter-full-examples-v1
Flutter Codegen: Full Examples
Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal
that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps,
there's no step history or diff structure here -- each row is a single, standalone
goal -> complete file example.
This is the whole-code counterpart to flutter-diff-steps-v1, intended for
training/evaluating a baseline that generates the entire file in one shot, to
compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.thema-panhellenic-exams
Thema
Thema is an open benchmark based on the Greek Panhellenic university entrance examinations (Πανελλαδικές Εξετάσεις ΓΕΛ). Models answer authentic exam questions in Greek. Responses are graded on the national 0 to 20 scale and can be converted to the admission points used by Greek university departments.
Website · Code · Leaderboard · Method
Dataset summary
Current release
Exam years
2023 to 2026
Published subjects
10 of 10
Complete exam… See the full description on the dataset page: https://huggingface.co/datasets/mpvasilis/thema-panhellenic-exams.marketing_exam1medical-exams-LEK-EN-2013-2024
Dataset Card for medical-exams-LEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LEK-EN-2013-2024.AIO2025-Final-Exam-Datasetversion https://git-lfs.github.com/spec/v1
oid sha256:072d081dec61aad02275a7ed608c136a28f1de6672c7ff3a520bdbf7b0952e53
size 2225
cambridge_exams_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://ilexir.co.uk/datasets/index.html
Original Dataset Paper:
Menglin Xia, Ekaterina Kochmar, and Ted Briscoe. 2016. Text Readability… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cambridge_exams_en.
