datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
copycolors_mcqaThis dataset consists of formatted n-way multiple choice questions, where n is in [2,10]. The task itself is simply to copy the prototypical color from the context and produce the corresponding color's answer choice letter.
The "prototypical colors" dataset instances themselves come from Memory Colors (Norland et al. 2021) and corypaik/coda (instances whose object_group is 0, indicating participants agreed on a prototypical color of that object).
alba_mcq
ALBA MCQ
Multiple Choice Version of the ALBA benchmark, a Portuguese language benchmark for proficiency in pt-PT linguistic-related tasks.
For more details, see the ALBA paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/alba_mcq.git-history-mcq-ru
git-history-mcq-ru
805 вопросов с вариантами ответа по истории трёх открытых репозиториев
(digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс
8 672 ответа пяти моделей и 4 878 разборов этих ответов.
Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда
и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю.
Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять
вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.copycolors_mcqaThis dataset consists of a formatted version of the Memory Colors dataset (Norland et al. 2021), formatted for n-way multiple choice where n is in [2,11].
aplikacje-prawnicze-mcq
Polish Legal Apprenticeship Entrance Exams — adwokacka/radcowska, notarialna, komornicza (2007–2025)
Native-Polish, single-choice (A/B/C) legal MCQ benchmark built from the official entrance
examinations for the Polish legal apprenticeships, published by the Ministry of Justice:
adwokacka + radcowska (advocate + legal counsel — a single shared test from 2009 on;
two separate exams in 2007),
notarialna (notary),
komornicza (court-enforcement officer / bailiff).
Each item… See the full description on the dataset page: https://huggingface.co/datasets/bartoszkobylinski1/aplikacje-prawnicze-mcq.arc-synth-mcqa
ARC-Style Synthetic Science MCQs (teacher: Qwen3.8-27B)
Synthetic multiple-choice science questions generated for ARC-Challenge
fine-tuning, released for reproducibility of the companion model. Every file
that was used in training is included, along with the full audit trail.
Files
file
rows
what it is
clean_all.jsonl
6,861
v2 pool: generated, blind-label-verified, deduped, ARC-form-gated
clean_std.jsonl
4,630
non-negation subset of the above… See the full description on the dataset page: https://huggingface.co/datasets/minjujeon/arc-synth-mcqa.BlackSwanSuite-MCQ
Black Swan Suite
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events Aditya Chinchure*, Sahithya Ravi*, Raymond Ng, Vered Shwartz, Boyang Li, Leonid Sigal (* equal) 🎉 Accepted at CVPR 2025 arXiv | Website
Dataset Information
Black Swan has three variants of questions. Please find the data in the appropriate repositories:
BlackSwanSuite-Gen -- link
BlackSwanSuite-MCQ (this)
BlackSwanSuite-YN -- link
This dataset contains questions for MCQ… See the full description on the dataset page: https://huggingface.co/datasets/UBC-ViL/BlackSwanSuite-MCQ.longmemeval-pooled-mcq
LongMemEval Pooled MCQ — context/question sets for KV-cache (cartridge) training
Conversational experiences (context to distill into a compact KV representation) paired with
exact-match 4-option MCQ validation questions, derived from the distractor sessions of
LongMemEval (longmemeval_m, cleaned release). No LLM judge
needed: answers are single tokens, graded by letter match.
config
records
contents
pooled3_experiences
88
conversations (~245k tok total), 3 questions… See the full description on the dataset page: https://huggingface.co/datasets/hhy13/longmemeval-pooled-mcq.NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET
Nepali Devanagari SFT Dataset — Final Clean Release
A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.
Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.
Dataset at a Glance
Property
Value
Total rows
100,000
Total conversation messages
200,000
Human messages
100,000
GPT messages
100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.polsci-exams-mcqmcq_test_2TempoMed-Bench-MCQ
TempoMed-Bench-MCQ
Data Overview
TempoMed-Bench-MCQ is a multiple-choice question benchmark designed to evaluate temporal awareness in medical large language models. Each instance is constructed from a pair of medical guidelines: an up-to-date guideline and an oudated guideline. The question asks about the recommendation according to the up-to-date guideline, while the answer choices include the up-to-date recommendation, the outdated recommendation, plausible distractors… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2509/TempoMed-Bench-MCQ.mcq-gen-trainmcq-gen-testmcqa_ladin_italian
Italian-Ladin MCQA Dataset
This is a translated multiple-choice question answering dataset in the Ladin language (Val Badia Variant) from the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_all_italian', 'choices_all_ladin', 'max_choices (number of answer options)', 'answer (correct choice)'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:… See the full description on the dataset page: https://huggingface.co/datasets/ulinnuha/mcqa_ladin_italian.mcqa_ladin_italian_manual
Italian-Ladin MCQA Dataset (Golden)
This is a manually created multiple-choice question answering dataset in the Ladin language (Val Badia Variant) paired with the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_italian', 'choices_ladin', 'answer (correct choice)', 'max_choices (number of answer options)'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:… See the full description on the dataset page: https://huggingface.co/datasets/ulinnuha/mcqa_ladin_italian_manual.mcqa_ladin_italian
Italian-Ladin MCQA Dataset
This is a translated multiple-choice question answering dataset in the Ladin language (Val Badia Variant) from the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_all_italian', 'choices_all_ladin', 'max_choices', 'answer'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
To be announced.
Synthetic_Dataset_For_MCQAmcqa_ladin_italian_manual
Italian-Ladin MCQA Dataset (Golden)
This is a manually created multiple-choice question answering dataset in the Ladin language (Val Badia Variant) paired with the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_italian', 'choices_ladin', 'answer', 'max_choices'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
To be announced.
