datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Traditional-Chinese-Medicine-Multiple_choice_question
Discription
This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.multiple-choice-questions
Questões de Múltipla Escolha - Base de dados (PT-BR)
Contextualização
Este repositório contém uma base de dados (data.json) com questões de múltipla escolha, a qual foi utilizada principalmente no desenvolvimento de modelos de recuperação de informação.
Descrição do conjunto de dados
O conjunto de dados é composto por questões de múltipla escolha, abrangendo uma variedade de temas dentro da área da Ciência da Computação. Cada questão é estruturada em formato… See the full description on the dataset page: https://huggingface.co/datasets/mateus-hamade/multiple-choice-questions.Vietnamese-Openorca-Multiplechoice-gg-translatedkorean-bar-exam-moj-multiple-choice
Korean Bar Exam Multiple-Choice Questions and Answers (MOJ)
대한민국 법무부가 공개한 변호사시험 선택형(다지선다) 기출문제와 공식 정답을 문항 단위로 정리한 데이터셋입니다.
이 데이터셋은 해설 데이터가 아니라 문제 -> 정답 번호 학습/평가에 맞춰져 있습니다.
SFT에 바로 쓸 파일
SFT에 가장 바로 쓰기 좋은 파일은 다음입니다.
data/questions.csv
핵심 컬럼:
question_text: 문제 원문과 선택지를 포함한 전체 텍스트
stem: 문제 본문
choices_json: 선택지 배열(JSON 문자열)
answer: 공식 정답 번호
round, year, subject, question_no: 회차, 연도, 과목, 문항 번호
source_article_url, source_file_url, source_license: 출처와 라이선스 추적 정보… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-moj-multiple-choice.genetic-counselor-multiple-choiceA collection of multiple-choice questions intended for students preparing for the
American Board of Genetic Counseling (ABGC) Certification Examination.
Also see the genetic-counselor-freeform-questions evaluation set.
A genetic counselor must be prepared to answer questions about inheritance of traits,
medical statistics, testing, empathetic and ethical conversations with patients,
and observing symptoms.
For evaluation only
The goal of this dataset is to evaluate LLMs and… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/genetic-counselor-multiple-choice.abc-multiple-choice
abc-multiple-choice Dataset
abc-multiple-choice は、競技クイズの大会「abc」で使用された4択問題を元に作成された、多肢選択式の質問応答データセットです。
データセットの詳細については、下記の発表資料を参照してください。
鈴木正敏. 4択クイズを題材にした多肢選択式日本語質問応答データセットの構築. 言語処理学会第30回年次大会 (NLP2024) 併設ワークショップ 日本語言語資源の構築と利用性の向上 (JLR2024), 2024. [PDF]
下記の GitHub リポジトリで、本データセットを用いた評価実験のスクリプトを管理しています。
https://github.com/cl-tohoku/abc-multiple-choice
ライセンス
本データセットのクイズ問題の著作権は abc/EQIDEN 実行委員会 に帰属します。
本データセットは研究目的での利用許諾を得ているものです。商用目的での利用は不可とします。
Multiple-Choice-Questions
🇰🇿 Kazakh Multiple Choice Question Answering
📖 Overview
This dataset consists of 600 high-quality multiple-choice questions (MCQs) in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
600
Total Words (approx.)
19,423
Avg. Words per Sample
32
Word Count Distribution (Per Field)
The following table details the distribution of word counts across different fields… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Multiple-Choice-Questions.parsinlu-multiple-choice-alpaca-style
ParsiNLU Multiple Choice in Alpaca Style
This dataset is an Alpaca-style and instruction-included version of the ParsiNLU original dataset.
Multiple_choices-30k
Dataset Description
Combined datasets for training LLMs on Qwestion_Answering and Multiple_Choices.
These datasets are all for general knowledge, commonsense reasoning, and reading comprehension abilities of AI models.
Source Data
allenai/ai2_arc -----------------"https://huggingface.co/datasets/allenai/ai2_arc"
allenai/openbookqa ----------"https://huggingface.co/datasets/allenai/openbookqa"
tau/commonsense_qa… See the full description on the dataset page: https://huggingface.co/datasets/HashTag766/Multiple_choices-30k.
