datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaokao-sft-chinese-strict-abcd-v3
Gaokao SFT Chinese Strict ABCD V3
This dataset is the cleaned Chinese SFT release that keeps only single-choice samples where A, B, C, and D all have explicit option-level analysis.
Composition
Total samples: 88466
Train samples: 86670
Validation samples: 1796
Subject Counts
{
"biology": 33104,
"chemistry": 35796,
"english": 174,
"general_exam": 7982,
"geography": 888,
"history": 229,
"physics": 9986,
"politics": 307
}
Fields
id… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd-v3.VSMRC-mrc-ABCD
VSMRC/mrc (bản tách cột A/B/C/D)
Dataset này là gì
Đây là bản định dạng lại (reformatted / derived) của dataset gốc VSMRC/mrc — phần multiple-choice reading comprehension trong bộ VSMRC (Vietnamese Text Segmentation and Multiple-Choice Reading Comprehension Dataset), do nhóm tác giả tại Đại học
Công nghệ, ĐHQGHN công bố.
Dataset gốc đã có sẵn cột choices (list Python) và correctchoice (số nguyên 0-3) — bản này chỉ map lại thành các cột A, B, C, D, answer cho khớp… See the full description on the dataset page: https://huggingface.co/datasets/p-storm/VSMRC-mrc-ABCD.gaokao-sft-chinese-strict-abcd
Gaokao SFT Chinese Balanced
This is the strict balanced Chinese SFT dataset version.
Only multiple-choice samples with explicit A/B/C/D option-level explanations are kept in this balanced release.
Composition
Total samples: 645
Train samples: 632
Validation samples: 13
Subject Counts
{
"biology": 199,
"chemistry": 170,
"english": 13,
"geography": 34,
"history": 118,
"physics": 77,
"politics": 34
}
Fields
id
lang
subject
source… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd.abc-multiple-choice
abc-multiple-choice Dataset
abc-multiple-choice は、競技クイズの大会「abc」で使用された4択問題を元に作成された、多肢選択式の質問応答データセットです。
データセットの詳細については、下記の発表資料を参照してください。
鈴木正敏. 4択クイズを題材にした多肢選択式日本語質問応答データセットの構築. 言語処理学会第30回年次大会 (NLP2024) 併設ワークショップ 日本語言語資源の構築と利用性の向上 (JLR2024), 2024. [PDF]
下記の GitHub リポジトリで、本データセットを用いた評価実験のスクリプトを管理しています。
https://github.com/cl-tohoku/abc-multiple-choice
ライセンス
本データセットのクイズ問題の著作権は abc/EQIDEN 実行委員会 に帰属します。
本データセットは研究目的での利用許諾を得ているものです。商用目的での利用は不可とします。
