datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Zerde-QA-50K
🇰🇿 Zerde-QA-50K
A large-scale synthetic Kazakh question-answer dataset for instruction tuning and NLP research.Created and maintained by kurumikz. Free to use with attribution.
📌 Overview
Zerde-QA-50K is a synthetically generated open-domain QA dataset written entirely in the Kazakh language (kk), consisting of 51,422 high-quality question-answer pairs spanning 20+ academic and professional domains.
Each record follows a clean {question, answer} structure… See the full description on the dataset page: https://huggingface.co/datasets/AbaiUniversity/Zerde-QA-50K.aba-official-curriculum-sft
ABA Official Curriculum SFT
Structured supervision dataset derived from official QABA curriculum sources for:
ABAT
QASP-S
QBA
Files
official_lessons.jsonl
official_qa.jsonl
official_mcq.jsonl
official_curriculum_sft.jsonl
official_curriculum_train.jsonl
official_curriculum_eval.jsonl
manifest.json
Intended use
This dataset is intended for:
instruction tuning on official ABA curriculum content
grounded lesson planning
grounded question answering
grounded… See the full description on the dataset page: https://huggingface.co/datasets/nopoh44/aba-official-curriculum-sft.
