datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mongolian-mcq-dataset
Mongolian MCQ Dataset with Sources
This dataset contains Mongolian multiple-choice questions across school and general-knowledge subjects. Each row includes answer choices, the correct answer, an explanation, and source metadata.
Dataset contents
File
Rows
mongolian_ap_chemistry_mcq_100.jsonl
100
mongolian_ap_physics_slightly_harder_mcq_100.jsonl
100
mongolian_biology_highschool_wikibooks_mcq_100.jsonl
100… See the full description on the dataset page: https://huggingface.co/datasets/Asakuu/mongolian-mcq-dataset.Mongolian-LLM-Benchmark
Mongolian LLM Benchmark
A multi-task evaluation benchmark for large language models on the Mongolian language (Cyrillic script). Six task configurations cover open-ended QA, multiple-choice, code generation, instruction following, math, and culturally grounded knowledge.
Configurations
Config
Rows
Format
Key fields
01_culture
150
Multiple choice (A–D)
prompt, options, answer, source_url
02_math
150
Numeric / short answer
prompt, answer, accepted_formats… See the full description on the dataset page: https://huggingface.co/datasets/Bokhbat/Mongolian-LLM-Benchmark.mongolian-llm-benchmark
Mongolian LLM Benchmark
A combined Mongolian-language benchmark dataset for evaluating large language models.
Aggregated from 19 community datasets on HuggingFace, normalized to a single unified schema.
Stat
Value
Total rows
47,974
Language
Mongolian (mn)
Source datasets
19
Question types
QA, MCQ, Problem Solving, Code, DPO, Instruction Following
Source datasets
Dataset
Rows
Category
TRUMO12/LLM_QA
10,000
general… See the full description on the dataset page: https://huggingface.co/datasets/toorgil/mongolian-llm-benchmark.
