datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mongolian-speech-datasetmongolian-stt-dataset
Mongolian Speech Dataset (v24 corpus)
Mongolian (Cyrillic Khalkha) read speech for ASR fine-tuning: 146.9 hours
across Common Voice v24, FLEURS, and MBSpeech.
2026-07-30 — two changes, read this if you pulled before that date.
YouTube-sourced audio removed. 598 clips (559 train / 39 validation, ~1.1 h)
are gone. Every remaining row is read speech from a redistributable public corpus.
This repo now hosts the v24 corpus. It previously held the v20 blend
(57,320 train / 3,017… See the full description on the dataset page: https://huggingface.co/datasets/Blgn94/mongolian-stt-dataset.cleaned-mongolian-datasetMongolian-pretrain-dataset
Mongolian Pretraining Dataset
Dataset Information
Language: Mongolian (Traditional Mongolian script)
Size: ~12GB
Format: Plain text (.txt)
Use Case: Language model pretraining
Description
This dataset contains Mongolian text data for training language models on low-resource languages. The data uses Traditional Mongolian script and covers 45 core characters identified through frequency analysis.
Code: The Huffman transliteration framework implementation is… See the full description on the dataset page: https://huggingface.co/datasets/CMLI-NLP/Mongolian-pretrain-dataset.mongolian-mcq-dataset
Mongolian MCQ Dataset with Sources
This dataset contains Mongolian multiple-choice questions across school and general-knowledge subjects. Each row includes answer choices, the correct answer, an explanation, and source metadata.
Dataset contents
File
Rows
mongolian_ap_chemistry_mcq_100.jsonl
100
mongolian_ap_physics_slightly_harder_mcq_100.jsonl
100
mongolian_biology_highschool_wikibooks_mcq_100.jsonl
100… See the full description on the dataset page: https://huggingface.co/datasets/Asakuu/mongolian-mcq-dataset.mongolian_ocr_textsMongolian_audiosmongolian_news
Online Mongolian News Dataset
This dataset was scraped from an online news portal in Mongolia. It contains news stories and their headlines. It is ideal for a summarization task (making headlines from story content).
monsub-mongolian-asrmozilla_mongolian4Mongolian-LLM-Benchmark
Mongolian LLM Benchmark
A multi-task evaluation benchmark for large language models on the Mongolian language (Cyrillic script). Six task configurations cover open-ended QA, multiple-choice, code generation, instruction following, math, and culturally grounded knowledge.
Configurations
Config
Rows
Format
Key fields
01_culture
150
Multiple choice (A–D)
prompt, options, answer, source_url
02_math
150
Numeric / short answer
prompt, answer, accepted_formats… See the full description on the dataset page: https://huggingface.co/datasets/Bokhbat/Mongolian-LLM-Benchmark.common-voice-scripted-speech-24.0-mongolianmongolian-text-datasetmongolian-commonvoice-stt-translated-fullmongolian-commonvoice-stt-translatedmongolian-nermozilla_mongolian3mongolian-llm-benchmark
Mongolian LLM Benchmark
A combined Mongolian-language benchmark dataset for evaluating large language models.
Aggregated from 19 community datasets on HuggingFace, normalized to a single unified schema.
Stat
Value
Total rows
47,974
Language
Mongolian (mn)
Source datasets
19
Question types
QA, MCQ, Problem Solving, Code, DPO, Instruction Following
Source datasets
Dataset
Rows
Category
TRUMO12/LLM_QA
10,000
general… See the full description on the dataset page: https://huggingface.co/datasets/toorgil/mongolian-llm-benchmark.mongolian-chat-datasetmongolian-qa-datasetmongolian-dpo-ultrafeedback
mongolian-dpo-ultrafeedback
Mongolian (Cyrillic) DPO preference pairs, machine-translated from
HuggingFaceH4/ultrafeedback_binarized with facebook/nllb-200-3.3B and filtered
for Cyrillic ratio, minimum length, and chosen/rejected length balance.
Schema
column
type
description
prompt
string
user prompt (Mongolian Cyrillic)
chosen
string
preferred response
rejected
string
dispreferred response
Stats
Rows: 35062
Source:… See the full description on the dataset page: https://huggingface.co/datasets/Bokhbat/mongolian-dpo-ultrafeedback.mongolian-englishalpaca-mongolian-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-mongolian-cleaned.mongolian-ner-demomongolian-dpo-orca
mongolian-dpo-orca
Mongolian (Cyrillic) DPO preference pairs, machine-translated from
Intel/orca_dpo_pairs with facebook/nllb-200-3.3B and filtered
for Cyrillic ratio, minimum length, and chosen/rejected length balance.
Schema
column
type
description
prompt
string
user prompt (Mongolian Cyrillic)
chosen
string
preferred response
rejected
string
dispreferred response
Stats
Rows: 9664
Source: Intel/orca_dpo_pairs
Translator:… See the full description on the dataset page: https://huggingface.co/datasets/Bokhbat/mongolian-dpo-orca.mongolian-dpo-datasetmongolian-ner-demomongolian-nerMongolian_FakeNews_Comprehendo_datasetgpt_rationales_for_mongolian_news
