mongolian
mongolian-speech-datasetmongolian-stt-dataset
Mongolian Speech Dataset (v24 corpus)
Mongolian (Cyrillic Khalkha) read speech for ASR fine-tuning: 146.9 hours
across Common Voice v24, FLEURS, and MBSpeech.
2026-07-30 — two changes, read this if you pulled before that date.
YouTube-sourced audio removed. 598 clips (559 train / 39 validation, ~1.1 h)
are gone. Every remaining row is read speech from a redistributable public corpus.
This repo now hosts the v24 corpus. It previously held the v20 blend
(57,320 train / 3,017… See the full description on the dataset page: https://huggingface.co/datasets/Blgn94/mongolian-stt-dataset.whisper-mongolian-finalcleaned-mongolian-datasetmongolian_instruments_dataset
Mongolian Traditional Instruments Dataset for Music Source Separation
Diploma project: Music Source Separation Using Deep Learning: A Study on Modern and Mongolian Traditional Instruments
Бүтэц
raw/ — YouTube-аас татсан түүхий аудио (4 зэмсэг)
morin_khuur/ — 719 wav (морин хуур)
yatga/ — 511 wav (ятга)
limbe/ — 643 wav (лимбэ)
tovshuur/ — 484 wav (товшуур)
clean_morin_khuur_v2/ — Audio QC v2 pipeline-аар цэвэр solo морин хуур сегмент (~1 GB)
synth_dataset/ — Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Tuugu/mongolian_instruments_dataset.Mongolian-pretrain-dataset
Mongolian Pretraining Dataset
Dataset Information
Language: Mongolian (Traditional Mongolian script)
Size: ~12GB
Format: Plain text (.txt)
Use Case: Language model pretraining
Description
This dataset contains Mongolian text data for training language models on low-resource languages. The data uses Traditional Mongolian script and covers 45 core characters identified through frequency analysis.
Code: The Huffman transliteration framework implementation is… See the full description on the dataset page: https://huggingface.co/datasets/CMLI-NLP/Mongolian-pretrain-dataset.
