datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mbspeech-mn
MBSpeech Mongolian (cleaned)
A quality-filtered, normalised subset of MBSpeech Mongolian (Bible read-speech), prepared for training
Mongolian (Khalkha Cyrillic) text-to-speech with
oron-tts.
Built by oron-cleaner. Every threshold
was calibrated on this corpus, and every number and column on this page is read
from the shipped data rather than asserted.
from datasets import load_dataset
ds = load_dataset("btsee/mbspeech-mn-clean", split="train")
print(ds[0]["text"]… See the full description on the dataset page: https://huggingface.co/datasets/btsee/mbspeech-mn.mbspeech
mbspeech
Mongolian speech recognition dataset
Dataset Statistics
Total samples: 3,846Total duration: 6h 37m 43s (6.63 h)
Per-split breakdown
Split
Samples
Total Duration
Avg Duration
train
3,846
6h 37m 43s (6.63 h)
6.20 s
synthetic_mbspeech_dataset_edgetts
