CoolFace
20 results

mongolian

shunyalabs /mongolian-speech-datasetaudio1K<n<10K2 likes173 downloads1y agoHugging FaceBlgn94 /mongolian-stt-dataset Mongolian Speech Dataset (v24 corpus) Mongolian (Cyrillic Khalkha) read speech for ASR fine-tuning: 146.9 hours across Common Voice v24, FLEURS, and MBSpeech. 2026-07-30 — two changes, read this if you pulled before that date. YouTube-sourced audio removed. 598 clips (559 train / 39 validation, ~1.1 h) are gone. Every remaining row is read speech from a redistributable public corpus. This repo now hosts the v24 corpus. It previously held the v20 blend (57,320 train / 3,017… See the full description on the dataset page: https://huggingface.co/datasets/Blgn94/mongolian-stt-dataset.audioautomatic-speech-recognition10K<n<100K0 likes171 downloads2mo agoHugging FaceCafet /whisper-mongolian-final1K<n<10K1 likes169 downloads2y agoHugging FaceBillyyy /cleaned-mongolian-datasettext100K<n<1M2 likes100 downloads2y agoHugging FaceTuugu /mongolian_instruments_dataset Mongolian Traditional Instruments Dataset for Music Source Separation Diploma project: Music Source Separation Using Deep Learning: A Study on Modern and Mongolian Traditional Instruments Бүтэц raw/ — YouTube-аас татсан түүхий аудио (4 зэмсэг) morin_khuur/ — 719 wav (морин хуур) yatga/ — 511 wav (ятга) limbe/ — 643 wav (лимбэ) tovshuur/ — 484 wav (товшуур) clean_morin_khuur_v2/ — Audio QC v2 pipeline-аар цэвэр solo морин хуур сегмент (~1 GB) synth_dataset/ — Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Tuugu/mongolian_instruments_dataset.audioaudio-to-audio1K<n<10K1 likes89 downloads4mo agoHugging FaceCMLI-NLP /Mongolian-pretrain-dataset Mongolian Pretraining Dataset Dataset Information Language: Mongolian (Traditional Mongolian script) Size: ~12GB Format: Plain text (.txt) Use Case: Language model pretraining Description This dataset contains Mongolian text data for training language models on low-resource languages. The data uses Traditional Mongolian script and covers 45 core characters identified through frequency analysis. Code: The Huffman transliteration framework implementation is… See the full description on the dataset page: https://huggingface.co/datasets/CMLI-NLP/Mongolian-pretrain-dataset.text100K<n<1M6 likes88 downloads1y agoHugging Face