CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shunyalabs /mongolian-speech-datasetaudio1K<n<10K2 likes183 downloads1y agoHugging Face02Blgn94 /mongolian-stt-dataset Mongolian Speech Dataset (v24 corpus) Mongolian (Cyrillic Khalkha) read speech for ASR fine-tuning: 146.9 hours across Common Voice v24, FLEURS, and MBSpeech. 2026-07-30 — two changes, read this if you pulled before that date. YouTube-sourced audio removed. 598 clips (559 train / 39 validation, ~1.1 h) are gone. Every remaining row is read speech from a redistributable public corpus. This repo now hosts the v24 corpus. It previously held the v20 blend (57,320 train / 3,017… See the full description on the dataset page: https://huggingface.co/datasets/Blgn94/mongolian-stt-dataset.audioautomatic-speech-recognition10K<n<100K0 likes168 downloads2mo agoHugging Face03Billyyy /cleaned-mongolian-datasettext100K<n<1M2 likes99 downloads2y agoHugging Face04CMLI-NLP /Mongolian-pretrain-dataset Mongolian Pretraining Dataset Dataset Information Language: Mongolian (Traditional Mongolian script) Size: ~12GB Format: Plain text (.txt) Use Case: Language model pretraining Description This dataset contains Mongolian text data for training language models on low-resource languages. The data uses Traditional Mongolian script and covers 45 core characters identified through frequency analysis. Code: The Huffman transliteration framework implementation is… See the full description on the dataset page: https://huggingface.co/datasets/CMLI-NLP/Mongolian-pretrain-dataset.text100K<n<1M6 likes92 downloads1y agoHugging Face05Asakuu /mongolian-mcq-dataset Mongolian MCQ Dataset with Sources This dataset contains Mongolian multiple-choice questions across school and general-knowledge subjects. Each row includes answer choices, the correct answer, an explanation, and source metadata. Dataset contents File Rows mongolian_ap_chemistry_mcq_100.jsonl 100 mongolian_ap_physics_slightly_harder_mcq_100.jsonl 100 mongolian_biology_highschool_wikibooks_mcq_100.jsonl 100… See the full description on the dataset page: https://huggingface.co/datasets/Asakuu/mongolian-mcq-dataset.textquestion-answering1K<n<10K0 likes82 downloads4mo agoHugging Face06Oyussi /mongolian_ocr_textsimage10K<n<100K0 likes74 downloads6h agoHugging Face07Sodkhuu /Mongolian_audiosaudio1K<n<10K1 likes71 downloads1y agoHugging Face08robertritz /mongolian_news Online Mongolian News Dataset This dataset was scraped from an online news portal in Mongolia. It contains news stories and their headlines. It is ideal for a summarization task (making headlines from story content). textsummarization100K<n<1M7 likes70 downloads2y agoHugging Face09Tsedee /monsub-mongolian-asraudio10K<n<100K0 likes41 downloads6mo agoHugging Face10Vijish /mozilla_mongolian4audio1K<n<10K0 likes35 downloads3y agoHugging Face11Bokhbat /Mongolian-LLM-Benchmark Mongolian LLM Benchmark A multi-task evaluation benchmark for large language models on the Mongolian language (Cyrillic script). Six task configurations cover open-ended QA, multiple-choice, code generation, instruction following, math, and culturally grounded knowledge. Configurations Config Rows Format Key fields 01_culture 150 Multiple choice (A–D) prompt, options, answer, source_url 02_math 150 Numeric / short answer prompt, answer, accepted_formats… See the full description on the dataset page: https://huggingface.co/datasets/Bokhbat/Mongolian-LLM-Benchmark.textquestion-answeringn<1K0 likes29 downloads4mo agoHugging Face12onlysainaa /common-voice-scripted-speech-24.0-mongoliantext10K<n<100K0 likes26 downloads8mo agoHugging Face13Ganaa0614 /mongolian-text-datasettext10K<n<100K0 likes25 downloads5mo agoHugging Face14Ganaa0614 /mongolian-commonvoice-stt-translated-fullaudio10K<n<100K1 likes25 downloads5mo agoHugging Face15Ganaa0614 /mongolian-commonvoice-stt-translatedaudio1K<n<10K0 likes24 downloads6mo agoHugging Face16bayartsogt /mongolian-nertext10K<n<100K0 likes22 downloads4y agoHugging Face17Vijish /mozilla_mongolian3audio1K<n<10K0 likes22 downloads3y agoHugging Face18toorgil /mongolian-llm-benchmark Mongolian LLM Benchmark A combined Mongolian-language benchmark dataset for evaluating large language models. Aggregated from 19 community datasets on HuggingFace, normalized to a single unified schema. Stat Value Total rows 47,974 Language Mongolian (mn) Source datasets 19 Question types QA, MCQ, Problem Solving, Code, DPO, Instruction Following Source datasets Dataset Rows Category TRUMO12/LLM_QA 10,000 general… See the full description on the dataset page: https://huggingface.co/datasets/toorgil/mongolian-llm-benchmark.textquestion-answering10K<n<100K0 likes22 downloads4mo agoHugging Face19dulguun222 /mongolian-chat-datasettext1K<n<10K0 likes18 downloads1y agoHugging Face20Ganaa0614 /mongolian-qa-datasettext100K<n<1M0 likes15 downloads5mo agoHugging Face21Bokhbat /mongolian-dpo-ultrafeedback mongolian-dpo-ultrafeedback Mongolian (Cyrillic) DPO preference pairs, machine-translated from HuggingFaceH4/ultrafeedback_binarized with facebook/nllb-200-3.3B and filtered for Cyrillic ratio, minimum length, and chosen/rejected length balance. Schema column type description prompt string user prompt (Mongolian Cyrillic) chosen string preferred response rejected string dispreferred response Stats Rows: 35062 Source:… See the full description on the dataset page: https://huggingface.co/datasets/Bokhbat/mongolian-dpo-ultrafeedback.texttext-generation10K<n<100K0 likes15 downloads5mo agoHugging Face22baaska7 /mongolian-englishtext10K<n<100K0 likes15 downloads4mo agoHugging Face23saillab /alpaca-mongolian-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-mongolian-cleaned.text10K<n<100K2 likes14 downloads2y agoHugging Face24onon214 /mongolian-ner-demotext10K<n<100K0 likes13 downloads4y agoHugging Face25Bokhbat /mongolian-dpo-orca mongolian-dpo-orca Mongolian (Cyrillic) DPO preference pairs, machine-translated from Intel/orca_dpo_pairs with facebook/nllb-200-3.3B and filtered for Cyrillic ratio, minimum length, and chosen/rejected length balance. Schema column type description prompt string user prompt (Mongolian Cyrillic) chosen string preferred response rejected string dispreferred response Stats Rows: 9664 Source: Intel/orca_dpo_pairs Translator:… See the full description on the dataset page: https://huggingface.co/datasets/Bokhbat/mongolian-dpo-orca.texttext-generation1K<n<10K0 likes12 downloads5mo agoHugging Face26Bokhbat /mongolian-dpo-datasettext1K<n<10K0 likes11 downloads5mo agoHugging Face27bayartsogt /mongolian-ner-demotext10K<n<100K0 likes9 downloads4y agoHugging Face28Buyandelger /mongolian-nertext10K<n<100K0 likes9 downloads4y agoHugging Face29vaibhav1 /Mongolian_FakeNews_Comprehendo_datasettextn<1K0 likes9 downloads2y agoHugging Face30vaibhav1 /gpt_rationales_for_mongolian_newstextn<1K0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.