pkupie/mc2_corpus
MC^2: A Multilingual Corpus of Minority Languages in China We present MC^2, a Multilingual Corpus of Minority Languages in China, which is the largest open-source corpus so far. This corpus encompasses four languages, namely Tibetan, Uyghur, Kazakh written in the Kazakh Arabic script, and Mongolian written in the traditional Mongolian script. Please read our paper for more information: MC^2: Towards Transparent and Culturally-Aware NLP for Minority Languages in China (ACL 2024).… See the full description on the dataset page: https://huggingface.co/datasets/pkupie/mc2_corpus.
This repository belongs to pkupie on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
