pkupie/mc2_corpus
MC^2: A Multilingual Corpus of Minority Languages in China We present MC^2, a Multilingual Corpus of Minority Languages in China, which is the largest open-source corpus so far. This corpus encompasses four languages, namely Tibetan, Uyghur, Kazakh written in the Kazakh Arabic script, and Mongolian written in the traditional Mongolian script. Please read our paper for more information: MC^2: Towards Transparent and Culturally-Aware NLP for Minority Languages in China (ACL 2024).… See the full description on the dataset page: https://huggingface.co/datasets/pkupie/mc2_corpus.
1997
No card is published for this repository, or it could not be fetched from Hugging Face right now.
