datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tatar_detoxThis dataset was created from s-nlp/ru_paradetox (https://huggingface.co/datasets/s-nlp/ru_paradetox).
Samples was translated to Tatar using facebook/nllb-200-3.3B (https://huggingface.co/facebook/nllb-200-3.3B).
Caution: no post-processing or check-ups was perfom on this dataset.
Nllb-3.3b trained with LoRA adapter on this dataset achieved 0.41% on PAN 2025.
tatar-exams
Tatar-exams - part of TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages
This repo contains the Tatar lnaguage part of TUMLU-mini dataset from TUMLU paper. Code repository is hosted on GitHub.
Dataset
TurkicMMLU spans 8 Turkic languages, with plans to add more.
Azerbaijani
Crimean Tatar
Karakalpak
Kazakh
Tatar
Turkish
Uyghur
Uzbek
Kyrgyz
All questions are at middle- and high-school level. All questions are native, i.e., not… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tatar-exams.MPEP-Tatartatar_monocorpus
