datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tatar-wiki-corpus
Dataset Card for Tatar Wiki Corpus
Dataset Details
Dataset Description
A comprehensive cleaned corpus of Tatar Wikipedia and Wikibooks with over 467,000 articles. This dataset is ideal for training language models, text classification, information retrieval, and various NLP tasks for the Tatar language.
Curated by: TatarNLPWorld Community
Language(s) (NLP): Tatar (tt)
License: cc-by-sa-4.0 – see Licensing & Legal Notice below.… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-wiki-corpus.tatarstan-toponyms
Dataset Card for Toponyms of Tatarstan
Dataset Details
Dataset Description
A comprehensive dataset of 9,688 toponyms (place names) from Tatarstan and Tatar-populated regions, curated by TatarNLPWorld as part of the Tat2Vec project. Each entry provides detailed linguistic, geographical, and etymological information about Tatar and Russian place names. The dataset is specifically designed for linguistic research, onomastic studies, and training NLP… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatarstan-toponyms.
