datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tatar-web-corpus
Dataset Card for Tatar Web Corpus
Dataset Details
Dataset Description
The largest open corpus for the Tatar language with over 1 million documents collected from news websites, social media, articles, books, and Wikipedia. Designed for various NLP tasks including language modeling, text classification, information extraction, and search.
Curated by: TatarNLPWorld Community
Language(s) (NLP): Tatar (tt)
License: other – see Licensing & Legal Notice… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus.tatar-news-analysis-multilabel
Dataset Card for Tatar News Multilabel Classification
Dataset Details
Dataset Description
The Tatar News Multilabel Classification Dataset contains 55,709 Tatar language news articles annotated with 13 distinct topic labels in a multi-label setting (each article can have multiple labels). Each entry includes the full article content, title, label indices, multi-hot label vector, number of labels, original single category, source URL, publication… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multilabel.tatar-english-russian-corpus
Dataset Card: Tatar-English-Russian Parallel Corpus
Dataset Details
Dataset Description
This dataset is a parallel corpus containing 14,983 sentences in three languages: Tatar, English, and Russian. It combines two distinct sources:
KickItLikeShika/english-tatar-translation (7,746 entries) – an existing English-Tatar dataset with Russian translations added
yasalma/tt-en-language-corpus (7,615 entries) – another English-Tatar corpus with newly added… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-english-russian-corpus.tatar-wiki-corpus
Dataset Card for Tatar Wiki Corpus
Dataset Details
Dataset Description
A comprehensive cleaned corpus of Tatar Wikipedia and Wikibooks with over 467,000 articles. This dataset is ideal for training language models, text classification, information retrieval, and various NLP tasks for the Tatar language.
Curated by: TatarNLPWorld Community
Language(s) (NLP): Tatar (tt)
License: cc-by-sa-4.0 – see Licensing & Legal Notice below.… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-wiki-corpus.tatar-web-corpus-v3
Dataset Card for Tatar Web Corpus
Dataset Details
Dataset Description
The Tatar Web Corpus is the largest open-source corpus of the Tatar language (Turkic family), containing 2,465,867 documents (approximately 251 million tokens) collected from publicly available web sources. It covers news portals, social media, blogs, literary websites, and other domains. The corpus underwent soft deduplication to remove exact duplicates while preserving… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus-v3.tatar-news-cluster
Dataset Card for Tatar News Clustered Dataset
Dataset Details
Dataset Description
The Tatar News Clustered Dataset is a comprehensive collection of 57,340 Tatar language news articles with topic categories, curated by TatarNLPWorld as part of the Tat2Vec project. The dataset includes full article content, titles, source names, publication dates, and 282 topic categories. It is designed for multi-class text classification, topic modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-cluster.tatar-news-analysis-multiclass
Dataset Card for Tatar News Multiclass Classification
Dataset Details
Dataset Description
The Tatar News Multiclass Classification Dataset contains 86,963 Tatar language news articles classified into 9 distinct topic categories. Each entry includes the full article content, title, category (numeric label and text label), source URL, publication date, and content length. The dataset is specifically designed for training and evaluating multi-class… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multiclass.tatar-folklore-corpus
Dataset Card for Tatar Text Corpus with Rich Metadata
Dataset Details
Dataset Description
This dataset is a collection of 326 Tatar language texts with extensive metadata, curated by TatarNLPWorld. Each record includes the full text and a structured metadata object containing fields such as title, author, year, source, genre, category, and more. The dataset is designed for NLP research on the Tatar language, including text classification, language… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-folklore-corpus.bulak-ultimate-tatar-corpus
Ultimate cleaned Tatar Corpus
The training corpus of the boolak project (a Tatar language model trained from
scratch), frozen right before tokenization. One row = one document. The text
is cleaned, filtered and deduplicated, but not tokenized: this is the exact
input of to_ids.py.
Version: v20260824 — built on 2026-08-24
from datasets import load_dataset
ds = load_dataset("eldiablo92/bulak-ultimate-tatar-corpus", split="train", revision="v20260824") # pin the version
ds =… See the full description on the dataset page: https://huggingface.co/datasets/eldiablo92/bulak-ultimate-tatar-corpus.vk-groups
VK Groups Dataset: Tatar Social Media Posts
Dataset Description
This dataset contains posts from five popular VKontakte (VK) communities targeting Tatar-speaking audiences. It includes content across different themes: religion, relationships, humor, advice, and cooking. The dataset is designed for various NLP tasks including text classification, sentiment analysis, topic modeling, and engagement prediction.
Curated by: TatarNLPWorld
Language(s): Tatar (Cyrillic script)… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/vk-groups.
