datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tatartts_vibevoice-formattatara_kogasa_touhou
Dataset of tatara_kogasa/多々良小傘/타타라코가사 (Touhou)
This is the dataset of tatara_kogasa/多々良小傘/타타라코가사 (Touhou), containing 500 images and their tags.
The core tags of this character are blue_hair, short_hair, red_eyes, blue_eyes, heterochromia, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages
Name
Images
Size… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/tatara_kogasa_touhou.TatarTTS
TatarTTS Dataset
Paper: TatarTTS: An Open-Source Text-to-Speech Synthesis Dataset for the Tatar Language
GitHub: https://github.com/IS2AI/TatarTTS
Description: TatarTTS is an open-source text-to-speech dataset for the Tatar language.
The dataset comprises ~70 hours of transcribed audio recordings, featuring two professional speakers (one male and one female).
Citation:
The project was developed in academic collaboration between ISSAI and Institute of Applied Semiotics of Tatarstan… See the full description on the dataset page: https://huggingface.co/datasets/issai/TatarTTS.sampled-tatar-datasetSampled Tatar dataset based on https://huggingface.co/datasets/HuggingFaceFW/fineweb-2
tatar-speech-commands
An Open-Source Tatar Speech Commands Dataset
Paper: Paper
An Open-Source Tatar Speech Commands Dataset for IoT and Robotics Applications
GitHub: https://github.com/IS2AI/TatarSCR
Description:
The dataset covers 35 commands used in robotics, IoT, and smart systems. In total, the dataset contains 3,547 one-second utterances from 153 people. The utterances were saved in the WAV format with a sampling rate of 16 kHz.
Citation: The project was developed in academic collaboration between… See the full description on the dataset page: https://huggingface.co/datasets/issai/tatar-speech-commands.tatar-russian-parallel-corporaТатарско-русский параллельный корпус.
@inproceedings{
title={Tatar parallel corpus},
author={Academy of Siences of the Recpublic of Tatarstan, Institute of Applied Semiotics.},
year={2023}
}
tatar-news-analysis-binary
Dataset Card for Tatar News Analysis Binary
Dataset Details
Dataset Description
A binary text classification dataset for Tatar news articles, designed to support tasks such as sentiment analysis, topic detection, or category classification (e.g., positive/negative, relevant/irrelevant). The dataset contains short news excerpts in the Tatar language with binary labels, collected from publicly available online sources. It is intended for training and… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-binary.tatar-news-analysis-multilabel
Dataset Card for Tatar News Multilabel Classification
Dataset Details
Dataset Description
The Tatar News Multilabel Classification Dataset contains 55,709 Tatar language news articles annotated with 13 distinct topic labels in a multi-label setting (each article can have multiple labels). Each entry includes the full article content, title, label indices, multi-hot label vector, number of labels, original single category, source URL, publication… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multilabel.tatar-web-corpus
Dataset Card for Tatar Web Corpus
Dataset Details
Dataset Description
The largest open corpus for the Tatar language with over 1 million documents collected from news websites, social media, articles, books, and Wikipedia. Designed for various NLP tasks including language modeling, text classification, information extraction, and search.
Curated by: TatarNLPWorld Community
Language(s) (NLP): Tatar (tt)
License: other – see Licensing & Legal Notice… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus.tatar-ocr-benchmark
Tatar OCR Benchmark
This dataset is a page-level OCR benchmark for 251 Tatar documents.
Each item contains:
the original page image,
structured OCR/layout annotations in JSONL,
a visual control image for fast human verification (original | reconstructed).
Full original page images are included.
Acknowledgements
We express our appreciation to Yandex LLC for support in dataset collection.
What Is Included
Export folder structure:
images/...… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tatar-ocr-benchmark.tatar-english-russian-corpus
Dataset Card: Tatar-English-Russian Parallel Corpus
Dataset Details
Dataset Description
This dataset is a parallel corpus containing 14,983 sentences in three languages: Tatar, English, and Russian. It combines two distinct sources:
KickItLikeShika/english-tatar-translation (7,746 entries) – an existing English-Tatar dataset with Russian translations added
yasalma/tt-en-language-corpus (7,615 entries) – another English-Tatar corpus with newly added… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-english-russian-corpus.english-tatar-translation
Synthetic English-Tatar Dataset
Dataset was built using DeepSeek R1 model, the it was built as part of the research project in Low Resource Machine Translation Workshop (EACL26) https://www.loresmt.org/
Citation
@inproceedings{khamis-2026-navigating,
title = "Navigating Data Scarcity in Low-Resource {E}nglish-{T}atar Translation using {LLM} Fine-Tuning",
author = "Khamis, Ahmed Khaled",
editor = "Ojha, Atul Kr. and
Liu, Chao-hong and
Vylomova… See the full description on the dataset page: https://huggingface.co/datasets/KickItLikeShika/english-tatar-translation.tatar-wiki-corpus
Dataset Card for Tatar Wiki Corpus
Dataset Details
Dataset Description
A comprehensive cleaned corpus of Tatar Wikipedia and Wikibooks with over 467,000 articles. This dataset is ideal for training language models, text classification, information retrieval, and various NLP tasks for the Tatar language.
Curated by: TatarNLPWorld Community
Language(s) (NLP): Tatar (tt)
License: cc-by-sa-4.0 – see Licensing & Legal Notice below.… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-wiki-corpus.tatar-web-corpus-v3
Dataset Card for Tatar Web Corpus
Dataset Details
Dataset Description
The Tatar Web Corpus is the largest open-source corpus of the Tatar language (Turkic family), containing 2,465,867 documents (approximately 251 million tokens) collected from publicly available web sources. It covers news portals, social media, blogs, literary websites, and other domains. The corpus underwent soft deduplication to remove exact duplicates while preserving… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus-v3.bulak-ultimate-tatar-corpus
Ultimate cleaned Tatar Corpus
The training corpus of the boolak project (a Tatar language model trained from
scratch), frozen right before tokenization. One row = one document. The text
is cleaned, filtered and deduplicated, but not tokenized: this is the exact
input of to_ids.py.
Version: v20260824 — built on 2026-08-24
from datasets import load_dataset
ds = load_dataset("eldiablo92/bulak-ultimate-tatar-corpus", split="train", revision="v20260824") # pin the version
ds =… See the full description on the dataset page: https://huggingface.co/datasets/eldiablo92/bulak-ultimate-tatar-corpus.tatar-news-cluster
Dataset Card for Tatar News Clustered Dataset
Dataset Details
Dataset Description
The Tatar News Clustered Dataset is a comprehensive collection of 57,340 Tatar language news articles with topic categories, curated by TatarNLPWorld as part of the Tat2Vec project. The dataset includes full article content, titles, source names, publication dates, and 282 topic categories. It is designed for multi-class text classification, topic modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-cluster.tatar-news-analysis-multiclass
Dataset Card for Tatar News Multiclass Classification
Dataset Details
Dataset Description
The Tatar News Multiclass Classification Dataset contains 86,963 Tatar language news articles classified into 9 distinct topic categories. Each entry includes the full article content, title, category (numeric label and text label), source URL, publication date, and content length. The dataset is specifically designed for training and evaluating multi-class… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multiclass.tatar-folklore-corpus
Dataset Card for Tatar Text Corpus with Rich Metadata
Dataset Details
Dataset Description
This dataset is a collection of 326 Tatar language texts with extensive metadata, curated by TatarNLPWorld. Each record includes the full text and a structured metadata object containing fields such as title, author, year, source, genre, category, and more. The dataset is designed for NLP research on the Tatar language, including text classification, language… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-folklore-corpus.alpaca-tatar-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-tatar-cleaned.generated-ud-tatar
Data Splits
Split
Description
train
Automatically translated and silver-annotated sentences derived from UD training sources
dev
Silver-annotated evaluation data derived from UD test resources
test
Turkic UD parallel corpora - https://github.com/ud-turkic/parallel
Dataset Creation
Source Data
This dataset is a derived work based on resources from:
Universal Dependencies treebanks - https://github.com/UniversalDependencies
Turkic UD… See the full description on the dataset page: https://huggingface.co/datasets/turkicnlp/generated-ud-tatar.tatarstan-toponyms
Dataset Card for Toponyms of Tatarstan
Dataset Details
Dataset Description
A comprehensive dataset of 9,688 toponyms (place names) from Tatarstan and Tatar-populated regions, curated by TatarNLPWorld as part of the Tat2Vec project. Each entry provides detailed linguistic, geographical, and etymological information about Tatar and Russian place names. The dataset is specifically designed for linguistic research, onomastic studies, and training NLP… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatarstan-toponyms.Tatar-Mega-ASRtatar_translation_datasetAuthors and Citation
The dataset has been developed in Institute of Applied Semiotics of Tatarstan Academy of Sciences (https://www.antat.ru/ru/ips/)
Telegram channel: https://t.me/ipsanrt
tatartatar-morphology-benchmark
Tatar Morphology Benchmark
This repository contains evaluation results for morphological analysis models trained on the Tatar Morphological Corpus.
Models Evaluated
mBERT
RuBERT
DistilBERT
LSTM
Turkish BERT
XLM-R
Key Results (Test Set Accuracy)
Model
Accuracy
F1 (micro)
mBERT
0.9905
0.9905
RuBERT
0.9861
0.9861
DistilBERT
0.9850
0.9850
XLM-R
0.9837
0.9837
LSTM
0.9440
0.9440
Turkish BERT
0.8769
0.8769
All results are based on a test… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-morphology-benchmark.tatari_kogasa_touhou
Dataset of tatari_kogasa/祟小傘 (Touhou)
This is the dataset of tatari_kogasa/祟小傘 (Touhou), containing 27 images and their tags.
The core tags of this character are blue_hair, red_eyes, blue_eyes, heterochromia, breasts, short_hair, medium_breasts, large_breasts, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/tatari_kogasa_touhou.tatartts-male-snac-24khz-tatTatar_IQA_dsalpaca_tatar_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_tatar_taco.audiobooks-tatartts__snac-24khz-tat
