CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bedanar /tatartts_vibevoice-formattext10K<n<100K0 likes249 downloads1y agoHugging Face02tagay1n /sampled-tatar-datasetSampled Tatar dataset based on https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 tabular1K<n<10K1 likes42 downloads2y agoHugging Face03AigizK /tatar-russian-parallel-corporaТатарско-русский параллельный корпус. @inproceedings{ title={Tatar parallel corpus}, author={Academy of Siences of the Recpublic of Tatarstan, Institute of Applied Semiotics.}, year={2023} } text100K<n<1M3 likes37 downloads3y agoHugging Face04TatarNLPWorld /tatar-news-analysis-binarygated Dataset Card for Tatar News Analysis Binary Dataset Details Dataset Description A binary text classification dataset for Tatar news articles, designed to support tasks such as sentiment analysis, topic detection, or category classification (e.g., positive/negative, relevant/irrelevant). The dataset contains short news excerpts in the Tatar language with binary labels, collected from publicly available online sources. It is intended for training and… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-binary.texttext-classification1K<n<10K0 likes36 downloads4d agoHugging Face05TatarNLPWorld /tatar-web-corpusgated Dataset Card for Tatar Web Corpus Dataset Details Dataset Description The largest open corpus for the Tatar language with over 1 million documents collected from news websites, social media, articles, books, and Wikipedia. Designed for various NLP tasks including language modeling, text classification, information extraction, and search. Curated by: TatarNLPWorld Community Language(s) (NLP): Tatar (tt) License: other – see Licensing & Legal Notice… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus.texttext-classification1M<n<10M0 likes34 downloads26d agoHugging Face06TatarNLPWorld /tatar-news-analysis-multilabelgated Dataset Card for Tatar News Multilabel Classification Dataset Details Dataset Description The Tatar News Multilabel Classification Dataset contains 55,709 Tatar language news articles annotated with 13 distinct topic labels in a multi-label setting (each article can have multiple labels). Each entry includes the full article content, title, label indices, multi-hot label vector, number of labels, original single category, source URL, publication… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multilabel.tabulartext-classification10K<n<100K0 likes34 downloads26d agoHugging Face07KickItLikeShika /english-tatar-translation Synthetic English-Tatar Dataset Dataset was built using DeepSeek R1 model, the it was built as part of the research project in Low Resource Machine Translation Workshop (EACL26) https://www.loresmt.org/ Citation @inproceedings{khamis-2026-navigating, title = "Navigating Data Scarcity in Low-Resource {E}nglish-{T}atar Translation using {LLM} Fine-Tuning", author = "Khamis, Ahmed Khaled", editor = "Ojha, Atul Kr. and Liu, Chao-hong and Vylomova… See the full description on the dataset page: https://huggingface.co/datasets/KickItLikeShika/english-tatar-translation.text1K<n<10K0 likes33 downloads6mo agoHugging Face08TatarNLPWorld /tatar-english-russian-corpusgated Dataset Card: Tatar-English-Russian Parallel Corpus Dataset Details Dataset Description This dataset is a parallel corpus containing 14,983 sentences in three languages: Tatar, English, and Russian. It combines two distinct sources: KickItLikeShika/english-tatar-translation (7,746 entries) – an existing English-Tatar dataset with Russian translations added yasalma/tt-en-language-corpus (7,615 entries) – another English-Tatar corpus with newly added… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-english-russian-corpus.tabulartranslation10K<n<100K0 likes32 downloads29d agoHugging Face09kamachkin /Tatar-Mega-ASRaudio100K<n<1M0 likes27 downloads2mo agoHugging Face10TatarNLPWorld /tatar-wiki-corpusgated Dataset Card for Tatar Wiki Corpus Dataset Details Dataset Description A comprehensive cleaned corpus of Tatar Wikipedia and Wikibooks with over 467,000 articles. This dataset is ideal for training language models, text classification, information retrieval, and various NLP tasks for the Tatar language. Curated by: TatarNLPWorld Community Language(s) (NLP): Tatar (tt) License: cc-by-sa-4.0 – see Licensing & Legal Notice below.… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-wiki-corpus.tabulartext-classification100K<n<1M0 likes25 downloads26d agoHugging Face11TatarNLPWorld /tatar-web-corpus-v3gated Dataset Card for Tatar Web Corpus Dataset Details Dataset Description The Tatar Web Corpus is the largest open-source corpus of the Tatar language (Turkic family), containing 2,465,867 documents (approximately 251 million tokens) collected from publicly available web sources. It covers news portals, social media, blogs, literary websites, and other domains. The corpus underwent soft deduplication to remove exact duplicates while preserving… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus-v3.texttext-generation1M<n<10M0 likes25 downloads26d agoHugging Face12TatarNLPWorld /tatar-news-clustergated Dataset Card for Tatar News Clustered Dataset Dataset Details Dataset Description The Tatar News Clustered Dataset is a comprehensive collection of 57,340 Tatar language news articles with topic categories, curated by TatarNLPWorld as part of the Tat2Vec project. The dataset includes full article content, titles, source names, publication dates, and 282 topic categories. It is designed for multi-class text classification, topic modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-cluster.texttext-classification10K<n<100K0 likes24 downloads26d agoHugging Face13eldiablo92 /bulak-ultimate-tatar-corpusgated Ultimate cleaned Tatar Corpus The training corpus of the boolak project (a Tatar language model trained from scratch), frozen right before tokenization. One row = one document. The text is cleaned, filtered and deduplicated, but not tokenized: this is the exact input of to_ids.py. Version: v20260824 — built on 2026-08-24 from datasets import load_dataset ds = load_dataset("eldiablo92/bulak-ultimate-tatar-corpus", split="train", revision="v20260824") # pin the version ds =… See the full description on the dataset page: https://huggingface.co/datasets/eldiablo92/bulak-ultimate-tatar-corpus.tabulartext-generation1M<n<10M0 likes24 downloads25d agoHugging Face14saillab /alpaca-tatar-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-tatar-cleaned.text10K<n<100K0 likes23 downloads2y agoHugging Face15TatarNLPWorld /tatar-news-analysis-multiclassgated Dataset Card for Tatar News Multiclass Classification Dataset Details Dataset Description The Tatar News Multiclass Classification Dataset contains 86,963 Tatar language news articles classified into 9 distinct topic categories. Each entry includes the full article content, title, category (numeric label and text label), source URL, publication date, and content length. The dataset is specifically designed for training and evaluating multi-class… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multiclass.tabulartext-classification10K<n<100K0 likes22 downloads26d agoHugging Face16TatarNLPWorld /tatar-folklore-corpusgated Dataset Card for Tatar Text Corpus with Rich Metadata Dataset Details Dataset Description This dataset is a collection of 326 Tatar language texts with extensive metadata, curated by TatarNLPWorld. Each record includes the full text and a structured metadata object containing fields such as title, author, year, source, genre, category, and more. The dataset is designed for NLP research on the Tatar language, including text classification, language… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-folklore-corpus.texttext-generationn<1K0 likes22 downloads26d agoHugging Face17TatarNLPWorld /tatarstan-toponymsgated Dataset Card for Toponyms of Tatarstan Dataset Details Dataset Description A comprehensive dataset of 9,688 toponyms (place names) from Tatarstan and Tatar-populated regions, curated by TatarNLPWorld as part of the Tat2Vec project. Each entry provides detailed linguistic, geographical, and etymological information about Tatar and Russian place names. The dataset is specifically designed for linguistic research, onomastic studies, and training NLP… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatarstan-toponyms.tabulartoken-classification1K<n<10K0 likes20 downloads26d agoHugging Face18IPSAN /tatar_translation_datasetgatedAuthors and Citation The dataset has been developed in Institute of Applied Semiotics of Tatarstan Academy of Sciences (https://www.antat.ru/ru/ips/) Telegram channel: https://t.me/ipsanrt text100K<n<1M4 likes16 downloads1y agoHugging Face19skdwoqwqje /tatartext10K<n<100K1 likes14 downloads1y agoHugging Face20saillab /alpaca_tatar_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_tatar_taco.text10K<n<100K0 likes12 downloads2y agoHugging Face21romgor /tatar-instruction-dataset-small Tatar Instruction-Following Dataset (2022 Small) Dataset Description This dataset is a representative sample from a larger, proprietary dataset created by Gerwin AI in 2022. It was originally used for a pioneering project to full fine-tune the OpenAI GPT-3 davinci model, enabling it to generate coherent and contextually relevant text in the Tatar language, a low-resource language for which no such capabilities previously existed. The dataset consists of 540 examples in… See the full description on the dataset page: https://huggingface.co/datasets/romgor/tatar-instruction-dataset-small.textn<1K2 likes9 downloads1y agoHugging Face22afwull /tatar_detoxThis dataset was created from s-nlp/ru_paradetox (https://huggingface.co/datasets/s-nlp/ru_paradetox). Samples was translated to Tatar using facebook/nllb-200-3.3B (https://huggingface.co/facebook/nllb-200-3.3B). Caution: no post-processing or check-ups was perfom on this dataset. Nllb-3.3b trained with LoRA adapter on this dataset achieved 0.41% on PAN 2025. text10K<n<100K0 likes7 downloads10mo agoHugging Face23Sakalti /Saishin-tatartextn<1K0 likes6 downloads2y agoHugging Face24TatarNLPWorld /tatarstan-toponyms-qagated Dataset Card for Tatarstan Toponyms QA Dataset A question-answering dataset about toponyms (place names) of Tatarstan, containing 38,696 QA pairs in Russian and Tatar languages. The dataset covers various aspects of geographical names including their type, coordinates, sources, etymology, administrative region, location, and physical characteristics. Dataset Details Dataset Description This dataset provides extractive question-answering pairs about toponyms… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatarstan-toponyms-qa.text10K<n<100K0 likes6 downloads7mo agoHugging Face25yasalma /tatar-examsgated Tatar-exams - part of TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages This repo contains the Tatar lnaguage part of TUMLU-mini dataset from TUMLU paper. Code repository is hosted on GitHub. Dataset TurkicMMLU spans 8 Turkic languages, with plans to add more. Azerbaijani Crimean Tatar Karakalpak Kazakh Tatar Turkish Uyghur Uzbek Kyrgyz All questions are at middle- and high-school level. All questions are native, i.e., not… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tatar-exams.textmultiple-choice1K<n<10K0 likes5 downloads1y agoHugging Face26gaydmi /MPEP-Tatartextn<1K0 likes4 downloads2y agoHugging Face27alexkkir /tatar-wikipedia-corpustext1M<n<10M0 likes4 downloads9mo agoHugging Face28TatarNLPWorld /sovet_kinesh-vkgated Dataset Card for Sovet Kinesh VK Dataset Dataset Details Dataset Description This dataset contains posts and comments from the VK community "Совет Кинеш" (Sovet Kinesh), a Tatar-language community focused on cultural, social, and educational content. The data was collected to support NLP research and development for the Tatar language, particularly for tasks such as text classification, sentiment analysis, language modeling, and social media analysis. Curated… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/sovet_kinesh-vk.tabulartext-classification100K<n<1M0 likes3 downloads6mo agoHugging Face29TatarNLPWorld /vk-groupsgated VK Groups Dataset: Tatar Social Media Posts Dataset Description This dataset contains posts from five popular VKontakte (VK) communities targeting Tatar-speaking audiences. It includes content across different themes: religion, relationships, humor, advice, and cooking. The dataset is designed for various NLP tasks including text classification, sentiment analysis, topic modeling, and engagement prediction. Curated by: TatarNLPWorld Language(s): Tatar (Cyrillic script)… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/vk-groups.tabulartext-classification100K<n<1M0 likes3 downloads7mo agoHugging Face30IPSAN /tatar_monocorpusgatedtext100K<n<1M0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.