CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bedanar /tatartts_vibevoice-formattext10K<n<100K0 likes249 downloads1y agoHugging Face02CyberHarem /tatara_kogasa_touhou Dataset of tatara_kogasa/多々良小傘/타타라코가사 (Touhou) This is the dataset of tatara_kogasa/多々良小傘/타타라코가사 (Touhou), containing 500 images and their tags. The core tags of this character are blue_hair, short_hair, red_eyes, blue_eyes, heterochromia, which are pruned in this dataset. Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization). List of Packages Name Images Size… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/tatara_kogasa_touhou.text-to-image1K<n<10K0 likes75 downloads3y agoHugging Face03issai /TatarTTS TatarTTS Dataset Paper: TatarTTS: An Open-Source Text-to-Speech Synthesis Dataset for the Tatar Language GitHub: https://github.com/IS2AI/TatarTTS Description: TatarTTS is an open-source text-to-speech dataset for the Tatar language. The dataset comprises ~70 hours of transcribed audio recordings, featuring two professional speakers (one male and one female). Citation: The project was developed in academic collaboration between ISSAI and Institute of Applied Semiotics of Tatarstan… See the full description on the dataset page: https://huggingface.co/datasets/issai/TatarTTS.text-to-speech10K<n<100K3 likes50 downloads2y agoHugging Face04tagay1n /sampled-tatar-datasetSampled Tatar dataset based on https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 tabular1K<n<10K1 likes42 downloads2y agoHugging Face05issai /tatar-speech-commands An Open-Source Tatar Speech Commands Dataset Paper: Paper An Open-Source Tatar Speech Commands Dataset for IoT and Robotics Applications GitHub: https://github.com/IS2AI/TatarSCR Description: The dataset covers 35 commands used in robotics, IoT, and smart systems. In total, the dataset contains 3,547 one-second utterances from 153 people. The utterances were saved in the WAV format with a sampling rate of 16 kHz. Citation: The project was developed in academic collaboration between… See the full description on the dataset page: https://huggingface.co/datasets/issai/tatar-speech-commands.audioaudio-classification1K<n<10K0 likes42 downloads2y agoHugging Face06AigizK /tatar-russian-parallel-corporaТатарско-русский параллельный корпус. @inproceedings{ title={Tatar parallel corpus}, author={Academy of Siences of the Recpublic of Tatarstan, Institute of Applied Semiotics.}, year={2023} } text100K<n<1M3 likes39 downloads3y agoHugging Face07TatarNLPWorld /tatar-news-analysis-binarygated Dataset Card for Tatar News Analysis Binary Dataset Details Dataset Description A binary text classification dataset for Tatar news articles, designed to support tasks such as sentiment analysis, topic detection, or category classification (e.g., positive/negative, relevant/irrelevant). The dataset contains short news excerpts in the Tatar language with binary labels, collected from publicly available online sources. It is intended for training and… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-binary.texttext-classification1K<n<10K0 likes35 downloads3d agoHugging Face08TatarNLPWorld /tatar-news-analysis-multilabelgated Dataset Card for Tatar News Multilabel Classification Dataset Details Dataset Description The Tatar News Multilabel Classification Dataset contains 55,709 Tatar language news articles annotated with 13 distinct topic labels in a multi-label setting (each article can have multiple labels). Each entry includes the full article content, title, label indices, multi-hot label vector, number of labels, original single category, source URL, publication… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multilabel.tabulartext-classification10K<n<100K0 likes34 downloads25d agoHugging Face09TatarNLPWorld /tatar-web-corpusgated Dataset Card for Tatar Web Corpus Dataset Details Dataset Description The largest open corpus for the Tatar language with over 1 million documents collected from news websites, social media, articles, books, and Wikipedia. Designed for various NLP tasks including language modeling, text classification, information extraction, and search. Curated by: TatarNLPWorld Community Language(s) (NLP): Tatar (tt) License: other – see Licensing & Legal Notice… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus.texttext-classification1M<n<10M0 likes33 downloads25d agoHugging Face10yasalma /tatar-ocr-benchmarkgated Tatar OCR Benchmark This dataset is a page-level OCR benchmark for 251 Tatar documents. Each item contains: the original page image, structured OCR/layout annotations in JSONL, a visual control image for fast human verification (original | reconstructed). Full original page images are included. Acknowledgements We express our appreciation to Yandex LLC for support in dataset collection. What Is Included Export folder structure: images/...… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tatar-ocr-benchmark.imageimage-to-textn<1K1 likes32 downloads26d agoHugging Face11TatarNLPWorld /tatar-english-russian-corpusgated Dataset Card: Tatar-English-Russian Parallel Corpus Dataset Details Dataset Description This dataset is a parallel corpus containing 14,983 sentences in three languages: Tatar, English, and Russian. It combines two distinct sources: KickItLikeShika/english-tatar-translation (7,746 entries) – an existing English-Tatar dataset with Russian translations added yasalma/tt-en-language-corpus (7,615 entries) – another English-Tatar corpus with newly added… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-english-russian-corpus.tabulartranslation10K<n<100K0 likes32 downloads28d agoHugging Face12KickItLikeShika /english-tatar-translation Synthetic English-Tatar Dataset Dataset was built using DeepSeek R1 model, the it was built as part of the research project in Low Resource Machine Translation Workshop (EACL26) https://www.loresmt.org/ Citation @inproceedings{khamis-2026-navigating, title = "Navigating Data Scarcity in Low-Resource {E}nglish-{T}atar Translation using {LLM} Fine-Tuning", author = "Khamis, Ahmed Khaled", editor = "Ojha, Atul Kr. and Liu, Chao-hong and Vylomova… See the full description on the dataset page: https://huggingface.co/datasets/KickItLikeShika/english-tatar-translation.text1K<n<10K0 likes31 downloads6mo agoHugging Face13TatarNLPWorld /tatar-wiki-corpusgated Dataset Card for Tatar Wiki Corpus Dataset Details Dataset Description A comprehensive cleaned corpus of Tatar Wikipedia and Wikibooks with over 467,000 articles. This dataset is ideal for training language models, text classification, information retrieval, and various NLP tasks for the Tatar language. Curated by: TatarNLPWorld Community Language(s) (NLP): Tatar (tt) License: cc-by-sa-4.0 – see Licensing & Legal Notice below.… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-wiki-corpus.tabulartext-classification100K<n<1M0 likes25 downloads25d agoHugging Face14TatarNLPWorld /tatar-web-corpus-v3gated Dataset Card for Tatar Web Corpus Dataset Details Dataset Description The Tatar Web Corpus is the largest open-source corpus of the Tatar language (Turkic family), containing 2,465,867 documents (approximately 251 million tokens) collected from publicly available web sources. It covers news portals, social media, blogs, literary websites, and other domains. The corpus underwent soft deduplication to remove exact duplicates while preserving… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus-v3.texttext-generation1M<n<10M0 likes25 downloads25d agoHugging Face15eldiablo92 /bulak-ultimate-tatar-corpusgated Ultimate cleaned Tatar Corpus The training corpus of the boolak project (a Tatar language model trained from scratch), frozen right before tokenization. One row = one document. The text is cleaned, filtered and deduplicated, but not tokenized: this is the exact input of to_ids.py. Version: v20260824 — built on 2026-08-24 from datasets import load_dataset ds = load_dataset("eldiablo92/bulak-ultimate-tatar-corpus", split="train", revision="v20260824") # pin the version ds =… See the full description on the dataset page: https://huggingface.co/datasets/eldiablo92/bulak-ultimate-tatar-corpus.tabulartext-generation1M<n<10M0 likes25 downloads24d agoHugging Face16TatarNLPWorld /tatar-news-clustergated Dataset Card for Tatar News Clustered Dataset Dataset Details Dataset Description The Tatar News Clustered Dataset is a comprehensive collection of 57,340 Tatar language news articles with topic categories, curated by TatarNLPWorld as part of the Tat2Vec project. The dataset includes full article content, titles, source names, publication dates, and 282 topic categories. It is designed for multi-class text classification, topic modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-cluster.texttext-classification10K<n<100K0 likes24 downloads25d agoHugging Face17TatarNLPWorld /tatar-news-analysis-multiclassgated Dataset Card for Tatar News Multiclass Classification Dataset Details Dataset Description The Tatar News Multiclass Classification Dataset contains 86,963 Tatar language news articles classified into 9 distinct topic categories. Each entry includes the full article content, title, category (numeric label and text label), source URL, publication date, and content length. The dataset is specifically designed for training and evaluating multi-class… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multiclass.tabulartext-classification10K<n<100K0 likes22 downloads25d agoHugging Face18TatarNLPWorld /tatar-folklore-corpusgated Dataset Card for Tatar Text Corpus with Rich Metadata Dataset Details Dataset Description This dataset is a collection of 326 Tatar language texts with extensive metadata, curated by TatarNLPWorld. Each record includes the full text and a structured metadata object containing fields such as title, author, year, source, genre, category, and more. The dataset is designed for NLP research on the Tatar language, including text classification, language… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-folklore-corpus.texttext-generationn<1K0 likes22 downloads25d agoHugging Face19saillab /alpaca-tatar-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-tatar-cleaned.text10K<n<100K0 likes21 downloads2y agoHugging Face20turkicnlp /generated-ud-tatar Data Splits Split Description train Automatically translated and silver-annotated sentences derived from UD training sources dev Silver-annotated evaluation data derived from UD test resources test Turkic UD parallel corpora - https://github.com/ud-turkic/parallel Dataset Creation Source Data This dataset is a derived work based on resources from: Universal Dependencies treebanks - https://github.com/UniversalDependencies Turkic UD… See the full description on the dataset page: https://huggingface.co/datasets/turkicnlp/generated-ud-tatar.0 likes20 downloads5mo agoHugging Face21TatarNLPWorld /tatarstan-toponymsgated Dataset Card for Toponyms of Tatarstan Dataset Details Dataset Description A comprehensive dataset of 9,688 toponyms (place names) from Tatarstan and Tatar-populated regions, curated by TatarNLPWorld as part of the Tat2Vec project. Each entry provides detailed linguistic, geographical, and etymological information about Tatar and Russian place names. The dataset is specifically designed for linguistic research, onomastic studies, and training NLP… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatarstan-toponyms.tabulartoken-classification1K<n<10K0 likes19 downloads25d agoHugging Face22kamachkin /Tatar-Mega-ASRaudio100K<n<1M0 likes17 downloads2mo agoHugging Face23IPSAN /tatar_translation_datasetgatedAuthors and Citation The dataset has been developed in Institute of Applied Semiotics of Tatarstan Academy of Sciences (https://www.antat.ru/ru/ips/) Telegram channel: https://t.me/ipsanrt text100K<n<1M4 likes16 downloads1y agoHugging Face24skdwoqwqje /tatartext10K<n<100K1 likes14 downloads1y agoHugging Face25TatarNLPWorld /tatar-morphology-benchmark Tatar Morphology Benchmark This repository contains evaluation results for morphological analysis models trained on the Tatar Morphological Corpus. Models Evaluated mBERT RuBERT DistilBERT LSTM Turkish BERT XLM-R Key Results (Test Set Accuracy) Model Accuracy F1 (micro) mBERT 0.9905 0.9905 RuBERT 0.9861 0.9861 DistilBERT 0.9850 0.9850 XLM-R 0.9837 0.9837 LSTM 0.9440 0.9440 Turkish BERT 0.8769 0.8769 All results are based on a test… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-morphology-benchmark.0 likes14 downloads6mo agoHugging Face26CyberHarem /tatari_kogasa_touhou Dataset of tatari_kogasa/祟小傘 (Touhou) This is the dataset of tatari_kogasa/祟小傘 (Touhou), containing 27 images and their tags. The core tags of this character are blue_hair, red_eyes, blue_eyes, heterochromia, breasts, short_hair, medium_breasts, large_breasts, which are pruned in this dataset. Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization). List of Packages… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/tatari_kogasa_touhou.text-to-imagen<1K0 likes12 downloads3y agoHugging Face27bedanar /tatartts-male-snac-24khz-tat10K<n<100K0 likes12 downloads1y agoHugging Face28LadiesMan69 /Tatar_IQA_ds0 likes11 downloads4mo agoHugging Face29saillab /alpaca_tatar_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_tatar_taco.text10K<n<100K0 likes10 downloads2y agoHugging Face30bedanar /audiobooks-tatartts__snac-24khz-tat100K<n<1M0 likes10 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.