CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TatarNLPWorld /tatar-web-corpusgated Dataset Card for Tatar Web Corpus Dataset Details Dataset Description The largest open corpus for the Tatar language with over 1 million documents collected from news websites, social media, articles, books, and Wikipedia. Designed for various NLP tasks including language modeling, text classification, information extraction, and search. Curated by: TatarNLPWorld Community Language(s) (NLP): Tatar (tt) License: other – see Licensing & Legal Notice… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus.texttext-classification1M<n<10M0 likes34 downloads26d agoHugging Face02TatarNLPWorld /tatar-news-analysis-multilabelgated Dataset Card for Tatar News Multilabel Classification Dataset Details Dataset Description The Tatar News Multilabel Classification Dataset contains 55,709 Tatar language news articles annotated with 13 distinct topic labels in a multi-label setting (each article can have multiple labels). Each entry includes the full article content, title, label indices, multi-hot label vector, number of labels, original single category, source URL, publication… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multilabel.tabulartext-classification10K<n<100K0 likes34 downloads26d agoHugging Face03TatarNLPWorld /tatar-english-russian-corpusgated Dataset Card: Tatar-English-Russian Parallel Corpus Dataset Details Dataset Description This dataset is a parallel corpus containing 14,983 sentences in three languages: Tatar, English, and Russian. It combines two distinct sources: KickItLikeShika/english-tatar-translation (7,746 entries) – an existing English-Tatar dataset with Russian translations added yasalma/tt-en-language-corpus (7,615 entries) – another English-Tatar corpus with newly added… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-english-russian-corpus.tabulartranslation10K<n<100K0 likes32 downloads29d agoHugging Face04TatarNLPWorld /tatar-wiki-corpusgated Dataset Card for Tatar Wiki Corpus Dataset Details Dataset Description A comprehensive cleaned corpus of Tatar Wikipedia and Wikibooks with over 467,000 articles. This dataset is ideal for training language models, text classification, information retrieval, and various NLP tasks for the Tatar language. Curated by: TatarNLPWorld Community Language(s) (NLP): Tatar (tt) License: cc-by-sa-4.0 – see Licensing & Legal Notice below.… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-wiki-corpus.tabulartext-classification100K<n<1M0 likes25 downloads26d agoHugging Face05TatarNLPWorld /tatar-web-corpus-v3gated Dataset Card for Tatar Web Corpus Dataset Details Dataset Description The Tatar Web Corpus is the largest open-source corpus of the Tatar language (Turkic family), containing 2,465,867 documents (approximately 251 million tokens) collected from publicly available web sources. It covers news portals, social media, blogs, literary websites, and other domains. The corpus underwent soft deduplication to remove exact duplicates while preserving… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus-v3.texttext-generation1M<n<10M0 likes25 downloads26d agoHugging Face06TatarNLPWorld /tatar-news-clustergated Dataset Card for Tatar News Clustered Dataset Dataset Details Dataset Description The Tatar News Clustered Dataset is a comprehensive collection of 57,340 Tatar language news articles with topic categories, curated by TatarNLPWorld as part of the Tat2Vec project. The dataset includes full article content, titles, source names, publication dates, and 282 topic categories. It is designed for multi-class text classification, topic modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-cluster.texttext-classification10K<n<100K0 likes24 downloads26d agoHugging Face07TatarNLPWorld /tatar-news-analysis-multiclassgated Dataset Card for Tatar News Multiclass Classification Dataset Details Dataset Description The Tatar News Multiclass Classification Dataset contains 86,963 Tatar language news articles classified into 9 distinct topic categories. Each entry includes the full article content, title, category (numeric label and text label), source URL, publication date, and content length. The dataset is specifically designed for training and evaluating multi-class… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multiclass.tabulartext-classification10K<n<100K0 likes22 downloads26d agoHugging Face08TatarNLPWorld /tatar-folklore-corpusgated Dataset Card for Tatar Text Corpus with Rich Metadata Dataset Details Dataset Description This dataset is a collection of 326 Tatar language texts with extensive metadata, curated by TatarNLPWorld. Each record includes the full text and a structured metadata object containing fields such as title, author, year, source, genre, category, and more. The dataset is designed for NLP research on the Tatar language, including text classification, language… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-folklore-corpus.texttext-generationn<1K0 likes22 downloads26d agoHugging Face09eldiablo92 /bulak-ultimate-tatar-corpusgated Ultimate cleaned Tatar Corpus The training corpus of the boolak project (a Tatar language model trained from scratch), frozen right before tokenization. One row = one document. The text is cleaned, filtered and deduplicated, but not tokenized: this is the exact input of to_ids.py. Version: v20260824 — built on 2026-08-24 from datasets import load_dataset ds = load_dataset("eldiablo92/bulak-ultimate-tatar-corpus", split="train", revision="v20260824") # pin the version ds =… See the full description on the dataset page: https://huggingface.co/datasets/eldiablo92/bulak-ultimate-tatar-corpus.tabulartext-generation1M<n<10M0 likes20 downloads26d agoHugging Face10TatarNLPWorld /vk-groupsgated VK Groups Dataset: Tatar Social Media Posts Dataset Description This dataset contains posts from five popular VKontakte (VK) communities targeting Tatar-speaking audiences. It includes content across different themes: religion, relationships, humor, advice, and cooking. The dataset is designed for various NLP tasks including text classification, sentiment analysis, topic modeling, and engagement prediction. Curated by: TatarNLPWorld Language(s): Tatar (Cyrillic script)… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/vk-groups.tabulartext-classification100K<n<1M0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.