CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01failed09 /bashkir-frequency-index Bashkir Frequency Index v11.5 Word-frequency index for Bashkir, computed over a large monolingual Bashkir-language dataset, for NLP, spellchecking and lexical research. Overview Word-frequency index for the Bashkir language computed over a large monolingual Bashkir-language dataset. Non-Bashkir admixture, borrowed vocabulary and scanning artifacts were reduced with automated language filtering. The public configuration (count ≥ 3) is the recommended default;… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.tabulartext-classification1M<n<10M0 likes235 downloads1d agoHugging Face02failed09 /bashkir-wikipedia-parallel Bashkir-Russian Wikipedia Parallel Corpus Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered for machine translation. Overview Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored for semantic alignment with multilingual sentence encoders (Meta LASER3, Google LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.tabulartranslation100K<n<1M0 likes180 downloads5d agoHugging Face03failed09 /bashkir-multilingual-phrasebooks Bashkir-Russian Phrasebook Corpus Edited Bashkir-Russian words, expressions and conversational phrases from university phrasebooks, annotated by entry type. Overview Edited Bashkir-Russian pairs derived from the original bashkorttele/trilingual-parallel-phrasebooks-bgpu dataset, published by Bashkorttele from phrasebooks of M. Akmulla Bashkir State Pedagogical University. The cleaned configuration is the deduplicated default; reviewed is the edited edition… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-multilingual-phrasebooks.texttranslation10K<n<100K0 likes157 downloads5d agoHugging Face04failed09 /bashkir-wikipedia-monolingual Bashkir Wikipedia Monolingual Corpus Cleaned sentence-level Bashkir text from Wikipedia for pretraining, tokenizer training and linguistic research. Overview Sentence-level text extracted from the Bashkir Wikipedia dump (bawiki-20260801), cleaned and filtered with automated language identification. The cleaned configuration is the recommended default for language modelling, tokenization and linguistic research; precleaned is an earlier, lighter extraction kept… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-monolingual.texttext-generation1M<n<10M0 likes104 downloads5d agoHugging Face05failed09 /bashkir-ngram-index Bashkir Word N-gram Index v11.5 Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and trigrams for spellchecking, OCR post-processing and lightweight language modelling. Overview Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The release provides unigram, bigram and trigram indexes for corpus processing, spellchecking, OCR post-processing, autocomplete and lightweight language-model experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.tabulartext-classification10M<n<100M0 likes102 downloads1d agoHugging Face06AigizK /bashkir-russian-parallel-corpora Dataset Card for "bashkir-russian-parallel-corpora" How the dataset was assembled. find the text in two languages. it can be a translated book or an internet page (wikipedia, news site) our algorithm tries to match Bashkir sentences with their translation in Russian We give these pairs to people to check @inproceedings{ title={Bashkir-Russian parallel corpora}, author={Iskander Shakirov, Aigiz Kunafin}, year={2023} } texttranslation1M<n<10M16 likes89 downloads2y agoHugging Face07BashkirNLPWorld /bashkir-web-corpusgated Dataset Card for Bashkir Web Corpus Dataset Details Dataset Description The Bashkir Web Corpus is a collection of 71,567 documents and approximately 46.9 million tokens in the Bashkir language (a Turkic language spoken in Bashkortostan, Russia). The corpus was compiled from 16 Bashkir‑language online sources, including news websites, literary magazines, social media, books, and Wikipedia. It is designed for language modeling, text classification… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-web-corpus.texttext-generation10K<n<100K0 likes43 downloads25d agoHugging Face08BashkirNLPWorld /bashkir-wiki-corpusgated Dataset Card for Bashkir Wikipedia Corpus Dataset Details Dataset Description The Bashkir Wikipedia Corpus is a collection of 43,926 articles from Bashkir Wikipedia and Wikibooks, totaling approximately 10.6 million tokens and 8.9 million words. The data has been extracted from official Wikimedia dumps and processed to provide clean, well‑structured text suitable for NLP tasks. The corpus includes article titles, full content, categories, source… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-wiki-corpus.texttext-generation10K<n<100K0 likes43 downloads25d agoHugging Face09BashkirNLPWorld /bashkir-russian-parallelgated Dataset Card for Bashkir-Russian Parallel Corpus Dataset Details Dataset Description Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart. The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.tabulartext-generation1M<n<10M0 likes27 downloads5d agoHugging Face10BashkirNLPWorld /bashkir-news-binarygated Dataset Card for Bashkir News Binary Classification Dataset Dataset Details Dataset Description This dataset contains 16,994 Bashkir-language news and analytical articles labeled for binary classification: news (label=1) vs analytics (label=0). The dataset is perfectly balanced with 8,497 examples in each class. It was created to support NLP research and applications for the Bashkir language, a low-resource Turkic language. Curated by: Arabov… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-binary.tabulartext-classification10K<n<100K0 likes26 downloads25d agoHugging Face11BashkirNLPWorld /bashkir-news-multiclassgated Dataset Card for Bashkir News Multiclass Classification Dataset Dataset Details Dataset Description This dataset contains 17,897 Bashkir-language news and analytical articles annotated with 19 thematic categories for multiclass text classification tasks. Each article belongs to exactly one category. The categories range from news and society to culture, education, and sports. The dataset was created to support NLP research and application… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multiclass.texttext-classification10K<n<100K0 likes20 downloads25d agoHugging Face12BashkirNLPWorld /bashkir-news-multilabelgated Dataset Card for Bashkir News Multilabel Classification Dataset Dataset Details Dataset Description This dataset contains 22,318 Bashkir-language news and analytical articles annotated with 14 thematic labels for multi-label text classification tasks. Each article can belong to several categories simultaneously. The average number of labels per article is 3.6. The dataset is designed to support NLP research and applications for the Bashkir language… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multilabel.tabulartext-classification10K<n<100K0 likes20 downloads25d agoHugging Face13BashkirNLPWorld /bashkir-news-clustergated Dataset Card for Bashkir News Cluster Dataset Dataset Details Dataset Description This dataset contains 24,428 Bashkir-language news and analytical articles collected from various online sources. It is intended for clustering, representation learning, and unsupervised NLP tasks. Each text is accompanied by metadata such as title, source, date, and original category. The corpus is part of the BashkirNLP project, aiming to support low-resource… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-cluster.textfeature-extraction10K<n<100K0 likes20 downloads25d agoHugging Face14BashkirNLPWorld /bashkir-lexicongated Dataset Card for Bashkir Lexicon (Machine Fund) Dataset Details Dataset Description Bashkir Lexicon (Machine Fund) is a machine-readable lexical database of the Bashkir language, compiled from the open digital resource Machine Fund of the Bashkir Language (mfbl2.ru). It contains 34,008 unique lexical entries covering dialectal word forms with their part of speech, dialect, subdialect, literary norm, and Russian translation. The dataset preserves… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-lexicon.texttext-generation10K<n<100K0 likes19 downloads8d agoHugging Face15DmitriyV /bashkir-russian-parallel-corpora-filtered Dataset Card for "bashkir-russian-parallel-corpora" How the dataset was assembled. find the text in two languages. it can be a translated book or an internet page (wikipedia, news site) our algorithm tries to match Bashkir sentences with their translation in Russian We give these pairs to people to check @inproceedings{ title={Bashkir-Russian parallel corpora}, author={Iskander Shakirov, Aigiz Kunafin}, year={2023} } texttranslation1M<n<10M0 likes16 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.