CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01failed09 /bashkir-frequency-index Bashkir Frequency Index v11.5 Word-frequency index for Bashkir, computed over a large monolingual Bashkir-language dataset, for NLP, spellchecking and lexical research. Overview Word-frequency index for the Bashkir language computed over a large monolingual Bashkir-language dataset. Non-Bashkir admixture, borrowed vocabulary and scanning artifacts were reduced with automated language filtering. The public configuration (count ≥ 3) is the recommended default;… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.tabulartext-classification1M<n<10M0 likes240 downloads4d agoHugging Face02failed09 /bashkir-wikipedia-parallel Bashkir-Russian Wikipedia Parallel Corpus Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered for machine translation. Overview Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored for semantic alignment with multilingual sentence encoders (Meta LASER3, Google LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.tabulartranslation100K<n<1M0 likes151 downloads7d agoHugging Face03failed09 /bashkir-ngram-index Bashkir Word N-gram Index v11.5 Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and trigrams for spellchecking, OCR post-processing and lightweight language modelling. Overview Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The release provides unigram, bigram and trigram indexes for corpus processing, spellchecking, OCR post-processing, autocomplete and lightweight language-model experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.tabulartext-classification10M<n<100M0 likes105 downloads4d agoHugging Face04BashkirNLPWorld /bashkir-news-binarygated Dataset Card for Bashkir News Binary Classification Dataset Dataset Details Dataset Description This dataset contains 16,994 Bashkir-language news and analytical articles labeled for binary classification: news (label=1) vs analytics (label=0). The dataset is perfectly balanced with 8,497 examples in each class. It was created to support NLP research and applications for the Bashkir language, a low-resource Turkic language. Curated by: Arabov… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-binary.tabulartext-classification10K<n<100K0 likes26 downloads28d agoHugging Face05BashkirNLPWorld /bashkir-news-multilabelgated Dataset Card for Bashkir News Multilabel Classification Dataset Dataset Details Dataset Description This dataset contains 22,318 Bashkir-language news and analytical articles annotated with 14 thematic labels for multi-label text classification tasks. Each article can belong to several categories simultaneously. The average number of labels per article is 3.6. The dataset is designed to support NLP research and applications for the Bashkir language… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multilabel.tabulartext-classification10K<n<100K0 likes20 downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.