CoolFace
19 results

bashkir

failed09 /bashkir-frequency-index Bashkir Frequency Index v11.5 Word-frequency index for Bashkir, computed over a large monolingual Bashkir-language dataset, for NLP, spellchecking and lexical research. Overview Word-frequency index for the Bashkir language computed over a large monolingual Bashkir-language dataset. Non-Bashkir admixture, borrowed vocabulary and scanning artifacts were reduced with automated language filtering. The public configuration (count ≥ 3) is the recommended default;… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.tabulartext-classification1M<n<10M0 likes235 downloads1d agoHugging Facefailed09 /bashkir-wikipedia-parallel Bashkir-Russian Wikipedia Parallel Corpus Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered for machine translation. Overview Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored for semantic alignment with multilingual sentence encoders (Meta LASER3, Google LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.tabulartranslation100K<n<1M0 likes180 downloads5d agoHugging Facefailed09 /bashkir-multilingual-phrasebooks Bashkir-Russian Phrasebook Corpus Edited Bashkir-Russian words, expressions and conversational phrases from university phrasebooks, annotated by entry type. Overview Edited Bashkir-Russian pairs derived from the original bashkorttele/trilingual-parallel-phrasebooks-bgpu dataset, published by Bashkorttele from phrasebooks of M. Akmulla Bashkir State Pedagogical University. The cleaned configuration is the deduplicated default; reviewed is the edited edition… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-multilingual-phrasebooks.texttranslation10K<n<100K0 likes157 downloads5d agoHugging Facefailed09 /bashkir-wikipedia-monolingual Bashkir Wikipedia Monolingual Corpus Cleaned sentence-level Bashkir text from Wikipedia for pretraining, tokenizer training and linguistic research. Overview Sentence-level text extracted from the Bashkir Wikipedia dump (bawiki-20260801), cleaned and filtered with automated language identification. The cleaned configuration is the recommended default for language modelling, tokenization and linguistic research; precleaned is an earlier, lighter extraction kept… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-monolingual.texttext-generation1M<n<10M0 likes104 downloads5d agoHugging Facefailed09 /bashkir-ngram-index Bashkir Word N-gram Index v11.5 Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and trigrams for spellchecking, OCR post-processing and lightweight language modelling. Overview Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The release provides unigram, bigram and trigram indexes for corpus processing, spellchecking, OCR post-processing, autocomplete and lightweight language-model experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.tabulartext-classification10M<n<100M0 likes102 downloads1d agoHugging FaceAigizK /bashkir-russian-parallel-corpora Dataset Card for "bashkir-russian-parallel-corpora" How the dataset was assembled. find the text in two languages. it can be a translated book or an internet page (wikipedia, news site) our algorithm tries to match Bashkir sentences with their translation in Russian We give these pairs to people to check @inproceedings{ title={Bashkir-Russian parallel corpora}, author={Iskander Shakirov, Aigiz Kunafin}, year={2023} } texttranslation1M<n<10M16 likes89 downloads2y agoHugging Face