bashkir
Datasets
All datasets matching “bashkir”bashkir-frequency-index
Bashkir Frequency Index v11.5
Word-frequency index for Bashkir, computed over a large monolingual
Bashkir-language dataset, for NLP, spellchecking and lexical research.
Overview
Word-frequency index for the Bashkir language computed over a large monolingual
Bashkir-language dataset. Non-Bashkir admixture, borrowed vocabulary
and scanning artifacts were reduced with automated language filtering. The
public configuration (count ≥ 3) is the recommended default;… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.bashkir-wikipedia-parallel
Bashkir-Russian Wikipedia Parallel Corpus
Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered
for machine translation.
Overview
Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding
Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored
for semantic alignment with multilingual sentence encoders (Meta LASER3, Google
LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.bashkir-multilingual-phrasebooks
Bashkir-Russian Phrasebook Corpus
Edited Bashkir-Russian words, expressions and conversational phrases from
university phrasebooks, annotated by entry type.
Overview
Edited Bashkir-Russian pairs derived from the original
bashkorttele/trilingual-parallel-phrasebooks-bgpu
dataset, published by Bashkorttele from phrasebooks of M. Akmulla Bashkir State
Pedagogical University. The cleaned configuration is the deduplicated default;
reviewed is the edited edition… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-multilingual-phrasebooks.bashkir-wikipedia-monolingual
Bashkir Wikipedia Monolingual Corpus
Cleaned sentence-level Bashkir text from Wikipedia for pretraining, tokenizer
training and linguistic research.
Overview
Sentence-level text extracted from the Bashkir Wikipedia dump (bawiki-20260801),
cleaned and filtered with automated language identification. The cleaned
configuration is the recommended default for language modelling, tokenization and
linguistic research; precleaned is an earlier, lighter extraction kept… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-monolingual.bashkir-ngram-index
Bashkir Word N-gram Index v11.5
Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and
trigrams for spellchecking, OCR post-processing and lightweight language modelling.
Overview
Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The
release provides unigram, bigram and trigram indexes for corpus processing,
spellchecking, OCR post-processing, autocomplete and lightweight language-model
experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.bashkir-russian-parallel-corpora
Dataset Card for "bashkir-russian-parallel-corpora"
How the dataset was assembled.
find the text in two languages. it can be a translated book or an internet page (wikipedia, news site)
our algorithm tries to match Bashkir sentences with their translation in Russian
We give these pairs to people to check
@inproceedings{
title={Bashkir-Russian parallel corpora},
author={Iskander Shakirov, Aigiz Kunafin},
year={2023}
}
