metuKKhud/bashqort-raw
Bashqort Raw Corpus Description This dataset contains raw Bashkir text collected for continual training of large language models (LLMs). It is part of the project "Adapting Open-Source LLMs for the Bashkir Language", which aims to evaluate adaptation methods proposed by LlamaTurk (Toraman, 2024) and Persian adaptation (Mahdizadeh Sani et al., 2024). The corpus is assembled from multiple sources to provide a diverse linguistic foundation for language modeling.… See the full description on the dataset page: https://huggingface.co/datasets/metuKKhud/bashqort-raw.
0104
../
bash_news_articles-00000-of-00001.parquetdownload
bashgazet_articles-00000-of-00001.parquetdownload
neftcity_articles-00000-of-00001.parquetdownload
public_domain-00000-of-00001.parquetdownload
texts_bashdram-00000-of-00001.parquetdownload
texts_bashgazet-00000-of-00001.parquetdownload
texts_gsrb-00000-of-00001.parquetdownload
texts_jeshlek-00000-of-00001.parquetdownload
texts_kiskeufa-00000-of-00001.parquetdownload
texts_kulturarb-00000-of-00001.parquetdownload
texts_president_rb-00000-of-00001.parquetdownload
texts_tabin-00000-of-00001.parquetdownload
