CoolFace
Datasetpublic

metuKKhud/bashqort-raw

Bashqort Raw Corpus Description This dataset contains raw Bashkir text collected for continual training of large language models (LLMs). It is part of the project "Adapting Open-Source LLMs for the Bashkir Language", which aims to evaluate adaptation methods proposed by LlamaTurk (Toraman, 2024) and Persian adaptation (Mahdizadeh Sani et al., 2024). The corpus is assembled from multiple sources to provide a diverse linguistic foundation for language modeling.… See the full description on the dataset page: https://huggingface.co/datasets/metuKKhud/bashqort-raw.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes104downloads
../
filebash_news_articles-00000-of-00001.parquet36.7 MBdownload
filebashgazet_articles-00000-of-00001.parquet1.4 MBdownload
fileneftcity_articles-00000-of-00001.parquet441 KBdownload
filepublic_domain-00000-of-00001.parquet65 KBdownload
filetexts_bashdram-00000-of-00001.parquet575 KBdownload
filetexts_bashgazet-00000-of-00001.parquet83.3 MBdownload
filetexts_gsrb-00000-of-00001.parquet425 KBdownload
filetexts_jeshlek-00000-of-00001.parquet18.5 MBdownload
filetexts_kiskeufa-00000-of-00001.parquet163 KBdownload
filetexts_kulturarb-00000-of-00001.parquet1.8 MBdownload
filetexts_president_rb-00000-of-00001.parquet2.9 MBdownload
filetexts_tabin-00000-of-00001.parquet766 KBdownload

metuKKhud/bashqort-raw · main · files are served by the source, never re-hosted here