CoolFace
Datasetpublic

metuKKhud/bashqort-raw

Bashqort Raw Corpus Description This dataset contains raw Bashkir text collected for continual training of large language models (LLMs). It is part of the project "Adapting Open-Source LLMs for the Bashkir Language", which aims to evaluate adaptation methods proposed by LlamaTurk (Toraman, 2024) and Persian adaptation (Mahdizadeh Sani et al., 2024). The corpus is assembled from multiple sources to provide a diverse linguistic foundation for language modeling.… See the full description on the dataset page: https://huggingface.co/datasets/metuKKhud/bashqort-raw.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes104downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

metuKKhud/bashqort-raw · main · files are served by the source, never re-hosted here