hyper-efficient-system-llc/bosnian-corpus-v1
Bosnian Corpus v1.0 (cleaned) This dataset provides a cleaned and genre-annotated corpus of contemporary Bosnian, designed for quantitative linguistic analysis, information-theoretic research, entropy estimation, corpus linguistics, language modeling, and modern NLP tasks. The canonical release of the corpus is archived on Zenodo: Dataset DOI:https://doi.org/10.5281/zenodo.17757098 Corpus composition The corpus is constructed from three publicly available… See the full description on the dataset page: https://huggingface.co/datasets/hyper-efficient-system-llc/bosnian-corpus-v1.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face