CoolFace
Datasetpublic

hyper-efficient-system-llc/bosnian-corpus-v1

Bosnian Corpus v1.0 (cleaned) This dataset provides a cleaned and genre-annotated corpus of contemporary Bosnian, designed for quantitative linguistic analysis, information-theoretic research, entropy estimation, corpus linguistics, language modeling, and modern NLP tasks. The canonical release of the corpus is archived on Zenodo: Dataset DOI:https://doi.org/10.5281/zenodo.17757098 Corpus composition The corpus is constructed from three publicly available… See the full description on the dataset page: https://huggingface.co/datasets/hyper-efficient-system-llc/bosnian-corpus-v1.

sourceHugging Facecc-by-sa-4.0updated 1d agoView on Hugging Face
1likes71downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
hyper-efficient-system-llc/bosnian-corpus-v1 · CoolFace