CoolFace
Datasetpublic

hasankursun/bulgarian-corpus-33b

Bulgarian Corpus 33B (BC-33B) Dataset Summary The Bulgarian Corpus 33B (BC-33B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Bulgarian. Comprising approximately 33.4 Billion tokens (measured with Qwen 2.5/Llama-3 tokenizer), it represents one of the largest open-source resources for Bulgarian LLM pretraining. The dataset is engineered for a modern two-stage training pipeline: Pretrain Subset (~29.3B Tokens): A… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/bulgarian-corpus-33b.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
6likes711downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
hasankursun/bulgarian-corpus-33b · CoolFace