CoolFace
Datasetpublic

DimitarV/bulgarian-medical-cpt-100m

Bulgarian text for MOSS continued pretraining Exactly 100 million training tokens: 10M medical and 90M general Bulgarian. An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers. No model training has been performed as part of this dataset build. Medical data Exactly 10,000,000 training tokens and 50,000 additional validation tokens, including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.

sourceHugging Faceotherupdated 16d agoView on Hugging Face
0likes88downloads
settings

This repository belongs to DimitarV on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namebulgarian-medical-cpt-100m
visibilitypublic
licenceother
gatedno
ownerDimitarV
Account settings
DimitarV/bulgarian-medical-cpt-100m · CoolFace