CoolFace
Datasetpublic

DimitarV/bulgarian-medical-cpt-100m

Bulgarian text for MOSS continued pretraining Exactly 100 million training tokens: 10M medical and 90M general Bulgarian. An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers. No model training has been performed as part of this dataset build. Medical data Exactly 10,000,000 training tokens and 50,000 additional validation tokens, including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.

sourceHugging Faceotherupdated 16d agoView on Hugging Face
0likes88downloads
2 commits on main
aca712116d ago

Add 100M-token Bulgarian CPT option: 90M general and 10M medical

DimitarV
5d0c2c516d ago

initial commit

DimitarV