CoolFace
Datasetpublic

khairi/uniref50-replay-mix-v1

uniref50-replay-mix-v1 Stage-1 continued-pretraining corpus for eshmun-vocab: protein sequences (UniRef50) mixed with a general/biomedical/math/code text replay slice, so a vocab-extended LLM (e.g. khairi/qwen3-0.6b-protein-vocab-v0) learns protein-sequence statistics without catastrophically forgetting its pretrained language ability. Design rationale and target ratios: see docs/pretrain-dataset-mix.md in the eshmun-vocab repo. Objective: plain next-token prediction. Every row… See the full description on the dataset page: https://huggingface.co/datasets/khairi/uniref50-replay-mix-v1.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes134downloads
5 commits on main
28d04773mo ago

Update dataset card for full 10.68M-row build

khairi
d1f7c713mo ago

Upload dataset

khairi
ef1cde23mo ago

Add dataset card

khairi
ebcf70d3mo ago

Upload dataset

khairi
d069fcb3mo ago

initial commit

khairi