khairi/uniref50-replay-mix-v1
uniref50-replay-mix-v1 Stage-1 continued-pretraining corpus for eshmun-vocab: protein sequences (UniRef50) mixed with a general/biomedical/math/code text replay slice, so a vocab-extended LLM (e.g. khairi/qwen3-0.6b-protein-vocab-v0) learns protein-sequence statistics without catastrophically forgetting its pretrained language ability. Design rationale and target ratios: see docs/pretrain-dataset-mix.md in the eshmun-vocab repo. Objective: plain next-token prediction. Every row… See the full description on the dataset page: https://huggingface.co/datasets/khairi/uniref50-replay-mix-v1.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face