sup2ch/ru-en-code-curriculum
RuEn Code Curriculum RuEn Code Curriculum is a curated Russian-English dataset for continued pretraining (CPT) and supervised fine-tuning (SFT) of small code-oriented language models. This public release contains only records classified as redistributable. Local-training-only web and code sources used by the internal curriculum are intentionally excluded. Dataset summary Configuration Split Records Tokens sft train 53,278 13,997,239 sft reserve 19,225… See the full description on the dataset page: https://huggingface.co/datasets/sup2ch/ru-en-code-curriculum.
This repository belongs to sup2ch on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
