CoolFace
Datasetpublic

fan-shu/swe-davinci-ctx-midtrain

fan-shu/swe-davinci-ctx-midtrain daVinci-Dev ctx-native (PR-derived) mid-training data, linearized for training Qwen3-style code models. Rendered from GAIR/daVinci-Dev ctx-native/llm_enhanced_prs following the daVinci Task-5 Markdown layout (Repository Context / Issue / Pull Request / Relevant Files Found / Edits with search-replace blocks). Configs cpt: continued-pretraining. messages = [assistant: <full PR document>]; train on the whole document (loss_mask… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-davinci-ctx-midtrain.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes129downloads
Dataset Card

fan-shu/swe-davinci-ctx-midtrain

daVinci-Dev ctx-native (PR-derived) mid-training data, linearized for training Qwen3-style code models. Rendered from GAIR/daVinci-Dev ctx-native/llm_enhanced_prs following the daVinci Task-5 Markdown layout (Repository Context / Issue / Pull Request / Relevant Files Found / Edits with search-replace blocks).

Configs

  • cpt: continued-pretraining. messages = [assistant: <full PR document>]; train on the whole document (loss_mask assistant=true).
  • sft: SFT. messages = [user: <context>, assistant: <edits>]; loss on the edits only.

Source: https://huggingface.co/datasets/GAIR/daVinci-Dev (paper: https://arxiv.org/abs/2601.18418)