CoolFace
Modelpublic

Lego-X/qwen3_5_35b_a3b_cc_200k_rl

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
1likes186downloads
Model Card

Lego-RL-Qwen3.5-35B-A3B · Claude Code · 200K

<p align="center"> <a href="https://arxiv.org/abs/2608.17393">📖 Paper</a> • <a href="https://github.com/LegoX/Lego-RL">🧑‍💻 Code</a> • <a href="https://lego-rl.pages.dev">📚 Docs</a> • <a href="https://huggingface.co/collections/Lego-X/lego-rl">🤗 Models</a> • <a href="https://huggingface.co/datasets/Lego-X/Lego-RL-2699">🤗 Data</a> • <a href="https://legox.net">🏠 LegoX</a> </p>

*[Qwen3.5-35B-A3B](https://huggingface.co/Qwen/Qwen3.5-35B-A3B) trained with online RL inside the unmodified Claude Code harness, at 200K context, on 2,699 real repository issues whose own test suites produce the reward.*

SWE-bench Verified: 62.4 → 68.2 (+5.8) — no reward model, no reference-patch similarity, no harness rewrite.

This checkpoint is the Claude Code production run of Lego-RL (Faithful · Reliable · Observable), released as training step 110. The OpenHands SDK counterpart is `Lego-X/qwen3_5_35b_a3b_ohsdk_200k_rl`; the OpenCode run is `Lego-X/qwen3_5_35b_a3b_oc_200k_rl`.

The agent solves a real issue in a real repository inside a fresh sandbox, the task's own verifier suite decides {0, 1}, and the trajectory the harness actually produced — token ids, masks, log-probs and MoE expert routes captured inside the serving path — becomes the gradient step.

Why harness-native training

The scaffold is part of the environment, not the policy. The same weights score very differently depending on which harness runs them, so training under a rewritten control flow optimizes for a deployment you never ship:

ModelOpenHands SDKClaude CodeOpenCode
Qwen3.5-35B-A3B (starting point)64.062.457.2
Qwen3.6-35B-A3B (next-gen base)67.463.460.6
KAT-Coder-V2.5-Dev (post-trained Qwen3.6)67.066.864.8
Lego-RL-Qwen3.5-35B-A3B70.468.266.6

SWE-bench Verified (%), one shared protocol: temperature 0.7, 200 turns, 200K context.

Each Lego-RL column is a separate run trained in that harness from the same starting checkpoint, the same 2,699 tasks and the same 3 epochs. This repository is the Claude Code run (68.2). Across the three harnesses RL adds +6.4 / +5.8 / +9.4.

Quick start

1. Serve with vLLM

bash
vllm serve Lego-X/qwen3_5_35b_a3b_cc_200k_rl \
    --served-model-name vllm_model \
    --tensor-parallel-size 4 \
    --enable-expert-parallel \
    --max-model-len 262144 \
    --gpu-memory-utilization 0.9 \
    --enable-chunked-prefill --enable-prefix-caching \
    --dtype bfloat16 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --host 0.0.0.0 --port 8000
[!IMPORTANT] --tool-call-parser qwen3_coder is not optional. The model was rolled out and trained with this parser; serving it behind hermes (or any other parser) silently degrades tool-call formatting.

2. Drive it with Claude Code

Point Claude Code at the vLLM endpoint as an Anthropic-compatible provider, and give it a repository workspace plus a 200-turn / 200K budget. The RL policy learned to spend turns; short turn caps systematically truncate the second half of its trajectories and cost most of the gain.

3. Reproduce the evaluation

bash
git clone https://github.com/LegoX/Lego-RL.git && cd Lego-RL
bash scripts/setup_env.sh
cp scripts/eval/_template.env scripts/eval/configs/my_eval.env   # set MODEL_PATH, DATASET_PATH, kubeconfig
bash scripts/eval/eval.sh scripts/eval/configs/my_eval.env

Sandboxed execution and verifier rewards come from Harbor; see the evaluation docs.

Training

Starting checkpointQwen/Qwen3.5-35B-A3B (sparse MoE, 256 experts, 8 active)
HarnessClaude Code, unmodified — a thin adapter, not a fork
Released stepglobal_step_110
TasksLego-X/Lego-RL-2699 — 2,699 real repository issues, converted from GAIR/OpenSWE
Rewardeach task's own test suite, run in a fresh sandbox: {0, 1}. No reward model, no patch similarity, no LLM judge
AlgorithmGSPO (sequence-level surrogate), group-relative advantage over G = 8 rollouts per task
Batch64 prompts × 8 responses = 512 trials/step; 3 epochs = 126 steps
Optimizerlr 1e-6 constant, KL loss 1e-3 (low-var), clip [3e-4, 4e-4], rollout temperature 1.0
Context200K, 200 turns
BackendVeOmni FSDP, Ulysses SP = 8, R3 rollout routing replay; fully-async with partial rollout (staleness 1)
Hardware3 nodes × 8 GPUs (2 training + 1 rollout)

The training set is disjoint from SWE-bench Verified at both the repository and the instance level.

Intended use and limitations

Use it as an agent policy, not as a chat model: it was optimized inside a harness that hands it a repository, a shell, and file-editing tools.

  • —Harness. Trained in Claude Code. A separate policy is released for OpenHands SDK and OpenCode; each does best in the harness it was trained in.
  • —Budget. 200K context and 200 turns. Short budgets truncate it.
  • —Domain. Python-heavy repository issue-resolution, in the SWE-bench/OpenSWE distribution.
  • —Inherited base behavior. Safety, multilingual and general-knowledge behavior come from Qwen3.5-35B-A3B and were not targeted by this RL.
  • —Sandbox it. The policy writes files and executes shell commands on purpose. Run it in a container.

Related work in the LegoX series

Lego-RLharness-native RL for coding agents (this model)
SWE-Legothe SFT recipe
SWE-Reviewinference-time generate-review-revise
Terminal-Legotrajectory-quality filtering

Acknowledgement

Built on verl (trainer + rollout) and Harbor (sandboxed execution + verifier reward), with Claude Code as the harness.

Citation

bibtex
@misc{du2026legorlharnessnativereinforcementlearning,
  title={LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents},
  author={Yiming Du and Yuxin Jiang and Tao Yuan and Jianbo Dai and Shaowei Wang and Jierun Chen and Chaofan Tao and Xianzhi Yu and Lifeng Shang and Kam-Fai Wong and Xiaohui Li and Haoli Bai},
  year={2026},
  eprint={2608.17393},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2608.17393},
}