davidanugraha/R-Few-Qwen3.5-35B-A3B-SWE-Smith-r18
R-Few Qwen3.5-35B-A3B SWE mechanics campaign r18
This public repository is a portability snapshot of the separate R-Few Challenger and Solver checkpoint lineages from rfew-qwen35-35b-all198-counter-fix-20260904-r18.
It is not a completed R-Few effectiveness claim. Cycle 7 completed without an optimizer update and its terminal data is included. Skipped cycles do not create fake checkpoint numbers: checkpoint steps count successful optimizer updates, not outer-loop cycles.
Checkpoint lineage
Current continuation parents in this snapshot are Challenger step 4 and Solver step 2. Cycle 7 attempted 32 Challenger groups: 30 rectangular groups yielded 240 trajectories, while two groups were discarded after workspace-contamination errors. No valid synthetic task was produced. The Challenger replay had no within-group reward variance, and no mixed human/synthetic Solver replay was emitted.
Included directories
checkpoints/challenger/step-1/
checkpoints/challenger/step-2/
checkpoints/challenger/step-3/
checkpoints/challenger/step-4/
checkpoints/solver/step-1/
checkpoints/solver/step-2/
metadata/configs/
metadata/docs/
data/cycle-0007/challenger-groups/
data/cycle-0007/backbone/
data/cycle-0007/diagnostics/
data/cycle-0007/resume_cycle7_summary.jsonEach checkpoint directory contains the serving LoRA adapter, committed receipt, diagnostic metrics, and the compact FSDP continuation checkpoint with optimizer, scheduler, and RNG state. The Qwen base-model weights are not included.
Configuration represented by this snapshot
- Base:
Qwen3.5-35B-A3Bat immutable revision59d61f3ce65a6d9863b86d2e96597125219dc754. - Challenger: 32 prompt groups per cycle, K=8, GRPO, LoRA rank 32 / alpha 64, learning rate
1e-5, 114,688-token trainable context. - Solver: K=8, optional SWE-Gym RLOO preset, LoRA rank 32 / alpha 64, learning rate
1e-5, 65,536-token context. - Difficulty band:
[0.2, 0.7], which admits 2--5 successes at K=8. - Generation universe: all 198 broad-ready SWE-Smith repository worlds.
- Challenger prompts sample 0--5 human examples.
Important human-data limitation
The canonical source manifest contains 4,549 eligible SWE-Smith tasks, but the r18 mechanics config materializes only five unique human tasks total from 16 validated candidates. Those five are reused for Challenger conditioning and are the entire human side of every Solver difficulty pass.
This should not be confused with the paper's separate 0--5 per-prompt anchor sampling rule. Consequently, r18 is valid mechanics evidence but not the final full-human-budget comparison. A future comparison must define a sealed 4,549-task pool, sample only 0--5 of that pool into each Challenger prompt, and separately specify a fixed-budget human Solver schedule comparable across R-Few, Basic GRPO, and the proposed method.
The current implementation also withholds a Solver update unless both human and synthetic tasks survive in the same cycle. This is a conservative SWE adaptation, not a rule inherited from released R-Zero. Released R-Zero instead filters a variable task pool, trains fixed-size batches, drops incomplete final batches, and has no human/synthetic composition constraint.
Storage boundary
This repository preserves model continuation state and compact diagnostics. The larger campaign directory from Cycles 0--6 is intentionally archived separately to S3. Cycle 7's sealed Challenger groups, backbone, replay, audits, and summary are included here; repository workspaces and unrelated raw campaign artifacts are not. Modal also retains the original training artifact prefixes. See metadata/docs/MACHINE_MIGRATION.md for locations, hashes, and recovery steps.
Source-code warning
The campaign ran from an uncommitted working tree. The current checkout changed after Cycle 7 launched, so this checkpoint release is not a substitute for the curated Git commit described in the migration handoff. Do not claim exact source reproduction until that source closure is committed.
