scrubster/dr-stein-stage25-qwen15-planning-r16
scrubster/dr-stein-stage25-qwen15-planning-r16
PEFT/LoRA adapter for Qwen/Qwen2.5-1.5B-Instruct, fine-tuned on the slm-learning GAD-tool translation track.
- Trained on: 153 hand-curated
(instruction, gad CLI command)pairs - Adapter kind: lora (r=16, alpha=32)
- Target modules: qproj, kproj, vproj, oproj, gateproj, upproj, down_proj
- Run name:
stage25_qwen15_planning_r16 - Compute:
hf-jobs
Loading
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "scrubster/dr-stein-stage25-qwen15-planning-r16")Notes
Planning specialist: Qwen2.5-1.5B + LoRA r=16, trained on the reasoningtoresponse pairs extracted from the global / planning cohort. Per audit-2026-05-07-data-quality-honest.md, the global / planning cohort has 117k envelopes — the largest available signal in the whole corpus. But the prompt envelopes have no handoffid, so they cannot be joined to response envelopes by ID alone. Initial extraction via scripts/delta/preparedataset.py (strategy: reasoningtoresponse) yields 98 pairs from the 2026-05-06 export. Output at data/processed/planning-2026-05-07/. TO INCREASE PAIR YIELD before firing the run on Tier 1 remote:
- Operator runs
gad telemetry exportfrom gad-monorepo with the new handoffs/decisions/tasks adapters (commits dad42a4d / 9a8e6acd). Those adapters export the original handoff body envelopes which the planning cohort is missing — should add ~3000+ joinable pairs. - Re-run prepare_dataset.py against the fresh export.
- Re-run the planning specialist spec (this config) once train_n crosses ~1000.
Per slm-learning-079 (per-domain ladder) + slm-learning-080 (Tier 1 remote L4 ~$1.60).
