CoolFace
Modelpublic

scrubster/dr-stein-stage25-qwen15-planning-r16

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes5downloads
Model Card

scrubster/dr-stein-stage25-qwen15-planning-r16

PEFT/LoRA adapter for Qwen/Qwen2.5-1.5B-Instruct, fine-tuned on the slm-learning GAD-tool translation track.

  • —Trained on: 153 hand-curated (instruction, gad CLI command) pairs
  • —Adapter kind: lora (r=16, alpha=32)
  • —Target modules: qproj, kproj, vproj, oproj, gateproj, upproj, down_proj
  • —Run name: stage25_qwen15_planning_r16
  • —Compute: hf-jobs

Loading

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "scrubster/dr-stein-stage25-qwen15-planning-r16")

Notes

Planning specialist: Qwen2.5-1.5B + LoRA r=16, trained on the reasoningtoresponse pairs extracted from the global / planning cohort. Per audit-2026-05-07-data-quality-honest.md, the global / planning cohort has 117k envelopes — the largest available signal in the whole corpus. But the prompt envelopes have no handoffid, so they cannot be joined to response envelopes by ID alone. Initial extraction via scripts/delta/preparedataset.py (strategy: reasoningtoresponse) yields 98 pairs from the 2026-05-06 export. Output at data/processed/planning-2026-05-07/. TO INCREASE PAIR YIELD before firing the run on Tier 1 remote:

  1. 1.Operator runs gad telemetry export from gad-monorepo with the new handoffs/decisions/tasks adapters (commits dad42a4d / 9a8e6acd). Those adapters export the original handoff body envelopes which the planning cohort is missing — should add ~3000+ joinable pairs.
  2. 2.Re-run prepare_dataset.py against the fresh export.
  3. 3.Re-run the planning specialist spec (this config) once train_n crosses ~1000.

Per slm-learning-079 (per-domain ladder) + slm-learning-080 (Tier 1 remote L4 ~$1.60).