CoolFace
Modelpublic

lsteno/Qwen3-4B-Instruct-2507-RLM-RL-depth1-r4-a8-lr5e-7-s150-lora

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes11downloads
Model Card

Qwen3-4B RLM RLVR Depth-1 LoRA Adapter

LoRA adapter from the first rank/LR ablation run.

  • —Base model: Qwen/Qwen3-4B-Instruct-2507
  • —Run id: rlm-rlvr-qwen3-4b-depth1-llmonly-r4-a8-lr5e-7-s150
  • —Prompt variant: sanjaya_text_depth1_llm_only_v1
  • —Runtime depth: depth-1 LLM-only orchestration (max_depth = 0, plain llm_query subcalls enabled)
  • —LoRA rank: 4
  • —LoRA alpha: 8
  • —Learning rate: 5e-7
  • —Training steps: 150
  • —Final adapter source: run_default/broadcasts/step_150

The run_configs/ directory contains the exact trainer and orchestrator TOML files saved with the run.