CoolFace
Modelpublic

mehul79/qwen35-4b-orchestrator-lora

sourceHugging Faceupdated 29d agoView on Hugging Face
1likes48downloads
Model Card

Model Card for qwen35-qlora-v4

QLoRA adapter for unsloth/Qwen3.5-4B, fine-tuned as a tool-routing / orchestration policy: given a system prompt, a tool roster, and conversation history, it picks the next tool call or response. Fine-tuned with Unsloth.

Supersedes `v3`: lower LoRA rank, added dropout, and a single-epoch schedule to fix the overfitting v3 showed past epoch 1. This is the main revision.

Base model: the post-trained Qwen3.5-4B, not the pretrained base

unsloth/Qwen3.5-4B is Qwen's instruction-tuned release, a mirror of Qwen/Qwen3.5-4B. It is not Qwen/Qwen3.5-4B-Base, which is the pretrained-only checkpoint. The -Base suffix is the whole distinction in Qwen's naming, and loading the wrong one silently changes what this adapter means.

Three layers stack here:

Qwen3.5-4B pretrained
  └─ Qwen's post-training: ChatML turns, <think> reasoning, nested tool-call XML
       └─ this adapter: tool-routing policy and platform conventions

The training data depends entirely on the middle layer. Records are rendered by the base model's own chat template with the tool schemas passed as tools=, so the <|im_start|> turn markers, the # Tools system block, the <think> blocks, and the nested <tool_call><function=...><parameter=...></function></tool_call> XML all come from that template. No code in the training pipeline writes that syntax by hand. Run the same corpus against -Base and none of that scaffolding exists, so the adapter would have to learn the entire chat and tool-call format from scratch.

What the adapter adds on top is the target platform's conventions. Measured on a live vLLM server at temperature 0, same prompt and the same 22 tool schemas, the fine-tuned adapter emits Read{file_path, activity} while the un-adapted base emits Read{file_path}. The activity argument exists only in the training corpus.

Run summary

  • —LoRA r=64, alpha=128, dropout=0.05, 1 epoch, lr=1e-4
  • —3,601 records / 10,889 trainable assistant spans, split 3,417 train / 184 val
  • —Eval loss fell every single eval step, ending below train loss with no sign of divergence
  • —load_best_model_at_end picked the final step (214) as best, so the adapter on the Hub is the end-of-run weights
stepepochtrain_losseval_loss
250.121.0751.081
500.230.9521.017
1000.470.9210.970
1500.700.9090.951
1750.820.7930.947
214 (final)1.000.917 (run avg)0.944 (best)

Not yet scored against the frozen golden_v2.0 eval set (pre_eval/) as of this push — treat these numbers as training-time signal only, not a promotion decision.

Quick start

python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained("unsloth/Qwen3.5-4B", device_map="cuda")
model = PeftModel.from_pretrained(base, "mehul79/qwen35-4b-orchestrator-lora")
tokenizer = AutoTokenizer.from_pretrained("mehul79/qwen35-4b-orchestrator-lora")

Serving

Served as an adapter on top of the base, not merged. Verified end to end on vLLM 0.28.0:

bash
vllm serve unsloth/Qwen3.5-4B \
  --enable-lora --lora-modules qwen35=<adapter dir> --max-lora-rank 64 \
  --enable-auto-tool-choice --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --dtype bfloat16 --max-model-len 32768

Four things that are easy to get wrong:

  • —Use the XML tool parser. Qwen3.5 emits nested <function=...> XML, not Hermes JSON. A Hermes parser does not raise an error. It returns tool_calls: null and leaks the raw XML into message.content, which looks like a model that never calls a tool.
  • —Leave thinking on. The adapter was trained with <think> content on the trained turn, so passing enable_thinking: false is train/serve skew. Send a generous max_tokens as well, or reasoning eats the budget and the reply returns a null content with finish_reason: length.
  • —`--max-model-len 32768`. The fixed prefix, 22 tool schemas plus the system prompt, measures about 18.6k tokens on its own. An 18k window cannot hold a single request.
  • —Read `message.reasoning`. vLLM 0.28.0 puts the reasoning trace there, not in message.reasoning_content. The two names appear interchangeably across versions, and reading the wrong one gives an empty string with no error.

Training procedure

QLoRA SFT via Unsloth, with loss masked to assistant-turn spans only (the system prompt, tool schemas, and user/tool-result turns carry no gradient).

Framework versions

  • —Unsloth
  • —PEFT 0.20.0
  • —Transformers: 5.5.0
  • —Pytorch: 2.11.0+cu128
  • —Datasets: 4.3.0
  • —Tokenizers: 0.22.2

Citation

bibtex
@misc{unsloth,
  title  = {Unsloth},
  author = {Daniel Han and Michael Han and Unsloth team},
  url    = {http://github.com/unslothai/unsloth}
}