mehul79/qwen35-4b-orchestrator-lora
Model Card for qwen35-qlora-v4
QLoRA adapter for unsloth/Qwen3.5-4B, fine-tuned as a tool-routing / orchestration policy: given a system prompt, a tool roster, and conversation history, it picks the next tool call or response. Fine-tuned with Unsloth.
Supersedes `v3`: lower LoRA rank, added dropout, and a single-epoch schedule to fix the overfitting v3 showed past epoch 1. This is the main revision.
Base model: the post-trained Qwen3.5-4B, not the pretrained base
unsloth/Qwen3.5-4B is Qwen's instruction-tuned release, a mirror of Qwen/Qwen3.5-4B. It is not Qwen/Qwen3.5-4B-Base, which is the pretrained-only checkpoint. The -Base suffix is the whole distinction in Qwen's naming, and loading the wrong one silently changes what this adapter means.
Three layers stack here:
Qwen3.5-4B pretrained
└─ Qwen's post-training: ChatML turns, <think> reasoning, nested tool-call XML
└─ this adapter: tool-routing policy and platform conventionsThe training data depends entirely on the middle layer. Records are rendered by the base model's own chat template with the tool schemas passed as tools=, so the <|im_start|> turn markers, the # Tools system block, the <think> blocks, and the nested <tool_call><function=...><parameter=...></function></tool_call> XML all come from that template. No code in the training pipeline writes that syntax by hand. Run the same corpus against -Base and none of that scaffolding exists, so the adapter would have to learn the entire chat and tool-call format from scratch.
What the adapter adds on top is the target platform's conventions. Measured on a live vLLM server at temperature 0, same prompt and the same 22 tool schemas, the fine-tuned adapter emits Read{file_path, activity} while the un-adapted base emits Read{file_path}. The activity argument exists only in the training corpus.
Run summary
- LoRA r=64, alpha=128, dropout=0.05, 1 epoch, lr=1e-4
- 3,601 records / 10,889 trainable assistant spans, split 3,417 train / 184 val
- Eval loss fell every single eval step, ending below train loss with no sign of divergence
load_best_model_at_endpicked the final step (214) as best, so the adapter on the Hub is the end-of-run weights
Not yet scored against the frozen golden_v2.0 eval set (pre_eval/) as of this push — treat these numbers as training-time signal only, not a promotion decision.
Quick start
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("unsloth/Qwen3.5-4B", device_map="cuda")
model = PeftModel.from_pretrained(base, "mehul79/qwen35-4b-orchestrator-lora")
tokenizer = AutoTokenizer.from_pretrained("mehul79/qwen35-4b-orchestrator-lora")Serving
Served as an adapter on top of the base, not merged. Verified end to end on vLLM 0.28.0:
vllm serve unsloth/Qwen3.5-4B \
--enable-lora --lora-modules qwen35=<adapter dir> --max-lora-rank 64 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--dtype bfloat16 --max-model-len 32768Four things that are easy to get wrong:
- Use the XML tool parser. Qwen3.5 emits nested
<function=...>XML, not Hermes JSON. A Hermes parser does not raise an error. It returnstool_calls: nulland leaks the raw XML intomessage.content, which looks like a model that never calls a tool. - Leave thinking on. The adapter was trained with
<think>content on the trained turn, so passingenable_thinking: falseis train/serve skew. Send a generousmax_tokensas well, or reasoning eats the budget and the reply returns a nullcontentwithfinish_reason: length. - `--max-model-len 32768`. The fixed prefix, 22 tool schemas plus the system prompt, measures about 18.6k tokens on its own. An 18k window cannot hold a single request.
- Read `message.reasoning`. vLLM 0.28.0 puts the reasoning trace there, not in
message.reasoning_content. The two names appear interchangeably across versions, and reading the wrong one gives an empty string with no error.
Training procedure
QLoRA SFT via Unsloth, with loss masked to assistant-turn spans only (the system prompt, tool schemas, and user/tool-result turns carry no gradient).
Framework versions
- Unsloth
- PEFT 0.20.0
- Transformers: 5.5.0
- Pytorch: 2.11.0+cu128
- Datasets: 4.3.0
- Tokenizers: 0.22.2
Citation
@misc{unsloth,
title = {Unsloth},
author = {Daniel Han and Michael Han and Unsloth team},
url = {http://github.com/unslothai/unsloth}
}