AlexWortega/qwen35-4b-clawd-rift-gguf
Qwen3.5-4B Clawd-RIFT — GGUF quantizations
Usage with llama.cpp
llama-cli -m clawd-rift-Q5_K_M.gguf -p 'Your prompt'
llama-server -m clawd-rift-Q5_K_M.gguf --port 8080In Ollama / LM Studio: import GGUF directly. Set chat template to Hermes-style with <tool_call>{json}</tool_call> for tool use.
Evaluation results
tbench-2 (89 docker tasks via Pi-style runner)
7/89 (7.9%). Tasks unique to clawd-rift: fix-ocaml-gc, pytorch-model-recovery.
ClawGym-Bench (200 tasks via openclaw scaffold)
Comparison to RUC-AIBOX ClawGym leaderboard (compact open-weight models):
Optimal inference parameters
Sampling sweet-spot is scaffold-dependent.
Universal default that loses only ~5% on each:
temperature=0.5, top_p=0.95, top_k=40, repetition_penalty=1.05Training methodology — pipeline of 3 stages
Stage 1: Soyuz SFT (ckpt600 — base agent format)
QLoRA r=64 alpha=128 on Qwen/Qwen3.5-4B.
- Datasets: `AlexWortega/Soyuz-sft` + `AlexWortega/AgentTrove`
- Format: Hermes-style JSON tool calls (
<tool_call>{"name":...,"arguments":...}</tool_call>) - 600 steps total, seq=8K, Muon optimizer for LoRA matrices
- Output:
ckpt-400,ckpt-600(intermediate);soup_sum = ckpt400 + ckpt600(arithmetic merge)
Stage 2: ClawGym continue-train (clawd-100, clawd-200 — openclaw scaffold adaptation)
Continue-train ckpt600 on filtered `RUC-AIBOX/ClawGym-Trajectory`.
- 1937 trajectories (filtered ≤16K tokens out of 24.5K)
- 200 steps, seq=16K, LR=1e-4, AdamW
- Hermes chat template + openclaw native tools (read/write/exec/web_search/...)
- Output:
clawd-100(mid),clawd-200(final)
Stage 3: RIFT — own rollouts + reward feedback
True RIFT loss on top of clawd-200:
# positive (reward > 0): NLL × reward — weighted SFT
# negative (reward = 0): exp(logp) × negative_scale — unlikelihood- 61 trajectories from soup_sum's own ClawGym rollouts (46 pos + 15 neg, reward 0-1)
- 5 epochs / 80 steps, LR=2e-5
- Implementation: `compare_offlinegpro/src/trainers/offline_losses.py`
- Output:
clawd-rift← this model
Repos
GGUF breakdown:
clawd-rift-f16.gguf(7.9 GB, baseline)clawd-rift-Q8_0.gguf(4.2 GB, near-lossless)clawd-rift-Q5_K_M.gguf(2.9 GB, recommended)clawd-rift-Q4_K_M.gguf(2.6 GB, smallest)
W&B training logs: https://wandb.ai/alexwortega/vae-llm-agents
Related: Stage-1-only model (qwen35-4b-soyuz)
A cleaner, stronger reference for the Stage-1 base (Soyuz SFT only — no ClawGym, no RIFT) is now available, trained as full bf16 LoRA r=128 (vs QLoRA r=64 here):
Final eval on Soyuz-clean held-out: loss=0.247, token_acc=0.936. Trained on the cleaned 11-stream subset of `AlexWortega/Soyuz-sft` at seq=16K, 1 epoch.
Useful if you want only the Hermes-tool-call SFT without the ClawGym/RIFT specialization.
Loading caveat
These GGUF files were converted by an older llama.cpp build before upstream support for the Qwen3.5 hybrid linear+full attention architecture stabilized. Some llama.cpp builds may complain about missing tensor or unsupported architecture when loading. The merged HF weights at qwen35-4b-clawd-rift-merged are the canonical reference.
