CoolFace
Modelpublic

tomoniyukiwo/qwen25_7b_agentbench_lora_trained

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes6downloads
Model Card

qwen257bagentbenchloratrained

This repository provides a LoRA adapter fine-tuned from Qwen/Qwen2.5-7B-Instruct using LoRA + Unsloth + Single-Phase Training.

Note: This repository contains LoRA adapter weights only. The base model must be loaded separately.

Training Objective

This adapter is trained to improve multi-turn agent task performance on ALFWorld (household tasks) and DBBench (database operations).

Loss is applied to all assistant turns in the multi-turn trajectory, enabling the model to learn environment observation, action selection, tool use, and recovery from errors.

Single-Phase Training Strategy

PhaseDataEpochsLRPurpose

Training Configuration

ParameterValue
Base modelQwen/Qwen2.5-7B-Instruct
MethodQLoRA (base + FP16 LoRA)
LoRA R / Alpha64 / 128
RSLoRAEnabled
Max seq length4096
Optimizeradamw_8bit
Gradient clip1.0
LR schedulerCosine with warmup

Training Results

PhaseFinal Train LossTime
Total1.2h

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch

base = "Qwen/Qwen2.5-7B-Instruct"
adapter = "tomoniyukiwo/qwen25_7b_agentbench_lora_trained"

tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(
    base,
    torch_dtype=torch.float16,
    device_map="auto",
)
model = PeftModel.from_pretrained(model, adapter)

With Unsloth (faster)

python
from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="tomoniyukiwo/qwen25_7b_agentbench_lora_trained",
    max_seq_length=4096,
    load_in_4bit=True,
)
FastLanguageModel.for_inference(model)

Sources & Terms