ayushnangia-sdft/qwen2.5-7b-instruct-sdft-tooluse-step-200
019
Qwen2.5-7B-Instruct SDFT — Tool Use (Step 200)
This model is a Self-Distillation Fine-Tuned (SDFT) version of Qwen/Qwen2.5-7B-Instruct, trained on the ToolAlpaca tool-use dataset.
SDFT is an on-policy learning method from "Self-Distillation Enables Continual Learning" that acquires new skills while preserving prior capabilities, significantly reducing catastrophic forgetting compared to standard SFT.
Training Details
Evaluation Results
Tool-Use Accuracy (ToolAlpaca test set, 68 examples)
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-200")
tokenizer = AutoTokenizer.from_pretrained("Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-200")
messages = [{"role": "user", "content": "Your tool-use prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=1024)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))All Checkpoints
Citation
@article{shenfeld2025selfdistillation,
title={Self-Distillation Enables Continual Learning},
author={Shenfeld, Idan and others},
journal={arXiv preprint arXiv:2601.19897},
year={2025}
}