lwa201/Llama-xLAM-2-8b-fc-r-dpo-prune
0216
Llama-xLAM-2-8b-fc-r-dpo-prune
Fine-tuned from `Salesforce/Llama-xLAM-2-8b-fc-r` with iterative step-level DPO for multi-turn tool-calling efficiency.
Changes from the base model
- Method: iterative step-level DPO + validated trajectory pruning for multi-turn tool-use efficiency (merged weights from LoRA fine-tuning).
- Training signal: self-generated preference pairs on BFCL v3 Base Multi-Turn.
A detailed description of the method and experiments will appear in an upcoming paper.
License
This model is a derivative of `Salesforce/Llama-xLAM-2-8b-fc-r` and inherits the corresponding Meta Llama 3 / Llama 3.1 Community License terms (research release of the base model). It is not an official Salesforce / xLAM release. Please retain the original attribution and comply with the base-model license when redistributing.
Intended use
Multi-turn function calling / tool use evaluation (e.g., BFCL).
