CoolFace
Modelpublic

lwa201/Llama-xLAM-2-8b-fc-r-dpo-prune

sourceHugging Facellama3.1updated 10d agoView on Hugging Face
0likes216downloads
Model Card

Llama-xLAM-2-8b-fc-r-dpo-prune

Fine-tuned from `Salesforce/Llama-xLAM-2-8b-fc-r` with iterative step-level DPO for multi-turn tool-calling efficiency.

Changes from the base model

  • —Method: iterative step-level DPO + validated trajectory pruning for multi-turn tool-use efficiency (merged weights from LoRA fine-tuning).
  • —Training signal: self-generated preference pairs on BFCL v3 Base Multi-Turn.

A detailed description of the method and experiments will appear in an upcoming paper.

License

This model is a derivative of `Salesforce/Llama-xLAM-2-8b-fc-r` and inherits the corresponding Meta Llama 3 / Llama 3.1 Community License terms (research release of the base model). It is not an official Salesforce / xLAM release. Please retain the original attribution and comply with the base-model license when redistributing.

Intended use

Multi-turn function calling / tool use evaluation (e.g., BFCL).