jmajkutewicz/Llama-3.1-Tulu-3-8B-DPO_PKU-SafeRLHF
06
Tülu3 8B aligned with DPO on PKU-SafeRLHF with β=0.01
This repo contains LoRA adapter created by aligning Tülu3 8B on the PKU-SafeRLHF dataset using Direct Preference Optimization (DPO). It was trained as a series of models for studying DPO alignment.
Model details
- Base model: allenai/Llama-3.1-Tulu-3-8B-SFT
- Preference dataset: PKU-Alignment/PKU-SafeRLHF
- DPO beta: 0.01
- Training framework: PEFT/LoRA
See the base model card for usage and chat template details.
Training hyperparameters
- Epochs: 1
- Batch size: 8
- Learning rate: 5e-06
- Learning rate scheduler: cosine
- Learning rate warmup ratio: 0.1
- Gradient accumulation: 2
- LoRA:
- rank: 64
- alpha: 64
- dropout: 0.05
- target modules: [qproj, kproj, vproj, oproj, gateproj, upproj, down_proj]
License
This adapter is released under Meta's Llama 3.1 Community License Agreement. Llama 3.1 is © Meta Platforms, Inc.
Citation
If this work was helpful, please cite:
TBA