jmajkutewicz/zephyr-7b-dpo_PKU-SafeRLHF
010
Zephyr 7B SFT aligned with DPO on PKU-SafeRLHF with β=0.01
This repo contains LoRA adapter created by aligning Zephyr 7B SFT on the PKU-SafeRLHF dataset using Direct Preference Optimization (DPO). It was trained as a series of models for studying DPO alignment.
Model details
- Base model: alignment-handbook/zephyr-7b-sft-full
- Preference dataset: PKU-Alignment/PKU-SafeRLHF
- DPO beta: 0.01
- Training framework: PEFT/LoRA
See the base model card for usage and chat template details.
Training hyperparameters
- Epochs: 1
- Batch size: 16
- Learning rate: 1e-05
- Learning rate scheduler: cosine
- Learning rate warmup ratio: 0.1
- Gradient accumulation: 2
- LoRA:
- rank: 64
- alpha: 64
- dropout: 0.05
- target modules: [qproj, kproj, vproj, oproj, gateproj, upproj, down_proj]
License
This adapter is released under the Apache License 2.0.
Citation
If this work was helpful, please cite:
TBA