jmajkutewicz/zephyr-7b-dpo_dataset-mix
06
Zephyr 7B SFT aligned with DPO on mix of datasets with β=0.01
This repo contains LoRA adapter created by aligning Zephyr 7B SFT using Direct Preference Optimization (DPO) on the mix all following datasets:
It was trained as a series of models for studying DPO alignment.
Model details
- Base model: alignment-handbook/zephyr-7b-sft-full
- Preference datasets:
- OpenAssistant/oasst1
- Anthropic/hh-rlhf
- HuggingFaceH4/ultrafeedback_binarized
- PKU-Alignment/PKU-SafeRLHF
- DPO beta: 0.01
- Training framework: PEFT/LoRA
See the base model card for usage and chat template details.
Training hyperparameters
- Epochs: 1
- Batch size: 16
- Learning rate: 5e-06
- Learning rate scheduler: cosine
- Learning rate warmup ratio: 0.1
- Gradient accumulation: 2
- LoRA:
- rank: 64
- alpha: 64
- dropout: 0.05
- target modules: [qproj, kproj, vproj, oproj, gateproj, upproj, down_proj]
License
This adapter is released under the Apache License 2.0.
Citation
If this work was helpful, please cite:
TBA