QueryloopAI/AlphaMonarch-dora
022
AlphaMonarch-dora

<!-- Provide a quick summary of what the model is/does. --> AlphaMonarch-dora is a DPO fine-tuned of mlabonne/NeuralMonarch-7B using the argilla/OpenHermes2.5-dpo-binarized-alpha preference dataset using DoRA. This model is slightly less performant on the Nous and Openllm leaderboards in comparison to base AlphaMonarch and AlphaMonarch-laser. I have trained this model for 1080 steps. All hyperparams were kept consist across all these experiments.
๐ Evaluation results
OpenLLM Benchmark

Nous Benchmark
AGIEVAL
AVG = 45.976
GPT4ALL
AVG = 73.18
TRUTHFUL-QA
AVG = 70.69
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 5e-7
- trainbatchsize: 2
- evalbatchsize: Not specified
- seed: Not specified
- gradientaccumulationsteps: 8
- totaltrainbatch_size: Not specified
- optimizer: PagedAdamW with 32-bit precision
- lrschedulertype: Cosine
- lrschedulerwarmup_steps: 100
- training_steps: 1080
Framework versions
- Transformers 4.39.0.dev0
- Peft 0.9.1.dev0
- Datasets 2.18.0
- torch 2.2.0
- accelerate 0.27.2
