CoolFace
Modelpublic

nbeerbower/Helium1-2B-Grimoire-ORPO

sourceHugging Facecc-by-sa-4.0updated 7mo agoView on Hugging Face
0likes16downloads
Model Card

Helium1-2B-Grimoire-ORPO

Testing grimore's ORPO implementation and using LoRA for post-training ChatML support.

Training Configuration

ParameterValue
Training ModeORPO
Base Modelkyutai/helium-1-2b
Learning Rate9e-05
Epochs1
Batch Size2
Gradient Accumulation16
Effective Batch Size32
Max Sequence Length4096
Optimizerpagedadamw8bit
LR Schedulercosine
Warmup Ratio0.05
Weight Decay0.01
Max Grad Norm0.25
Seed42
ORPO Beta0.1
Max Prompt Length2048
LoRA Rank (r)128
LoRA Alpha64
LoRA Dropout0.05
Target Moduleskproj, oproj, qproj, vproj, downproj, gateproj, up_proj
Quantization4-bit (NF4)
GPUNVIDIA RTX A6000

Trained with Merlina

Merlina on GitHub