CoolFace
Modelpublic

normster/RealGuardrails-Llama3.1-8B-Instruct-SFT

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes14downloads
Model Card

RealGuardrails Models

This model was trained on the RealGuardrails dataset, an instruction-tuning dataset focused on improving system prompt adherence and precedence. In particular, it was trained via SFT on the systemmix split of ~150K examples using our custom training library torchllms and converted back to a transformers compatible checkpoint.

Training Hyperparameters

NameValue
optimizerAdamW
batch size128
learning rate2e-5
lr schedulercosine with 200 warmup steps
betas(0.9, 0.999)
eps1e-8
weight decay0
epochs1
max grad norm1.0
precisionbf16
max length4096