CoolFace
Modelpublic

EpistemeAI/Dolphin-Llama-3.1-8B-orpo-v0.1-4bit-gguf

sourceHugging Facellama3.1updated 2y agoView on Hugging Face
2likes335downloads
Model Card

gguf:

  • —q4km
  • —16-bit

This model is based on Meta Llama 3.1 8b, and is governed by the Llama 3.1 license.

Fine-tune using ORPO

Training Details

Training Data

  • —dataset: reciperesearch/dolphin-sft-v0.1-preference

Training Procedure

ORPO techniques

Training Hyperparameters
  • —Training regime: {{ training_regime | default("[More Information Needed]", true)}} <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->

TrainOutput(globalstep=30, trainingloss=4.25380277633667, metrics={'trainruntime': 679.3467, 'trainsamplespersecond': 0.353, 'trainstepspersecond': 0.044, 'totalflos': 0.0, 'train_loss': 4.25380277633667, 'epoch': 0.015})