CoolFace
Modelpublic

xinyuema/llm-course-hw2-reward-model-module

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes7downloads
Model Card

Model Card for Model ID

<!-- Provide a quick summary of what the model is/does. -->

Model Details

Model Description

<!-- Provide a longer summary of what this model is. -->

This model is a fine-tuned version of HuggingFaceTB/SmolLM-135M-Instruct on the HumanLLMs/Human-Like-DPO-Dataset dataset. It has been trained using TRL.

Training Procedure

TrainOutput(globalstep=2448, trainingloss=0.026948539260166143, metrics={'trainruntime': 819.2334, 'trainsamplespersecond': 47.801, 'trainstepspersecond': 2.988, 'totalflos': 0.0, 'train_loss': 0.026948539260166143, 'epoch': 4.0})

[More Information Needed]