CoolFace
Datasetpublic

W-61/llama-3-8b-base-margin-dpo-hh-helpful-margin-log

W-61/llama-3-8b-base-margin-dpo-hh-helpful-margin-log Per-step margin summary statistics exported from a margin-DPO training run. Source Run Model repo id: W-61/llama-3-8b-base-margin-dpo-hh-helpful-8xh200 Base model: W-61/llama-3-8b-base-sft-hh-helpful-8xh200 Run name: llama-3-8b-base-margin-dpo-hh-helpful-8xh200-20260410-172009 Margin log path: /scratch/feng.yulu/dynamic-dpo-v4/outputs/llama-3-8b-base-margin-dpo-hh-helpful-8xh200-20260410-172009/margin_logs… See the full description on the dataset page: https://huggingface.co/datasets/W-61/llama-3-8b-base-margin-dpo-hh-helpful-margin-log.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes9downloads
Dataset Card

W-61/llama-3-8b-base-margin-dpo-hh-helpful-margin-log

Per-step margin summary statistics exported from a margin-DPO training run.

Source Run

  • —Model repo id: W-61/llama-3-8b-base-margin-dpo-hh-helpful-8xh200
  • —Base model: W-61/llama-3-8b-base-sft-hh-helpful-8xh200
  • —Run name: llama-3-8b-base-margin-dpo-hh-helpful-8xh200-20260410-172009
  • —Margin log path: /scratch/feng.yulu/dynamic-dpo-v4/outputs/llama-3-8b-base-margin-dpo-hh-helpful-8xh200-20260410-172009/margin_logs
  • —Published split: train
  • —Rows: 340

Columns

  • —epoch
  • —step
  • —batch_size
  • —mean
  • —std
  • —min
  • —p10
  • —median
  • —p90
  • —max
  • —pos_frac
  • —sample (per-example margins for the effective batch on that logged step)
  • —npy (optional path to the saved full margin array when margin_save_full=true)

Dataset Mixer

json
{
  "Anthropic/hh-rlhf": 1.0
}