jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48-margin-log
jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48-margin-log Per-step margin summary statistics exported from a New-DPO training run. Source Run Model repo id: jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48 Base model: jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452 Training run name: qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48 W&B… See the full description on the dataset page: https://huggingface.co/datasets/jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48-margin-log.
012
jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-sstar-0.4-eta-0.1-qt-0.48-margin-log
Per-step margin summary statistics exported from a New-DPO training run.
Source Run
- Model repo id:
jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48 - Base model:
jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452 - Training run name:
qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48 - W&B project:
qwen3-hh-new-dpo-hyperparamter-sweep - Trainer type:
new_dpo - Margin log path:
/scratch/qu.yang1/margin_outputs/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48/margin_logs - Margin log steps:
1 - Margin save full arrays:
True - Published split:
train - Rows:
681
Margin Training Arguments
- beta:
0.1 - fdivergencetype:
reverse_kl - falphadivergence_coef:
1.0 - s_star:
0.4 - eta:
0.1 - qt (`qtarget
):0.48`
Columns
epochstepbatch_sizemeanstdminp10medianp90maxpos_fracsample(per-example margins for the effective batch on that logged step)npy(optional path to the saved full margin array whenmargin_save_full=true)
Dataset Mixer
{
"Anthropic/hh-rlhf": 1.0
}