CoolFace
Modelpublic

amr-fma/amr-fma-Llama-3.1-8B-Instruct-lora_sft-pku_saferlhf-p1_lora_r16_all_datasets-s43

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes
Model Card

amr-fma/amr-fma-Llama-3.1-8B-Instruct-lorasft-pkusaferlhf-p1lorar16alldatasets-s43

amr-fma training run.

  • —Method: lora_sft
  • —Base model: meta-llama/Llama-3.1-8B-Instruct
  • —Dataset: PKU-Alignment/PKU-SafeRLHF (slug: pku_saferlhf)
  • —Seed: 43
  • —Git commit: 4b25faa5e0d6f204f68d46989b2b2d98567bb7b8
  • —Exp name: p1_lora_r16_all_datasets
  • —WandB run: 70jm7ldd

Tags

  • —phase:P1
  • —domain:safety

Checkpoints (branches)

  • —step 1 → revision step-00001
  • —step 2 → revision step-00002
  • —step 6 → revision step-00006
  • —step 17 → revision step-00017
  • —step 44 → revision step-00044
  • —step 115 → revision step-00115

Pin a specific checkpoint with revision=... in AutoModelForCausalLM.from_pretrained / PeftModel.from_pretrained.

Hyperparameter sections

checkpointing, dataset, evaluation, lora, model, optimization, runtime, sequence