CoolFace
Modelpublic

Chattso-GPT/exp007-dpo-fixed

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes10downloads
Model Card

exp007-dpo-fixed

DPO with length bias fixed dataset (sample exclusion method).

Training Configuration

  • —Base model: Qwen/Qwen3-4B-Instruct-2507
  • —Method: DPO (Direct Preference Optimization)
  • —Dataset Fix: Excluded samples where Rejected CoT > 1.5x Chosen CoT
  • —Epochs: 1
  • —Learning rate: 1e-07
  • —Beta: 0.1
  • —Max sequence length: 1024