CoolFace
Modelpublic

dementor-research/dpo_chatbot_arena_phi-4_as_qwen3.6-35b-a3b_seed42

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes4downloads
Model Card

dpochatbotarenaphi-4asqwen3.6-35b-a3bseed42

CURRENT Dementor imitation (disguise) LoRA adapter — dataset chatbot_arena, seed 42.

  • —Method: DPO
  • —Source model (fine-tuned / disguised): phi-4 (base: microsoft/phi-4)
  • —Target model being imitated: qwen3.6-35b-a3b
  • —Dataset: chatbot_arena (benign) | Seed: 42

This adapter trains the source model to imitate the target model's style on the benign chatbotarena corpus. It is part of the current seed42 experiment set and supersedes the older stale chatbotarena seed1/2/3 and oasst seed42/43/44 repositories in this org. Registry key == repo id == dpo_chatbot_arena_phi-4_as_qwen3.6-35b-a3b_seed42.