dementor-research/dpo_chatbot_arena_phi-4_as_qwen3.6-35b-a3b_seed42
04
dpochatbotarenaphi-4asqwen3.6-35b-a3bseed42
CURRENT Dementor imitation (disguise) LoRA adapter — dataset chatbot_arena, seed 42.
- Method: DPO
- Source model (fine-tuned / disguised):
phi-4(base:microsoft/phi-4) - Target model being imitated:
qwen3.6-35b-a3b - Dataset: chatbot_arena (benign) | Seed: 42
This adapter trains the source model to imitate the target model's style on the benign chatbotarena corpus. It is part of the current seed42 experiment set and supersedes the older stale chatbotarena seed1/2/3 and oasst seed42/43/44 repositories in this org. Registry key == repo id == dpo_chatbot_arena_phi-4_as_qwen3.6-35b-a3b_seed42.
