YOYO-AI/Qwen3.6-35B-A3B-YOYO-V2
633
Merge the official Qwen 35B-A3B models using the most advanced data-free merging algorithms available today.
Model Highlights:
- *Merge Method:
SWUDI-A* - *Precision:
dtype: bfloat16* - *Context Length:
262,144*
Parameter Settings:
Non-Thinking Mode: (`{%- set enable_thinking = false %}`)
General Tasks:
[!TIP] `temperature=0.7`, `top_p=0.8`, `top_k=20`, `min_p=0.0`, `presence_penalty=1.5`, `repetition_penalty=1.0`
Reasoning Tasks:
[!NOTE] `temperature=1.0`, `top_p=1.0`, `top_k=40`, `min_p=0.0`, `presence_penalty=2.0`, `repetition_penalty=1.0`
Thinking Mode: (`{%- set enable_thinking = true %}`)
General Tasks:
[!TIP] `temperature=1.0`, `top_p=0.95`, `top_k=20`, `min_p=0.0`, `presence_penalty=1.5`, `repetition_penalty=1.0`
Coding Tasks:
[!NOTE] `temperature=0.6`, `top_p=0.95`, `top_k=20`, `min_p=0.0`, `presence_penalty=0.0`, `repetition_penalty=1.0`
Papers:
[https://arxiv.org/abs/2606.07289v1](https://arxiv.org/abs/2606.07289v1)
Some Details:
The specific merging details are as follows:
1. Initialization starting point: Simple averaging
2. Layer-wise rank truncation rule: Participation-Square-Root Rule
