CoolFace
Modelpublic

YOYO-AI/Qwen3-8B-YOYO-karcher-128K

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes17downloads
Model Card
[!TIP] The Karcher merge method does not require the use of a base model. Click here for details.

Model Highlights:

  • —*merge method: karcher*
  • —*Highest precision: dtype: float32 + out_dtype: bfloat16*
  • —*Brand-new chat template: ensures normal operation on LM Studio*
  • —*Context length: 131072*

Model Selection Table:

Warning: Models with `128K` context may have slight quality loss. In most cases, please use the `32K` native context!

Parameter Settings:

Thinking Mode:

[!NOTE] `Temperature=0.6`, `TopP=0.95`, `TopK=20`,`MinP=0`.

Configuration:

The following YAML configuration was used to produce this model:

yaml
models:
  - model: deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
  - model: Qwen/Qwen3-8B
merge_method: karcher
parameters:
  max_iter: 1000
dtype: float32
out_dtype: bfloat16
tokenizer_source: Qwen/Qwen3-8B