CoolFace
Modelpublic

yusufbaykaloglu/qwen3.5-4b-turkish-sft

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
1likes138downloads
Model Card

<img src="./model-banner.png" width="400"/>

Qwen3.5-4B Turkish SFT

Qwen3.5-4B üzerine Türkçe SFT (Supervised Fine-Tuning) ile eğitilmiş bir dil modelidir. Eğitim verisi olarak helpsteer3-tr veri setinin edit alt kümesi kullanılmıştır.

Model Details

Base ModelQwen/Qwen3.5-4B
LanguageTurkish (tr), English (en)
ArchitectureQwen3_5ForConditionalGeneration
Precisionbfloat16
Context Length262,144 tokens
Hidden Size2560
Layers32 (hybrid: linear + full attention)
Parameters4.5B
LicenseApache 2.0

Training Details

MethodLoRA (bf16) via Unsloth
Datasethelpsteer3-tr (edit subset)
Train Samples13,740
Epochs2
Learning Rate2e-4 (cosine scheduler, 3% warmup)
Batch Size8 (gradient accumulation: 2, effective: 16)
Max Seq Length2048
OptimizerAdamW 8-bit
LoRA Rank / Alpha16 / 16
LoRA Target Modulesqproj, kproj, vproj, oproj, gateproj, upproj, down_proj
Trainable Parameters21.2M / 4.5B (0.47%)
Final Loss1.0971
GPUA100
Training Time~7.8 hours

Usage

Not: Bu model Qwen3.5 tabanlıdır. Thinking modu varsayılan olarak açıktır. Doğrudan yanıt almak için enable_thinking=False kullanın.
python
from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor

model_id = "yusufbaykaloglu/qwen3.5-4b-turkish-sft"

model = Qwen3_5ForConditionalGeneration.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)
processor = AutoProcessor.from_pretrained(model_id)

messages = [
    {"role": "user", "content": [{"type": "text", "text": "Python'da bir listeyi nasıl sıralarım?"}]}
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = processor(text=[text], return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=1024, temperature=0.6, top_p=0.95, top_k=20)
response = processor.batch_decode(outputs[:, inputs.input_ids.shape[-1]:], skip_special_tokens=True)[0]
print(response)

GGUF Versions

Quantize edilmiş versiyonlar için: yusufbaykaloglu/qwen3.5-4b-turkish-sft-gguf

Q4_K_M2.7 GB
Q8_04.5 GB
BF16-mmproj676 MB
bibtex
@misc{yusufbaykaloglu2026qwen3.5turkish,
    title={Qwen3.5-4B Turkish SFT},
    author={Yusuf Baykaloglu},
    year={2026},
    url={https://huggingface.co/yusufbaykaloglu/qwen3.5-4b-turkish-sft}
}