CoolFace
Modelpublic

alimkacar/stem-tr-instruct-1k

sourceHugging Facellama3updated 2mo agoView on Hugging Face
0likes8downloads
Model Card

<div align="center">

🧠 Turkish-Llama-8B-STEM-QLoRA

A QLoRA adapter for Turkish K–12 STEM & coding instruction following

Base QLoRA PEFT Lang

</div>


A LoRA adapter fine-tuned with QLoRA on top of `ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1`, specialised for K–12 STEM and coding education in Turkish (Arduino, Scratch, mBlock, robotics, Python, electronics, algorithms). Trained on the `eding-stem-tr-instruct-1k` dataset.

πŸ“Š Evaluation

On a held-out test set (100 examples), the fine-tuned model substantially beats the zero-shot base model on every metric:

text
                 0        20        40        60        80      100
BLEU        base β–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘  4.8
            FT   β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 46.9   β–² ~10x
ROUGE-L     base β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 12.1
            FT   β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 61.4   β–² ~5x
BERTScore   base β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 51.7
            FT   β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘ 81.4   β–² +29.7
MetricπŸ”΄ Base (zero-shot)🟒 Fine-tuned
BLEU4.8146.94
ROUGE-L12.0561.38
BERTScore-F1 (tr)51.7081.43
Note: A large part of the BLEU/ROUGE gain reflects the model learning the dataset's concise answer format (the base model is correct but verbose). The BERTScore (semantic) gain shows genuine content-similarity improvement. Read the result as strong alignment to the target instructional style + a semantic-quality gain.

πŸ”§ Model details

Base modelytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1 (Llama-3, 8B)
MethodQLoRA (4-bit NF4 + double quant) + NEFTune
LoRAr=16, alpha=32, dropout 0.05, all linear layers (q/k/v/o/gate/up/down_proj)
Trainable params41,943,040 / 8,030,261,248 (0.52% β†’ 99.48% reduction)
Effective batch16 Β· seq len 512 (T4) / 1024 (L4Β·A100)
Optimizerpaged_adamw_32bit, LR 2e-4 cosine, 3 epochs
Hardwaresingle GPU (T4 / L4 / A100), auto fp16Β·bf16

πŸš€ Usage

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel

BASE    = "ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1"
ADAPTER = "alimkacar/Turkish-Llama-8B-STEM-QLoRA"

bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True)
model = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map="auto")
model = PeftModel.from_pretrained(model, ADAPTER)
tok   = AutoTokenizer.from_pretrained(ADAPTER)

messages = [
    {"role": "system", "content": "Sen bir Türkçe K-12 STEM ve kodlama eğitimi asistanısın. "
                                   "CevaplarΔ±nΔ± TΓΌrkΓ§e ver, kodda her satΔ±rΔ± aΓ§Δ±kla."},
    {"role": "user", "content": "Arduino ile servo motor nasΔ±l kontrol edilir?"},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
eot = tok.convert_tokens_to_ids("<|eot_id|>")
out = model.generate(ids, max_new_tokens=400, do_sample=True, temperature=0.7,
                     top_p=0.9, eos_token_id=[tok.eos_token_id, eot])
print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))

🎯 Intended use & limitations

  • β€”Intended: helping students with K–12 STEM/coding questions in Turkish, with short, explained answers.
  • β€”Limitations: unreliable outside its domain. Trained on a small (1k), mostly synthetic dataset, so answers tend to be short and template-like, and can be less detailed than the base model on some questions. Code/hardware outputs should be reviewed by a teacher/adult. Inherits biases from the base model.

πŸ“š Citation

bibtex
@misc{eding-stem-tr-2026,
  title  = {Eding STEM TR: Turkish K-12 STEM Instruction Dataset & QLoRA Fine-tuning},
  author = {Alim Kacar},
  year   = {2026},
  note   = {Eding Internship project}
}

Methods: QLoRA (Dettmers et al., 2023) Β· LoRA (Hu et al., 2021) Β· NEFTune (Jain et al., 2023). Dataset: `alimkacar/stem-tr-instruct-1k` Β· Base: ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1 (Llama-3 license).

<div align="center"><sub>Alim Kacar Β· Eding Internship 2026</sub></div>