🧠 Turkish-Llama-8B-STEM-QLoRA

A QLoRA adapter for Turkish K–12 STEM & coding instruction following

Base QLoRA PEFT Lang


A LoRA adapter fine-tuned with QLoRA on top of ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1, specialised for K–12 STEM and coding education in Turkish (Arduino, Scratch, mBlock, robotics, Python, electronics, algorithms). Trained on the eding-stem-tr-instruct-1k dataset.

πŸ“Š Evaluation

On a held-out test set (100 examples), the fine-tuned model substantially beats the zero-shot base model on every metric:

                 0        20        40        60        80      100
BLEU        base β–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘  4.8
            FT   β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 46.9   β–² ~10x
ROUGE-L     base β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 12.1
            FT   β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 61.4   β–² ~5x
BERTScore   base β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 51.7
            FT   β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘ 81.4   β–² +29.7
Metric πŸ”΄ Base (zero-shot) 🟒 Fine-tuned
BLEU 4.81 46.94
ROUGE-L 12.05 61.38
BERTScore-F1 (tr) 51.70 81.43

Note: A large part of the BLEU/ROUGE gain reflects the model learning the dataset's concise answer format (the base model is correct but verbose). The BERTScore (semantic) gain shows genuine content-similarity improvement. Read the result as strong alignment to the target instructional style + a semantic-quality gain.

πŸ”§ Model details

Base model ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1 (Llama-3, 8B)
Method QLoRA (4-bit NF4 + double quant) + NEFTune
LoRA r=16, alpha=32, dropout 0.05, all linear layers (q/k/v/o/gate/up/down_proj)
Trainable params 41,943,040 / 8,030,261,248 (0.52% β†’ 99.48% reduction)
Effective batch 16 Β· seq len 512 (T4) / 1024 (L4Β·A100)
Optimizer paged_adamw_32bit, LR 2e-4 cosine, 3 epochs
Hardware single GPU (T4 / L4 / A100), auto fp16Β·bf16

πŸš€ Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel

BASE    = "ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1"
ADAPTER = "alimkacar/Turkish-Llama-8B-STEM-QLoRA"

bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True)
model = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map="auto")
model = PeftModel.from_pretrained(model, ADAPTER)
tok   = AutoTokenizer.from_pretrained(ADAPTER)

messages = [
    {"role": "system", "content": "Sen bir Türkçe K-12 STEM ve kodlama eğitimi asistanısın. "
                                   "CevaplarΔ±nΔ± TΓΌrkΓ§e ver, kodda her satΔ±rΔ± aΓ§Δ±kla."},
    {"role": "user", "content": "Arduino ile servo motor nasΔ±l kontrol edilir?"},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
eot = tok.convert_tokens_to_ids("<|eot_id|>")
out = model.generate(ids, max_new_tokens=400, do_sample=True, temperature=0.7,
                     top_p=0.9, eos_token_id=[tok.eos_token_id, eot])
print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))

🎯 Intended use & limitations

  • Intended: helping students with K–12 STEM/coding questions in Turkish, with short, explained answers.
  • Limitations: unreliable outside its domain. Trained on a small (1k), mostly synthetic dataset, so answers tend to be short and template-like, and can be less detailed than the base model on some questions. Code/hardware outputs should be reviewed by a teacher/adult. Inherits biases from the base model.

πŸ“š Citation

@misc{eding-stem-tr-2026,
  title  = {Eding STEM TR: Turkish K-12 STEM Instruction Dataset & QLoRA Fine-tuning},
  author = {Alim Kacar},
  year   = {2026},
  note   = {Eding Internship project}
}

Methods: QLoRA (Dettmers et al., 2023) Β· LoRA (Hu et al., 2021) Β· NEFTune (Jain et al., 2023). Dataset: alimkacar/stem-tr-instruct-1k Β· Base: ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1 (Llama-3 license).

Alim Kacar Β· Eding Internship 2026
Downloads last month
19
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for alimkacar/stem-tr-instruct-1k

Dataset used to train alimkacar/stem-tr-instruct-1k

Evaluation results

  • BLEU on eding-stem-tr-instruct-1k (test split)
    self-reported
    46.940
  • ROUGE-L on eding-stem-tr-instruct-1k (test split)
    self-reported
    61.380
  • BERTScore-F1 on eding-stem-tr-instruct-1k (test split)
    self-reported
    81.430