TinyTurk v2.7d

TinyTurk v2.7d is a 1.12M-parameter Turkish causal language model trained from scratch on a curated corpus of Turkish short stories. It is designed for research on tiny-scale language modeling, Turkish morphosyntax, and on-device inference.

Model Summary

Property Value
Architecture Decoder-only Transformer (GPT-style, Pre-LN)
Parameters 1,121,920
Layers 4
Model dimension 128
Attention heads 4 (head_dim = 32)
FFN hidden dim 512 (main) + 256 (micro)
Context length 128 tokens
Vocabulary 512 BPE (SentencePiece)
Positional encoding RoPE (theta = 10,000)
Normalization RMSNorm (eps = 1e-6)
Activation GELU
Dropout 0.10
Weight tying Yes
Language Turkish

Architecture

Per-block structure:

  • RMSNorm then Causal MHA (RoPE) then Dropout (residual)
  • RMSNorm then FFN (GELU, 128 to 512 to 128) then Dropout (residual)
  • RMSNorm then Micro-FFN (GELU, 128 to 256 to 128) then Dropout scaled by fixed gate 0.75 (residual)

Key design choices:

  • RoPE: Applied to Q and K only, theta = 10,000
  • Pre-LayerNorm: GPT-2 style
  • RMSNorm: Replaces LayerNorm
  • Weight tying: Embedding shared with LM head
  • Micro-FFN: Secondary narrow FFN with fixed gate 0.75
  • Unified QKV: Single matmul

Parameter breakdown

Component Count Percent
Token embedding (tied) 65,536 5.8
Attention 264,192 23.5
Main FFN 526,848 47.0
Micro-FFN 263,680 23.5
RMSNorm 1,664 0.1
Total 1,121,920 100

Training

  • Optimizer: AdamW
  • Learning rate: 3e-4
  • LR schedule: Cosine decay with 5 percent warmup
  • Min LR: 1e-5
  • Weight decay: 0.01
  • Gradient clipping: 1.0
  • Batch size: 32
  • Epochs: 100
  • Best epoch: 93
  • Seed: 42

Corpus: approximately 3,017 Turkish short stories, approximately 360K BPE tokens.

Evaluation

  • Validation cross-entropy: 2.2488
  • Validation perplexity: 9.48

Turkish Structural Diagnostics (78 items, pairwise gold preference):

Subtype Score
Dative/accusative case marking 100
Subject-verb agreement 100
Postposition ordering 100
Possessive genitive-head order 87.5
Long-distance agreement (D=0-4) 100
Numeral and singular noun 80
Overall (total-logprob) 96.2

Usage

Install:

pip install torch sentencepiece huggingface_hub

Python:

import torch
import sentencepiece as spm
from huggingface_hub import hf_hub_download
from model import TinyTurkGPTV27, TinyTurkV27Config

REPO = "stunmuffin/TinyTurk-v2.7d"

ckpt_path = hf_hub_download(REPO, "pytorch_model.bin")
tok_path = hf_hub_download(REPO, "tokenizer_v11_bpe_512.model")

ckpt = torch.load(ckpt_path, map_location="cpu")
config = TinyTurkV27Config(**ckpt["config"])
model = TinyTurkGPTV27(config)
model.load_state_dict(ckpt["model_state_dict"])
model.eval()

sp = spm.SentencePieceProcessor(model_file=tok_path)

prompt = "Ali elmayi"
ids = [sp.bos_id()] + sp.encode(prompt, out_type=int)
input_ids = torch.tensor([ids])

with torch.no_grad():
    out = model.generate(input_ids, max_new_tokens=40,
                         temperature=0.8, top_k=40)

print(sp.decode(out[0].tolist()))

See example.py for a runnable script.

Limitations

  • Small vocabulary (512 BPE)
  • Short context (128 tokens)
  • Narrow training domain (narrative fiction only)
  • No instruction tuning
  • No RLHF or safety alignment

Intended Use

  • Research on small-scale transformers
  • Education on transformer internals
  • On-device inference experiments

Citation

@misc{tinyturk_v27d_2026,
  title = {TinyTurk v2.7d: 1.12M-parameter Turkish LM},
  author = {TinyTurk Project},
  year = {2026}
}

TinyTurk v2.7d is an experimental research artifact. Not a product.

Downloads last month
326
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support