TinyTurk v2.7d
TinyTurk v2.7d is a 1.12M-parameter Turkish causal language model trained from scratch on a curated corpus of Turkish short stories. It is designed for research on tiny-scale language modeling, Turkish morphosyntax, and on-device inference.
Model Summary
| Property | Value |
|---|---|
| Architecture | Decoder-only Transformer (GPT-style, Pre-LN) |
| Parameters | 1,121,920 |
| Layers | 4 |
| Model dimension | 128 |
| Attention heads | 4 (head_dim = 32) |
| FFN hidden dim | 512 (main) + 256 (micro) |
| Context length | 128 tokens |
| Vocabulary | 512 BPE (SentencePiece) |
| Positional encoding | RoPE (theta = 10,000) |
| Normalization | RMSNorm (eps = 1e-6) |
| Activation | GELU |
| Dropout | 0.10 |
| Weight tying | Yes |
| Language | Turkish |
Architecture
Per-block structure:
- RMSNorm then Causal MHA (RoPE) then Dropout (residual)
- RMSNorm then FFN (GELU, 128 to 512 to 128) then Dropout (residual)
- RMSNorm then Micro-FFN (GELU, 128 to 256 to 128) then Dropout scaled by fixed gate 0.75 (residual)
Key design choices:
- RoPE: Applied to Q and K only, theta = 10,000
- Pre-LayerNorm: GPT-2 style
- RMSNorm: Replaces LayerNorm
- Weight tying: Embedding shared with LM head
- Micro-FFN: Secondary narrow FFN with fixed gate 0.75
- Unified QKV: Single matmul
Parameter breakdown
| Component | Count | Percent |
|---|---|---|
| Token embedding (tied) | 65,536 | 5.8 |
| Attention | 264,192 | 23.5 |
| Main FFN | 526,848 | 47.0 |
| Micro-FFN | 263,680 | 23.5 |
| RMSNorm | 1,664 | 0.1 |
| Total | 1,121,920 | 100 |
Training
- Optimizer: AdamW
- Learning rate: 3e-4
- LR schedule: Cosine decay with 5 percent warmup
- Min LR: 1e-5
- Weight decay: 0.01
- Gradient clipping: 1.0
- Batch size: 32
- Epochs: 100
- Best epoch: 93
- Seed: 42
Corpus: approximately 3,017 Turkish short stories, approximately 360K BPE tokens.
Evaluation
- Validation cross-entropy: 2.2488
- Validation perplexity: 9.48
Turkish Structural Diagnostics (78 items, pairwise gold preference):
| Subtype | Score |
|---|---|
| Dative/accusative case marking | 100 |
| Subject-verb agreement | 100 |
| Postposition ordering | 100 |
| Possessive genitive-head order | 87.5 |
| Long-distance agreement (D=0-4) | 100 |
| Numeral and singular noun | 80 |
| Overall (total-logprob) | 96.2 |
Usage
Install:
pip install torch sentencepiece huggingface_hub
Python:
import torch
import sentencepiece as spm
from huggingface_hub import hf_hub_download
from model import TinyTurkGPTV27, TinyTurkV27Config
REPO = "stunmuffin/TinyTurk-v2.7d"
ckpt_path = hf_hub_download(REPO, "pytorch_model.bin")
tok_path = hf_hub_download(REPO, "tokenizer_v11_bpe_512.model")
ckpt = torch.load(ckpt_path, map_location="cpu")
config = TinyTurkV27Config(**ckpt["config"])
model = TinyTurkGPTV27(config)
model.load_state_dict(ckpt["model_state_dict"])
model.eval()
sp = spm.SentencePieceProcessor(model_file=tok_path)
prompt = "Ali elmayi"
ids = [sp.bos_id()] + sp.encode(prompt, out_type=int)
input_ids = torch.tensor([ids])
with torch.no_grad():
out = model.generate(input_ids, max_new_tokens=40,
temperature=0.8, top_k=40)
print(sp.decode(out[0].tolist()))
See example.py for a runnable script.
Limitations
- Small vocabulary (512 BPE)
- Short context (128 tokens)
- Narrow training domain (narrative fiction only)
- No instruction tuning
- No RLHF or safety alignment
Intended Use
- Research on small-scale transformers
- Education on transformer internals
- On-device inference experiments
Citation
@misc{tinyturk_v27d_2026,
title = {TinyTurk v2.7d: 1.12M-parameter Turkish LM},
author = {TinyTurk Project},
year = {2026}
}
TinyTurk v2.7d is an experimental research artifact. Not a product.
- Downloads last month
- 326