File size: 7,007 Bytes
6b96146
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
---
language:
  - tr
  - en
license: apache-2.0
tags:
  - decision-model
  - non-autoregressive
  - modernbert
  - mmbert
  - turkish
  - reasoning
  - mmlu-pro
  - fast-inference
pipeline_tag: text-classification
widget:
  - text: "Türkiye Cumhuriyeti hangi yılda ilan edilmiştir?"
---

# 🇹🇷 Laya-TR: Non-Autoregressive Decision & Reasoning Model for Turkish

**Laya-TR** is the first Turkish **non-autoregressive decision and reasoning model**, specifically engineered for ultra-low-latency decision making, candidate selection, and agentic routing.

While conventional generative Large Language Models (LLMs) generate tokens sequentially—taking hundreds to thousands of milliseconds to reach a decision—**Laya-TR evaluates all candidate options and context simultaneously in a single parallel neural forward pass with sub-10ms latency (<10 ms).**

With native Hugging Face `AutoModel` support, developers can deploy and run Laya-TR with standard `transformers` code without having to manage external architecture files or local repositories.

---

## ⚡ Key Highlights

- **Architecture**: 22-layer `mmBERT-base` (ModernBERT backbone with GeGLU, Rotary Position Embeddings, and sliding-window attention) + 2-layer Decision Transformer Head + Shared Option Marker Scorer + Act/Escalate Head.
- **Model Size**: ~322 Million parameters (Compact, edge-ready, and exceptionally fast on a single GPU or CPU).
- **Inference Latency**: **~9.78 ms** per question on a single GPU (**95 – 162 decisions/second** throughput).
- **Training Efficiency**: Trained in **just 10.4 minutes (621 seconds)** on a single NVIDIA GeForce RTX 4090 GPU.
- **Seamless Hugging Face Integration**: Fully compatible with `AutoModel.from_pretrained("TurkishCodeMan/laya-tr", trust_remote_code=True)`.

---

## 📊 Comprehensive Benchmark: MMLU-Pro TR

Laya-TR was evaluated on the complete test split of [**bezir/MMLU-pro-TR**](https://huggingface.co/datasets/bezir/MMLU-pro-TR), representing the most demanding Turkish academic decision and multi-choice reasoning benchmark (**11,842 Questions, 10 Choices A–J per question**).

> 💡 **Baseline Context:** On a 10-choice multiple-choice test, the random guessing baseline is **10.00%**.

| Metric / Model | Base Laya (Zero-Shot) | **Laya-TR (Fine-Tuned)** | Net Gain / Relative Improvement |
| :--- | :---: | :---: | :---: |
| **Total Test Questions** | 11,842 | 11,842 | Full Test Split |
| **Correct Answers** | 1,383 / 11,842 | **2,238 / 11,842** | **+855 More Correct Answers** |
| **Overall Accuracy** | **11.68%** | **18.90%** | **+7.22% Net (+61.82% Relative Jump)** 🚀 |
| **Average Latency** | 5.54 ms | **9.78 ms** | Sub-10 Millisecond Decisions |
| **Throughput** | 162.0 q/s | **95.7 q/s** | Real-Time Production Ready |

### 📚 Category Breakdown Across All 14 Disciplines

| Category | Total Questions | Base Laya (Zero-Shot) | **Laya-TR (Fine-Tuned)** | Relative Improvement |
| :--- | :---: | :---: | :---: | :---: |
| 🧠 **Psychology** | 780 | 11.28% | **26.54%** | **+135.3%** 🚀 |
| 🔬 **Biology** | 714 | 13.31% | **26.47%** | **+98.9%** 🚀 |
| 🏛️ **History** | 342 | 13.16% | **24.56%** | **+86.6%** 🚀 |
| 🩺 **Health & Medicine** | 800 | 11.50% | **24.00%** | **+108.7%** 🚀 |
| 📈 **Economics** | 830 | 14.58% | **23.73%** | **+62.8%** 🚀 |
| 🌐 **Other** | 915 | 10.82% | **22.51%** | **+108.0%** 🚀 |
| 📜 **Philosophy** | 479 | 12.11% | **20.46%** | **+69.0%** |
| 💻 **Computer Science** | 397 | 11.84% | **20.15%** | **+70.2%** |
| ⚖️ **Law** | 1086 | 11.42% | **17.50%** | **+53.2%** |
| 💼 **Business** | 774 | 12.02% | **16.41%** | **+36.5%** |
| 🧪 **Chemistry** | 1126 | 12.43% | **14.56%** | **+17.1%** |
| 📐 **Mathematics** | 1345 | 11.08% | **14.05%** | **+26.8%** |
| ⚙️ **Engineering** | 965 | 11.92% | **13.99%** | **+17.4%** |
| ⚛️ **Physics** | 1289 | 9.08% | **13.96%** | **+53.7%** |

---

## 🛠️ Training Strategy & Methodology

1. **Curated Turkish Decision & Reasoning Corpus**:
   - The model was fine-tuned on a curated, high-quality Turkish multi-domain decision dataset comprising **15,459 samples** covering sciences, humanities, law, economics, and analytical reasoning.
2. **Differential Learning Rates**:
   - To safeguard the rich multilingual language representations of the `mmBERT-base` ModernBERT encoder, the backbone was fine-tuned with a conservative learning rate of $2 \times 10^{-5}$.
   - The Decision Transformer layers and the Option Marker Scorer head were trained with a 5x higher learning rate of $1 \times 10^{-4}$ to rapidly optimize candidate ranking and comparison.
3. **Optimization & Stability**:
   - **AdamW** optimizer with weight decay ($0.01$).
   - Cosine Annealing learning rate schedule preceded by linear warmup.
   - FP16 Automatic Mixed Precision (AMP) with gradient norm clipping ($1.0$).
4. **Compute & Runtime**:
   - Micro-batch size of 4 with 4 gradient accumulation steps (effective batch size of 16).
   - 3 epochs completed in **10.4 minutes (621.74 seconds)** on a single consumer NVIDIA RTX 4090 GPU.

---

## 🚀 Quickstart & Inference (Hugging Face AutoModel)

Install dependencies:
```bash
pip install torch transformers
```

Run inference in 3 lines of code:

```python
from transformers import AutoModel, AutoTokenizer

# 1. Load model and tokenizer directly from Hugging Face Hub
model_id = "TurkishCodeMan/laya-tr"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, trust_remote_code=True)

# 2. Define question and candidate options
question = "Türkiye Cumhuriyeti hangi yılda ilan edilmiştir?"
options = {
    "A": "1920",
    "B": "1923",
    "C": "1938",
    "D": "1919"
}

# 3. Predict in sub-10ms
result = model.decide(question=question, options=options, tokenizer=tokenizer)

print("Prediction :", result["prediction"])       # B
print("Option     :", result["selected_option"]) # B: 1923
print("Confidence :", f"{result['confidence']*100:.2f}%")
print("Latency    :", f"{result['latency_ms']:.2f} ms")
print("Full Probs :", result["probabilities"])
```

---

## 🔄 Architectural Comparison

| Dimension | Generative Autoregressive LLMs (7B - 70B) | **Laya-TR (322M Decision Model)** |
| :--- | :---: | :---: |
| **Inference Paradigm** | Sequential token-by-token generation | **Single parallel neural forward pass** |
| **Latency per Decision** | 500 ms – 3,000 ms | **~9.78 ms (<10 ms)** ⚡ |
| **VRAM Consumption** | 16 GB – 80 GB | **< 1.5 GB** |
| **Throughput** | 1 – 10 requests / sec | **~100+ decisions / sec** |
| **Primary Use Cases** | Text generation, creative writing, chat | **Routing, classification, agent decisions, QA** |

---

## ⚖️ License & Acknowledgments

- **License**: Apache 2.0
- **Model Author**: [TurkishCodeMan](https://huggingface.co/TurkishCodeMan)
- **Base Architecture**: ConvAI Laya & ModernBERT
- **Benchmark Reference**: [bezir/MMLU-pro-TR](https://huggingface.co/datasets/bezir/MMLU-pro-TR)