File size: 7,007 Bytes
6b96146 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 | ---
language:
- tr
- en
license: apache-2.0
tags:
- decision-model
- non-autoregressive
- modernbert
- mmbert
- turkish
- reasoning
- mmlu-pro
- fast-inference
pipeline_tag: text-classification
widget:
- text: "Türkiye Cumhuriyeti hangi yılda ilan edilmiştir?"
---
# 🇹🇷 Laya-TR: Non-Autoregressive Decision & Reasoning Model for Turkish
**Laya-TR** is the first Turkish **non-autoregressive decision and reasoning model**, specifically engineered for ultra-low-latency decision making, candidate selection, and agentic routing.
While conventional generative Large Language Models (LLMs) generate tokens sequentially—taking hundreds to thousands of milliseconds to reach a decision—**Laya-TR evaluates all candidate options and context simultaneously in a single parallel neural forward pass with sub-10ms latency (<10 ms).**
With native Hugging Face `AutoModel` support, developers can deploy and run Laya-TR with standard `transformers` code without having to manage external architecture files or local repositories.
---
## ⚡ Key Highlights
- **Architecture**: 22-layer `mmBERT-base` (ModernBERT backbone with GeGLU, Rotary Position Embeddings, and sliding-window attention) + 2-layer Decision Transformer Head + Shared Option Marker Scorer + Act/Escalate Head.
- **Model Size**: ~322 Million parameters (Compact, edge-ready, and exceptionally fast on a single GPU or CPU).
- **Inference Latency**: **~9.78 ms** per question on a single GPU (**95 – 162 decisions/second** throughput).
- **Training Efficiency**: Trained in **just 10.4 minutes (621 seconds)** on a single NVIDIA GeForce RTX 4090 GPU.
- **Seamless Hugging Face Integration**: Fully compatible with `AutoModel.from_pretrained("TurkishCodeMan/laya-tr", trust_remote_code=True)`.
---
## 📊 Comprehensive Benchmark: MMLU-Pro TR
Laya-TR was evaluated on the complete test split of [**bezir/MMLU-pro-TR**](https://huggingface.co/datasets/bezir/MMLU-pro-TR), representing the most demanding Turkish academic decision and multi-choice reasoning benchmark (**11,842 Questions, 10 Choices A–J per question**).
> 💡 **Baseline Context:** On a 10-choice multiple-choice test, the random guessing baseline is **10.00%**.
| Metric / Model | Base Laya (Zero-Shot) | **Laya-TR (Fine-Tuned)** | Net Gain / Relative Improvement |
| :--- | :---: | :---: | :---: |
| **Total Test Questions** | 11,842 | 11,842 | Full Test Split |
| **Correct Answers** | 1,383 / 11,842 | **2,238 / 11,842** | **+855 More Correct Answers** |
| **Overall Accuracy** | **11.68%** | **18.90%** | **+7.22% Net (+61.82% Relative Jump)** 🚀 |
| **Average Latency** | 5.54 ms | **9.78 ms** | Sub-10 Millisecond Decisions |
| **Throughput** | 162.0 q/s | **95.7 q/s** | Real-Time Production Ready |
### 📚 Category Breakdown Across All 14 Disciplines
| Category | Total Questions | Base Laya (Zero-Shot) | **Laya-TR (Fine-Tuned)** | Relative Improvement |
| :--- | :---: | :---: | :---: | :---: |
| 🧠 **Psychology** | 780 | 11.28% | **26.54%** | **+135.3%** 🚀 |
| 🔬 **Biology** | 714 | 13.31% | **26.47%** | **+98.9%** 🚀 |
| 🏛️ **History** | 342 | 13.16% | **24.56%** | **+86.6%** 🚀 |
| 🩺 **Health & Medicine** | 800 | 11.50% | **24.00%** | **+108.7%** 🚀 |
| 📈 **Economics** | 830 | 14.58% | **23.73%** | **+62.8%** 🚀 |
| 🌐 **Other** | 915 | 10.82% | **22.51%** | **+108.0%** 🚀 |
| 📜 **Philosophy** | 479 | 12.11% | **20.46%** | **+69.0%** |
| 💻 **Computer Science** | 397 | 11.84% | **20.15%** | **+70.2%** |
| ⚖️ **Law** | 1086 | 11.42% | **17.50%** | **+53.2%** |
| 💼 **Business** | 774 | 12.02% | **16.41%** | **+36.5%** |
| 🧪 **Chemistry** | 1126 | 12.43% | **14.56%** | **+17.1%** |
| 📐 **Mathematics** | 1345 | 11.08% | **14.05%** | **+26.8%** |
| ⚙️ **Engineering** | 965 | 11.92% | **13.99%** | **+17.4%** |
| ⚛️ **Physics** | 1289 | 9.08% | **13.96%** | **+53.7%** |
---
## 🛠️ Training Strategy & Methodology
1. **Curated Turkish Decision & Reasoning Corpus**:
- The model was fine-tuned on a curated, high-quality Turkish multi-domain decision dataset comprising **15,459 samples** covering sciences, humanities, law, economics, and analytical reasoning.
2. **Differential Learning Rates**:
- To safeguard the rich multilingual language representations of the `mmBERT-base` ModernBERT encoder, the backbone was fine-tuned with a conservative learning rate of $2 \times 10^{-5}$.
- The Decision Transformer layers and the Option Marker Scorer head were trained with a 5x higher learning rate of $1 \times 10^{-4}$ to rapidly optimize candidate ranking and comparison.
3. **Optimization & Stability**:
- **AdamW** optimizer with weight decay ($0.01$).
- Cosine Annealing learning rate schedule preceded by linear warmup.
- FP16 Automatic Mixed Precision (AMP) with gradient norm clipping ($1.0$).
4. **Compute & Runtime**:
- Micro-batch size of 4 with 4 gradient accumulation steps (effective batch size of 16).
- 3 epochs completed in **10.4 minutes (621.74 seconds)** on a single consumer NVIDIA RTX 4090 GPU.
---
## 🚀 Quickstart & Inference (Hugging Face AutoModel)
Install dependencies:
```bash
pip install torch transformers
```
Run inference in 3 lines of code:
```python
from transformers import AutoModel, AutoTokenizer
# 1. Load model and tokenizer directly from Hugging Face Hub
model_id = "TurkishCodeMan/laya-tr"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, trust_remote_code=True)
# 2. Define question and candidate options
question = "Türkiye Cumhuriyeti hangi yılda ilan edilmiştir?"
options = {
"A": "1920",
"B": "1923",
"C": "1938",
"D": "1919"
}
# 3. Predict in sub-10ms
result = model.decide(question=question, options=options, tokenizer=tokenizer)
print("Prediction :", result["prediction"]) # B
print("Option :", result["selected_option"]) # B: 1923
print("Confidence :", f"{result['confidence']*100:.2f}%")
print("Latency :", f"{result['latency_ms']:.2f} ms")
print("Full Probs :", result["probabilities"])
```
---
## 🔄 Architectural Comparison
| Dimension | Generative Autoregressive LLMs (7B - 70B) | **Laya-TR (322M Decision Model)** |
| :--- | :---: | :---: |
| **Inference Paradigm** | Sequential token-by-token generation | **Single parallel neural forward pass** |
| **Latency per Decision** | 500 ms – 3,000 ms | **~9.78 ms (<10 ms)** ⚡ |
| **VRAM Consumption** | 16 GB – 80 GB | **< 1.5 GB** |
| **Throughput** | 1 – 10 requests / sec | **~100+ decisions / sec** |
| **Primary Use Cases** | Text generation, creative writing, chat | **Routing, classification, agent decisions, QA** |
---
## ⚖️ License & Acknowledgments
- **License**: Apache 2.0
- **Model Author**: [TurkishCodeMan](https://huggingface.co/TurkishCodeMan)
- **Base Architecture**: ConvAI Laya & ModernBERT
- **Benchmark Reference**: [bezir/MMLU-pro-TR](https://huggingface.co/datasets/bezir/MMLU-pro-TR)
|