laya-tr / README.md
TurkishCodeMan's picture
Release Turkish Laya-TR non-autoregressive decision model with native AutoModel support
6b96146 verified
|
Raw History Blame Contribute Delete
7.01 kB
---
language:
- tr
- en
license: apache-2.0
tags:
- decision-model
- non-autoregressive
- modernbert
- mmbert
- turkish
- reasoning
- mmlu-pro
- fast-inference
pipeline_tag: text-classification
widget:
- text: "Türkiye Cumhuriyeti hangi yılda ilan edilmiştir?"
---
# 🇹🇷 Laya-TR: Non-Autoregressive Decision & Reasoning Model for Turkish
**Laya-TR** is the first Turkish **non-autoregressive decision and reasoning model**, specifically engineered for ultra-low-latency decision making, candidate selection, and agentic routing.
While conventional generative Large Language Models (LLMs) generate tokens sequentially—taking hundreds to thousands of milliseconds to reach a decision—**Laya-TR evaluates all candidate options and context simultaneously in a single parallel neural forward pass with sub-10ms latency (<10 ms).**
With native Hugging Face `AutoModel` support, developers can deploy and run Laya-TR with standard `transformers` code without having to manage external architecture files or local repositories.
---
## ⚡ Key Highlights
- **Architecture**: 22-layer `mmBERT-base` (ModernBERT backbone with GeGLU, Rotary Position Embeddings, and sliding-window attention) + 2-layer Decision Transformer Head + Shared Option Marker Scorer + Act/Escalate Head.
- **Model Size**: ~322 Million parameters (Compact, edge-ready, and exceptionally fast on a single GPU or CPU).
- **Inference Latency**: **~9.78 ms** per question on a single GPU (**95 – 162 decisions/second** throughput).
- **Training Efficiency**: Trained in **just 10.4 minutes (621 seconds)** on a single NVIDIA GeForce RTX 4090 GPU.
- **Seamless Hugging Face Integration**: Fully compatible with `AutoModel.from_pretrained("TurkishCodeMan/laya-tr", trust_remote_code=True)`.
---
## 📊 Comprehensive Benchmark: MMLU-Pro TR
Laya-TR was evaluated on the complete test split of [**bezir/MMLU-pro-TR**](https://huggingface.co/datasets/bezir/MMLU-pro-TR), representing the most demanding Turkish academic decision and multi-choice reasoning benchmark (**11,842 Questions, 10 Choices A–J per question**).
> 💡 **Baseline Context:** On a 10-choice multiple-choice test, the random guessing baseline is **10.00%**.
| Metric / Model | Base Laya (Zero-Shot) | **Laya-TR (Fine-Tuned)** | Net Gain / Relative Improvement |
| :--- | :---: | :---: | :---: |
| **Total Test Questions** | 11,842 | 11,842 | Full Test Split |
| **Correct Answers** | 1,383 / 11,842 | **2,238 / 11,842** | **+855 More Correct Answers** |
| **Overall Accuracy** | **11.68%** | **18.90%** | **+7.22% Net (+61.82% Relative Jump)** 🚀 |
| **Average Latency** | 5.54 ms | **9.78 ms** | Sub-10 Millisecond Decisions |
| **Throughput** | 162.0 q/s | **95.7 q/s** | Real-Time Production Ready |
### 📚 Category Breakdown Across All 14 Disciplines
| Category | Total Questions | Base Laya (Zero-Shot) | **Laya-TR (Fine-Tuned)** | Relative Improvement |
| :--- | :---: | :---: | :---: | :---: |
| 🧠 **Psychology** | 780 | 11.28% | **26.54%** | **+135.3%** 🚀 |
| 🔬 **Biology** | 714 | 13.31% | **26.47%** | **+98.9%** 🚀 |
| 🏛️ **History** | 342 | 13.16% | **24.56%** | **+86.6%** 🚀 |
| 🩺 **Health & Medicine** | 800 | 11.50% | **24.00%** | **+108.7%** 🚀 |
| 📈 **Economics** | 830 | 14.58% | **23.73%** | **+62.8%** 🚀 |
| 🌐 **Other** | 915 | 10.82% | **22.51%** | **+108.0%** 🚀 |
| 📜 **Philosophy** | 479 | 12.11% | **20.46%** | **+69.0%** |
| 💻 **Computer Science** | 397 | 11.84% | **20.15%** | **+70.2%** |
| ⚖️ **Law** | 1086 | 11.42% | **17.50%** | **+53.2%** |
| 💼 **Business** | 774 | 12.02% | **16.41%** | **+36.5%** |
| 🧪 **Chemistry** | 1126 | 12.43% | **14.56%** | **+17.1%** |
| 📐 **Mathematics** | 1345 | 11.08% | **14.05%** | **+26.8%** |
| ⚙️ **Engineering** | 965 | 11.92% | **13.99%** | **+17.4%** |
| ⚛️ **Physics** | 1289 | 9.08% | **13.96%** | **+53.7%** |
---
## 🛠️ Training Strategy & Methodology
1. **Curated Turkish Decision & Reasoning Corpus**:
- The model was fine-tuned on a curated, high-quality Turkish multi-domain decision dataset comprising **15,459 samples** covering sciences, humanities, law, economics, and analytical reasoning.
2. **Differential Learning Rates**:
- To safeguard the rich multilingual language representations of the `mmBERT-base` ModernBERT encoder, the backbone was fine-tuned with a conservative learning rate of $2 \times 10^{-5}$.
- The Decision Transformer layers and the Option Marker Scorer head were trained with a 5x higher learning rate of $1 \times 10^{-4}$ to rapidly optimize candidate ranking and comparison.
3. **Optimization & Stability**:
- **AdamW** optimizer with weight decay ($0.01$).
- Cosine Annealing learning rate schedule preceded by linear warmup.
- FP16 Automatic Mixed Precision (AMP) with gradient norm clipping ($1.0$).
4. **Compute & Runtime**:
- Micro-batch size of 4 with 4 gradient accumulation steps (effective batch size of 16).
- 3 epochs completed in **10.4 minutes (621.74 seconds)** on a single consumer NVIDIA RTX 4090 GPU.
---
## 🚀 Quickstart & Inference (Hugging Face AutoModel)
Install dependencies:
```bash
pip install torch transformers
```
Run inference in 3 lines of code:
```python
from transformers import AutoModel, AutoTokenizer
# 1. Load model and tokenizer directly from Hugging Face Hub
model_id = "TurkishCodeMan/laya-tr"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, trust_remote_code=True)
# 2. Define question and candidate options
question = "Türkiye Cumhuriyeti hangi yılda ilan edilmiştir?"
options = {
"A": "1920",
"B": "1923",
"C": "1938",
"D": "1919"
}
# 3. Predict in sub-10ms
result = model.decide(question=question, options=options, tokenizer=tokenizer)
print("Prediction :", result["prediction"]) # B
print("Option :", result["selected_option"]) # B: 1923
print("Confidence :", f"{result['confidence']*100:.2f}%")
print("Latency :", f"{result['latency_ms']:.2f} ms")
print("Full Probs :", result["probabilities"])
```
---
## 🔄 Architectural Comparison
| Dimension | Generative Autoregressive LLMs (7B - 70B) | **Laya-TR (322M Decision Model)** |
| :--- | :---: | :---: |
| **Inference Paradigm** | Sequential token-by-token generation | **Single parallel neural forward pass** |
| **Latency per Decision** | 500 ms – 3,000 ms | **~9.78 ms (<10 ms)** ⚡ |
| **VRAM Consumption** | 16 GB – 80 GB | **< 1.5 GB** |
| **Throughput** | 1 – 10 requests / sec | **~100+ decisions / sec** |
| **Primary Use Cases** | Text generation, creative writing, chat | **Routing, classification, agent decisions, QA** |
---
## ⚖️ License & Acknowledgments
- **License**: Apache 2.0
- **Model Author**: [TurkishCodeMan](https://huggingface.co/TurkishCodeMan)
- **Base Architecture**: ConvAI Laya & ModernBERT
- **Benchmark Reference**: [bezir/MMLU-pro-TR](https://huggingface.co/datasets/bezir/MMLU-pro-TR)