TurkishCodeMan commited on
Commit
6b96146
·
verified ·
1 Parent(s): af5dd93

Release Turkish Laya-TR non-autoregressive decision model with native AutoModel support

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,147 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - tr
4
+ - en
5
+ license: apache-2.0
6
+ tags:
7
+ - decision-model
8
+ - non-autoregressive
9
+ - modernbert
10
+ - mmbert
11
+ - turkish
12
+ - reasoning
13
+ - mmlu-pro
14
+ - fast-inference
15
+ pipeline_tag: text-classification
16
+ widget:
17
+ - text: "Türkiye Cumhuriyeti hangi yılda ilan edilmiştir?"
18
+ ---
19
+
20
+ # 🇹🇷 Laya-TR: Non-Autoregressive Decision & Reasoning Model for Turkish
21
+
22
+ **Laya-TR** is the first Turkish **non-autoregressive decision and reasoning model**, specifically engineered for ultra-low-latency decision making, candidate selection, and agentic routing.
23
+
24
+ While conventional generative Large Language Models (LLMs) generate tokens sequentially—taking hundreds to thousands of milliseconds to reach a decision—**Laya-TR evaluates all candidate options and context simultaneously in a single parallel neural forward pass with sub-10ms latency (<10 ms).**
25
+
26
+ With native Hugging Face `AutoModel` support, developers can deploy and run Laya-TR with standard `transformers` code without having to manage external architecture files or local repositories.
27
+
28
+ ---
29
+
30
+ ## ⚡ Key Highlights
31
+
32
+ - **Architecture**: 22-layer `mmBERT-base` (ModernBERT backbone with GeGLU, Rotary Position Embeddings, and sliding-window attention) + 2-layer Decision Transformer Head + Shared Option Marker Scorer + Act/Escalate Head.
33
+ - **Model Size**: ~322 Million parameters (Compact, edge-ready, and exceptionally fast on a single GPU or CPU).
34
+ - **Inference Latency**: **~9.78 ms** per question on a single GPU (**95 – 162 decisions/second** throughput).
35
+ - **Training Efficiency**: Trained in **just 10.4 minutes (621 seconds)** on a single NVIDIA GeForce RTX 4090 GPU.
36
+ - **Seamless Hugging Face Integration**: Fully compatible with `AutoModel.from_pretrained("TurkishCodeMan/laya-tr", trust_remote_code=True)`.
37
+
38
+ ---
39
+
40
+ ## 📊 Comprehensive Benchmark: MMLU-Pro TR
41
+
42
+ Laya-TR was evaluated on the complete test split of [**bezir/MMLU-pro-TR**](https://huggingface.co/datasets/bezir/MMLU-pro-TR), representing the most demanding Turkish academic decision and multi-choice reasoning benchmark (**11,842 Questions, 10 Choices A–J per question**).
43
+
44
+ > 💡 **Baseline Context:** On a 10-choice multiple-choice test, the random guessing baseline is **10.00%**.
45
+
46
+ | Metric / Model | Base Laya (Zero-Shot) | **Laya-TR (Fine-Tuned)** | Net Gain / Relative Improvement |
47
+ | :--- | :---: | :---: | :---: |
48
+ | **Total Test Questions** | 11,842 | 11,842 | Full Test Split |
49
+ | **Correct Answers** | 1,383 / 11,842 | **2,238 / 11,842** | **+855 More Correct Answers** |
50
+ | **Overall Accuracy** | **11.68%** | **18.90%** | **+7.22% Net (+61.82% Relative Jump)** 🚀 |
51
+ | **Average Latency** | 5.54 ms | **9.78 ms** | Sub-10 Millisecond Decisions |
52
+ | **Throughput** | 162.0 q/s | **95.7 q/s** | Real-Time Production Ready |
53
+
54
+ ### 📚 Category Breakdown Across All 14 Disciplines
55
+
56
+ | Category | Total Questions | Base Laya (Zero-Shot) | **Laya-TR (Fine-Tuned)** | Relative Improvement |
57
+ | :--- | :---: | :---: | :---: | :---: |
58
+ | 🧠 **Psychology** | 780 | 11.28% | **26.54%** | **+135.3%** 🚀 |
59
+ | 🔬 **Biology** | 714 | 13.31% | **26.47%** | **+98.9%** 🚀 |
60
+ | 🏛️ **History** | 342 | 13.16% | **24.56%** | **+86.6%** 🚀 |
61
+ | 🩺 **Health & Medicine** | 800 | 11.50% | **24.00%** | **+108.7%** 🚀 |
62
+ | 📈 **Economics** | 830 | 14.58% | **23.73%** | **+62.8%** 🚀 |
63
+ | 🌐 **Other** | 915 | 10.82% | **22.51%** | **+108.0%** 🚀 |
64
+ | 📜 **Philosophy** | 479 | 12.11% | **20.46%** | **+69.0%** |
65
+ | 💻 **Computer Science** | 397 | 11.84% | **20.15%** | **+70.2%** |
66
+ | ⚖️ **Law** | 1086 | 11.42% | **17.50%** | **+53.2%** |
67
+ | 💼 **Business** | 774 | 12.02% | **16.41%** | **+36.5%** |
68
+ | 🧪 **Chemistry** | 1126 | 12.43% | **14.56%** | **+17.1%** |
69
+ | 📐 **Mathematics** | 1345 | 11.08% | **14.05%** | **+26.8%** |
70
+ | ⚙️ **Engineering** | 965 | 11.92% | **13.99%** | **+17.4%** |
71
+ | ⚛️ **Physics** | 1289 | 9.08% | **13.96%** | **+53.7%** |
72
+
73
+ ---
74
+
75
+ ## 🛠️ Training Strategy & Methodology
76
+
77
+ 1. **Curated Turkish Decision & Reasoning Corpus**:
78
+ - The model was fine-tuned on a curated, high-quality Turkish multi-domain decision dataset comprising **15,459 samples** covering sciences, humanities, law, economics, and analytical reasoning.
79
+ 2. **Differential Learning Rates**:
80
+ - To safeguard the rich multilingual language representations of the `mmBERT-base` ModernBERT encoder, the backbone was fine-tuned with a conservative learning rate of $2 \times 10^{-5}$.
81
+ - The Decision Transformer layers and the Option Marker Scorer head were trained with a 5x higher learning rate of $1 \times 10^{-4}$ to rapidly optimize candidate ranking and comparison.
82
+ 3. **Optimization & Stability**:
83
+ - **AdamW** optimizer with weight decay ($0.01$).
84
+ - Cosine Annealing learning rate schedule preceded by linear warmup.
85
+ - FP16 Automatic Mixed Precision (AMP) with gradient norm clipping ($1.0$).
86
+ 4. **Compute & Runtime**:
87
+ - Micro-batch size of 4 with 4 gradient accumulation steps (effective batch size of 16).
88
+ - 3 epochs completed in **10.4 minutes (621.74 seconds)** on a single consumer NVIDIA RTX 4090 GPU.
89
+
90
+ ---
91
+
92
+ ## 🚀 Quickstart & Inference (Hugging Face AutoModel)
93
+
94
+ Install dependencies:
95
+ ```bash
96
+ pip install torch transformers
97
+ ```
98
+
99
+ Run inference in 3 lines of code:
100
+
101
+ ```python
102
+ from transformers import AutoModel, AutoTokenizer
103
+
104
+ # 1. Load model and tokenizer directly from Hugging Face Hub
105
+ model_id = "TurkishCodeMan/laya-tr"
106
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
107
+ model = AutoModel.from_pretrained(model_id, trust_remote_code=True)
108
+
109
+ # 2. Define question and candidate options
110
+ question = "Türkiye Cumhuriyeti hangi yılda ilan edilmiştir?"
111
+ options = {
112
+ "A": "1920",
113
+ "B": "1923",
114
+ "C": "1938",
115
+ "D": "1919"
116
+ }
117
+
118
+ # 3. Predict in sub-10ms
119
+ result = model.decide(question=question, options=options, tokenizer=tokenizer)
120
+
121
+ print("Prediction :", result["prediction"]) # B
122
+ print("Option :", result["selected_option"]) # B: 1923
123
+ print("Confidence :", f"{result['confidence']*100:.2f}%")
124
+ print("Latency :", f"{result['latency_ms']:.2f} ms")
125
+ print("Full Probs :", result["probabilities"])
126
+ ```
127
+
128
+ ---
129
+
130
+ ## 🔄 Architectural Comparison
131
+
132
+ | Dimension | Generative Autoregressive LLMs (7B - 70B) | **Laya-TR (322M Decision Model)** |
133
+ | :--- | :---: | :---: |
134
+ | **Inference Paradigm** | Sequential token-by-token generation | **Single parallel neural forward pass** |
135
+ | **Latency per Decision** | 500 ms – 3,000 ms | **~9.78 ms (<10 ms)** ⚡ |
136
+ | **VRAM Consumption** | 16 GB – 80 GB | **< 1.5 GB** |
137
+ | **Throughput** | 1 – 10 requests / sec | **~100+ decisions / sec** |
138
+ | **Primary Use Cases** | Text generation, creative writing, chat | **Routing, classification, agent decisions, QA** |
139
+
140
+ ---
141
+
142
+ ## ⚖️ License & Acknowledgments
143
+
144
+ - **License**: Apache 2.0
145
+ - **Model Author**: [TurkishCodeMan](https://huggingface.co/TurkishCodeMan)
146
+ - **Base Architecture**: ConvAI Laya & ModernBERT
147
+ - **Benchmark Reference**: [bezir/MMLU-pro-TR](https://huggingface.co/datasets/bezir/MMLU-pro-TR)
config.json ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "LayaDecisionModel"
4
+ ],
5
+ "model_type": "laya",
6
+ "auto_map": {
7
+ "AutoConfig": "configuration_laya.LayaConfig",
8
+ "AutoModel": "modeling_laya.LayaDecisionModel"
9
+ },
10
+ "vocab_size": 256000,
11
+ "hidden_size": 768,
12
+ "intermediate_size": 1152,
13
+ "num_hidden_layers": 22,
14
+ "num_attention_heads": 12,
15
+ "hidden_activation": "gelu",
16
+ "norm_eps": 1e-05,
17
+ "norm_bias": false,
18
+ "attention_bias": false,
19
+ "mlp_bias": false,
20
+ "rope_theta": 160000.0,
21
+ "local_attention": 128,
22
+ "global_attn_every_n_layers": 3,
23
+ "max_position_embeddings": 8192,
24
+ "pad_token_id": 0,
25
+ "cls_token_id": 1,
26
+ "sep_token_id": 1,
27
+ "bos_token_id": 2,
28
+ "mask_token_id": 4,
29
+ "head_layers": 2,
30
+ "head_ff_dim": 3072,
31
+ "n_act": 2,
32
+ "num_question_types": 3,
33
+ "max_len": 1024,
34
+ "head_max_len": 256,
35
+ "torch_dtype": "float32"
36
+ }
configuration_laya.py ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from transformers import PretrainedConfig
2
+
3
+ class LayaConfig(PretrainedConfig):
4
+ model_type = "laya"
5
+
6
+ def __init__(
7
+ self,
8
+ vocab_size: int = 256000,
9
+ hidden_size: int = 768,
10
+ intermediate_size: int = 1152,
11
+ num_hidden_layers: int = 22,
12
+ num_attention_heads: int = 12,
13
+ hidden_activation: str = "gelu",
14
+ norm_eps: float = 1e-5,
15
+ norm_bias: bool = False,
16
+ attention_bias: bool = False,
17
+ mlp_bias: bool = False,
18
+ rope_theta: float = 160000.0,
19
+ local_attention: int = 128,
20
+ global_attn_every_n_layers: int = 3,
21
+ max_position_embeddings: int = 8192,
22
+ pad_token_id: int = 0,
23
+ cls_token_id: int = 1,
24
+ sep_token_id: int = 1,
25
+ bos_token_id: int = 2,
26
+ mask_token_id: int = 4,
27
+ head_layers: int = 2,
28
+ head_ff_dim: int = 3072,
29
+ n_act: int = 2,
30
+ num_question_types: int = 3,
31
+ max_len: int = 1024,
32
+ head_max_len: int = 256,
33
+ **kwargs
34
+ ):
35
+ super().__init__(
36
+ pad_token_id=pad_token_id,
37
+ bos_token_id=bos_token_id,
38
+ sep_token_id=sep_token_id,
39
+ **kwargs
40
+ )
41
+ self.vocab_size = vocab_size
42
+ self.hidden_size = hidden_size
43
+ self.intermediate_size = intermediate_size
44
+ self.num_hidden_layers = num_hidden_layers
45
+ self.num_attention_heads = num_attention_heads
46
+ self.hidden_activation = hidden_activation
47
+ self.norm_eps = norm_eps
48
+ self.norm_bias = norm_bias
49
+ self.attention_bias = attention_bias
50
+ self.mlp_bias = mlp_bias
51
+ self.rope_theta = rope_theta
52
+ self.local_attention = local_attention
53
+ self.global_attn_every_n_layers = global_attn_every_n_layers
54
+ self.max_position_embeddings = max_position_embeddings
55
+ self.cls_token_id = cls_token_id
56
+ self.mask_token_id = mask_token_id
57
+ self.head_layers = head_layers
58
+ self.head_ff_dim = head_ff_dim
59
+ self.n_act = n_act
60
+ self.num_question_types = num_question_types
61
+ self.max_len = max_len
62
+ self.head_max_len = head_max_len
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a545935a817d49ce9be6763cf4484b4aa6af47630ccd63c891ebd594265b1513
3
+ size 1287653720
modeling_laya.py ADDED
@@ -0,0 +1,369 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ Laya-TR: Non-Autoregressive Turkish Decision & Reasoning Model
3
+ Hugging Face PreTrainedModel uyumlu mimari tanımı.
4
+ """
5
+
6
+ import math
7
+ import time
8
+ from typing import Any, Dict, List, Optional, Tuple, Union
9
+
10
+ import torch
11
+ import torch.nn as nn
12
+ import torch.nn.functional as F
13
+ from transformers import PreTrainedModel, AutoTokenizer
14
+
15
+ try:
16
+ from .configuration_laya import LayaConfig
17
+ except ImportError:
18
+ from configuration_laya import LayaConfig
19
+
20
+
21
+ # -----------------------------------------------------------------------------
22
+ # 1. RoPE (Rotary Position Embeddings)
23
+ # -----------------------------------------------------------------------------
24
+ def rotate_half(x: torch.Tensor) -> torch.Tensor:
25
+ x1 = x[..., : x.shape[-1] // 2]
26
+ x2 = x[..., x.shape[-1] // 2 :]
27
+ return torch.cat((-x2, x1), dim=-1)
28
+
29
+
30
+ def apply_rotary_pos_emb(q: torch.Tensor, k: torch.Tensor, cos: torch.Tensor, sin: torch.Tensor) -> Tuple[torch.Tensor, torch.Tensor]:
31
+ orig_dtype = q.dtype
32
+ q_float = q.float()
33
+ k_float = k.float()
34
+ q_out = (q_float * cos) + (rotate_half(q_float) * sin)
35
+ k_out = (k_float * cos) + (rotate_half(k_float) * sin)
36
+ return q_out.to(orig_dtype), k_out.to(orig_dtype)
37
+
38
+
39
+ class ModernBertRotaryEmbedding(nn.Module):
40
+ def __init__(self, config: LayaConfig):
41
+ super().__init__()
42
+ self.dim = config.hidden_size // config.num_attention_heads
43
+ self.max_seq_len = config.max_position_embeddings
44
+ self.theta = config.rope_theta
45
+
46
+ inv_freq = 1.0 / (self.theta ** (torch.arange(0, self.dim, 2, dtype=torch.float32) / self.dim))
47
+ self.register_buffer("inv_freq", inv_freq, persistent=False)
48
+
49
+ def forward(self, x: torch.Tensor, seq_len: int) -> Tuple[torch.Tensor, torch.Tensor]:
50
+ t = torch.arange(seq_len, device=x.device, dtype=torch.float32)
51
+ freqs = torch.outer(t, self.inv_freq.to(device=x.device))
52
+ emb = torch.cat((freqs, freqs), dim=-1)
53
+ cos = emb.cos().unsqueeze(0).unsqueeze(1)
54
+ sin = emb.sin().unsqueeze(0).unsqueeze(1)
55
+ return cos.to(x.dtype), sin.to(x.dtype)
56
+
57
+
58
+ # -----------------------------------------------------------------------------
59
+ # 2. ModernBERT Embeddings & MLP
60
+ # -----------------------------------------------------------------------------
61
+ class ModernBertEmbeddings(nn.Module):
62
+ def __init__(self, config: LayaConfig):
63
+ super().__init__()
64
+ self.tok_embeddings = nn.Embedding(
65
+ config.vocab_size, config.hidden_size, padding_idx=config.pad_token_id
66
+ )
67
+ self.norm = nn.LayerNorm(config.hidden_size, eps=config.norm_eps, bias=config.norm_bias)
68
+
69
+ def forward(self, input_ids: torch.Tensor) -> torch.Tensor:
70
+ return self.norm(self.tok_embeddings(input_ids))
71
+
72
+
73
+ class ModernBertMLP(nn.Module):
74
+ def __init__(self, config: LayaConfig):
75
+ super().__init__()
76
+ self.Wi = nn.Linear(config.hidden_size, config.intermediate_size * 2, bias=config.mlp_bias)
77
+ self.Wo = nn.Linear(config.intermediate_size, config.hidden_size, bias=config.mlp_bias)
78
+
79
+ def forward(self, hidden_states: torch.Tensor) -> torch.Tensor:
80
+ input_gate, hidden = self.Wi(hidden_states).chunk(2, dim=-1)
81
+ return self.Wo(F.gelu(input_gate) * hidden)
82
+
83
+
84
+ # -----------------------------------------------------------------------------
85
+ # 3. ModernBERT Attention & Encoder Layer
86
+ # -----------------------------------------------------------------------------
87
+ class ModernBertAttention(nn.Module):
88
+ def __init__(self, config: LayaConfig, layer_idx: int):
89
+ super().__init__()
90
+ self.hidden_size = config.hidden_size
91
+ self.num_heads = config.num_attention_heads
92
+ self.head_dim = self.hidden_size // self.num_heads
93
+ self.layer_idx = layer_idx
94
+ self.is_global = (layer_idx % config.global_attn_every_n_layers == 0)
95
+ self.local_window = config.local_attention
96
+
97
+ self.Wqkv = nn.Linear(config.hidden_size, 3 * config.hidden_size, bias=config.attention_bias)
98
+ self.Wo = nn.Linear(config.hidden_size, config.hidden_size, bias=config.attention_bias)
99
+
100
+ def forward(
101
+ self,
102
+ hidden_states: torch.Tensor,
103
+ position_embeddings: Tuple[torch.Tensor, torch.Tensor],
104
+ attention_mask: Optional[torch.Tensor] = None
105
+ ) -> torch.Tensor:
106
+ B, S, _ = hidden_states.shape
107
+ cos, sin = position_embeddings
108
+
109
+ qkv = self.Wqkv(hidden_states)
110
+ q, k, v = qkv.chunk(3, dim=-1)
111
+
112
+ q = q.view(B, S, self.num_heads, self.head_dim).transpose(1, 2)
113
+ k = k.view(B, S, self.num_heads, self.head_dim).transpose(1, 2)
114
+ v = v.view(B, S, self.num_heads, self.head_dim).transpose(1, 2)
115
+
116
+ q, k = apply_rotary_pos_emb(q, k, cos, sin)
117
+
118
+ scale = 1.0 / math.sqrt(self.head_dim)
119
+ attn_scores = torch.matmul(q, k.transpose(-2, -1)) * scale
120
+
121
+ if not self.is_global and self.local_window > 0:
122
+ row_idx = torch.arange(S, device=hidden_states.device).unsqueeze(1)
123
+ col_idx = torch.arange(S, device=hidden_states.device).unsqueeze(0)
124
+ sliding_mask = (col_idx < (row_idx - self.local_window)) | (col_idx > (row_idx + self.local_window))
125
+ attn_scores = attn_scores.masked_fill(sliding_mask.unsqueeze(0).unsqueeze(0), -1e4)
126
+
127
+ if attention_mask is not None:
128
+ if attention_mask.dim() == 2:
129
+ pad_mask = attention_mask.bool().unsqueeze(1).unsqueeze(2)
130
+ else:
131
+ pad_mask = attention_mask.bool()
132
+ attn_scores = attn_scores.masked_fill(~pad_mask, -1e4)
133
+
134
+ attn_weights = F.softmax(attn_scores, dim=-1, dtype=torch.float32).to(q.dtype)
135
+ attn_out = torch.matmul(attn_weights, v)
136
+ attn_out = attn_out.transpose(1, 2).contiguous().view(B, S, self.hidden_size)
137
+
138
+ return self.Wo(attn_out)
139
+
140
+
141
+ class ModernBertEncoderLayer(nn.Module):
142
+ def __init__(self, config: LayaConfig, layer_idx: int):
143
+ super().__init__()
144
+ self.layer_idx = layer_idx
145
+ if layer_idx == 0:
146
+ self.attn_norm = nn.Identity()
147
+ else:
148
+ self.attn_norm = nn.LayerNorm(config.hidden_size, eps=config.norm_eps, bias=config.norm_bias)
149
+
150
+ self.attn = ModernBertAttention(config, layer_idx=layer_idx)
151
+ self.mlp_norm = nn.LayerNorm(config.hidden_size, eps=config.norm_eps, bias=config.norm_bias)
152
+ self.mlp = ModernBertMLP(config)
153
+
154
+ def forward(
155
+ self,
156
+ hidden_states: torch.Tensor,
157
+ position_embeddings: Tuple[torch.Tensor, torch.Tensor],
158
+ attention_mask: Optional[torch.Tensor] = None
159
+ ) -> torch.Tensor:
160
+ attn_out = self.attn(
161
+ self.attn_norm(hidden_states),
162
+ position_embeddings=position_embeddings,
163
+ attention_mask=attention_mask
164
+ )
165
+ hidden_states = hidden_states + attn_out
166
+ mlp_out = self.mlp(self.mlp_norm(hidden_states))
167
+ hidden_states = hidden_states + mlp_out
168
+ return hidden_states
169
+
170
+
171
+ class ModernBertEncoder(nn.Module):
172
+ def __init__(self, config: LayaConfig):
173
+ super().__init__()
174
+ self.config = config
175
+ self.embeddings = ModernBertEmbeddings(config)
176
+ self.rotary_emb = ModernBertRotaryEmbedding(config)
177
+ self.layers = nn.ModuleList([
178
+ ModernBertEncoderLayer(config, layer_idx=l)
179
+ for l in range(config.num_hidden_layers)
180
+ ])
181
+ self.final_norm = nn.LayerNorm(config.hidden_size, eps=config.norm_eps, bias=config.norm_bias)
182
+
183
+ def forward(self, input_ids: torch.Tensor, attention_mask: Optional[torch.Tensor] = None) -> torch.Tensor:
184
+ B, S = input_ids.shape
185
+ hidden_states = self.embeddings(input_ids)
186
+ position_embeddings = self.rotary_emb(hidden_states, seq_len=S)
187
+
188
+ for layer in self.layers:
189
+ hidden_states = layer(
190
+ hidden_states,
191
+ position_embeddings=position_embeddings,
192
+ attention_mask=attention_mask
193
+ )
194
+
195
+ hidden_states = self.final_norm(hidden_states)
196
+ return hidden_states
197
+
198
+
199
+ # -----------------------------------------------------------------------------
200
+ # 4. Decision Transformer Head, Scorer & Act Head
201
+ # -----------------------------------------------------------------------------
202
+ class DecisionTransformerHead(nn.Module):
203
+ def __init__(self, config: LayaConfig):
204
+ super().__init__()
205
+ d = config.hidden_size
206
+ nhead = config.num_attention_heads
207
+ d_ff = config.head_ff_dim
208
+
209
+ self.layers = nn.ModuleList([
210
+ nn.TransformerEncoderLayer(
211
+ d_model=d,
212
+ nhead=nhead,
213
+ dim_feedforward=d_ff,
214
+ dropout=0.1,
215
+ batch_first=True,
216
+ norm_first=True
217
+ )
218
+ for _ in range(config.head_layers)
219
+ ])
220
+
221
+ def forward(self, x: torch.Tensor, attention_mask: Optional[torch.Tensor] = None) -> torch.Tensor:
222
+ pad_mask = ~attention_mask.bool() if attention_mask is not None else None
223
+ for layer in self.layers:
224
+ x = layer(x, src_key_padding_mask=pad_mask)
225
+ return x
226
+
227
+
228
+ # -----------------------------------------------------------------------------
229
+ # 5. Hugging Face PreTrainedModel Uyumlu LayaDecisionModel
230
+ # -----------------------------------------------------------------------------
231
+ class LayaDecisionModel(PreTrainedModel):
232
+ config_class = LayaConfig
233
+ base_model_prefix = "laya"
234
+ supports_gradient_checkpointing = True
235
+
236
+ def __init__(self, config: LayaConfig):
237
+ super().__init__(config)
238
+ d = config.hidden_size
239
+
240
+ self.encoder = ModernBertEncoder(config)
241
+ self.type_emb = nn.Embedding(config.num_question_types, d)
242
+ self.head = DecisionTransformerHead(config) if config.head_layers > 0 else None
243
+
244
+ self.scorer = nn.Sequential(
245
+ nn.LayerNorm(d),
246
+ nn.Linear(d, d),
247
+ nn.GELU(),
248
+ nn.Linear(d, 1)
249
+ )
250
+
251
+ self.act_head = nn.Sequential(
252
+ nn.Linear(d + 4, 256),
253
+ nn.GELU(),
254
+ nn.Linear(256, config.n_act)
255
+ )
256
+
257
+ self.register_buffer("temperature", torch.ones(3))
258
+ self.post_init()
259
+
260
+ def forward(
261
+ self,
262
+ input_ids: torch.Tensor,
263
+ attention_mask: torch.Tensor,
264
+ marker_pos: torch.Tensor,
265
+ marker_mask: torch.Tensor,
266
+ qtype: torch.Tensor
267
+ ) -> Tuple[torch.Tensor, torch.Tensor]:
268
+ h = self.encoder(input_ids=input_ids, attention_mask=attention_mask)
269
+ h = h + self.type_emb(qtype)[:, None, :]
270
+
271
+ if self.head is not None:
272
+ h = self.head(h, attention_mask=attention_mask)
273
+
274
+ idx = marker_pos.clamp(min=0)[:, :, None].expand(-1, -1, h.size(-1))
275
+ m = torch.gather(h, 1, idx)
276
+
277
+ logits = self.scorer(m).squeeze(-1).float()
278
+ logits = logits.masked_fill(~marker_mask, -1e4)
279
+
280
+ p = torch.softmax(logits.detach(), dim=-1)
281
+ k = marker_mask.sum(-1).clamp(min=2).float()
282
+ ent = -(p * torch.log(p.clamp_min(1e-9))).sum(-1) / torch.log(k)
283
+
284
+ if p.size(-1) >= 2:
285
+ top2 = p.topk(2, dim=-1).values
286
+ else:
287
+ top1 = p.topk(1, dim=-1).values
288
+ top2 = torch.cat([top1, torch.zeros_like(top1)], dim=-1)
289
+
290
+ feats = torch.stack([top2[:, 0], top2[:, 0] - top2[:, 1], ent, k / 255.0], dim=-1)
291
+ pooled = h[:, 0]
292
+ act_input = torch.cat([pooled, feats.to(dtype=h.dtype)], dim=-1)
293
+ act_logits = self.act_head(act_input)
294
+
295
+ return logits, act_logits
296
+
297
+ @torch.no_grad()
298
+ def decide(
299
+ self,
300
+ question: str,
301
+ options: Union[List[str], Dict[str, str]],
302
+ tokenizer: Optional[AutoTokenizer] = None,
303
+ context: Optional[str] = None
304
+ ) -> Dict[str, Any]:
305
+ """
306
+ Kullanıcıların tek satırda AutoModel üzerinden sub-10ms karar almasını sağlar.
307
+ """
308
+ if tokenizer is None:
309
+ tokenizer = AutoTokenizer.from_pretrained("jhu-clsp/mmBERT-base")
310
+
311
+ t0 = time.perf_counter()
312
+ device = next(self.parameters()).device
313
+
314
+ if isinstance(options, dict):
315
+ opt_labels = list(options.keys())
316
+ opt_texts = [f"{k}: {v}" if v else k for k, v in options.items()]
317
+ else:
318
+ opt_labels = [chr(65 + i) for i in range(len(options))]
319
+ opt_texts = [f"{lbl}: {opt}" for lbl, opt in zip(opt_labels, options)]
320
+
321
+ mask_tok = tokenizer.mask_token
322
+ head_ids = tokenizer(f"choice question: {question}", add_special_tokens=False)["input_ids"]
323
+
324
+ opt_ids = []
325
+ for text in opt_texts:
326
+ opt_ids.append(tokenizer(f"{mask_tok} {text}", add_special_tokens=False)["input_ids"])
327
+
328
+ cls_id = tokenizer.cls_token_id or 1
329
+ sep_id = tokenizer.sep_token_id or 1
330
+
331
+ seq = [cls_id] + head_ids + [sep_id]
332
+ markers = []
333
+ for o_ids in opt_ids:
334
+ markers.append(len(seq))
335
+ seq.extend(o_ids)
336
+ seq.append(sep_id)
337
+
338
+ if context:
339
+ ctx_ids = tokenizer(str(context), add_special_tokens=False)["input_ids"][:512]
340
+ seq.extend(ctx_ids)
341
+ seq.append(sep_id)
342
+
343
+ input_ids = torch.tensor([seq], dtype=torch.long, device=device)
344
+ attention_mask = torch.ones_like(input_ids)
345
+ marker_pos = torch.tensor([markers], dtype=torch.long, device=device)
346
+ marker_mask = torch.ones_like(marker_pos, dtype=torch.bool)
347
+ qtype = torch.tensor([0], dtype=torch.long, device=device)
348
+
349
+ logits, act_logits = self(
350
+ input_ids=input_ids,
351
+ attention_mask=attention_mask,
352
+ marker_pos=marker_pos,
353
+ marker_mask=marker_mask,
354
+ qtype=qtype
355
+ )
356
+
357
+ probs = F.softmax(logits[0], dim=-1).cpu().tolist()
358
+ best_idx = int(torch.argmax(logits[0]).item())
359
+ elapsed_ms = (time.perf_counter() - t0) * 1000
360
+
361
+ prob_map = {lbl: round(p, 4) for lbl, p in zip(opt_labels, probs)}
362
+
363
+ return {
364
+ "prediction": opt_labels[best_idx],
365
+ "selected_option": opt_texts[best_idx],
366
+ "confidence": round(probs[best_idx], 4),
367
+ "probabilities": prob_map,
368
+ "latency_ms": round(elapsed_ms, 2)
369
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:609d8f4c067cd3950f88594c5a802616cea245823836ef5848ee4fc40aab5b6f
3
+ size 34363188
tokenizer_config.json ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "bos_token": "<bos>",
4
+ "clean_up_tokenization_spaces": false,
5
+ "cls_token": "<bos>",
6
+ "eos_token": "<eos>",
7
+ "extra_special_tokens": [
8
+ "<start_of_turn>",
9
+ "<end_of_turn>"
10
+ ],
11
+ "is_local": false,
12
+ "local_files_only": false,
13
+ "mask_token": "<mask>",
14
+ "model_input_names": [
15
+ "input_ids",
16
+ "attention_mask"
17
+ ],
18
+ "model_max_length": 8192,
19
+ "pad_token": "<pad>",
20
+ "padding_side": "right",
21
+ "sep_token": "<eos>",
22
+ "spaces_between_special_tokens": false,
23
+ "tokenizer_class": "TokenizersBackend",
24
+ "unk_token": "<unk>"
25
+ }