File size: 17,412 Bytes
d97b72b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d3ca140
 
149fb12
d97b72b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b0657a2
d97b72b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b0657a2
 
d97b72b
b0657a2
d97b72b
b0657a2
 
 
d97b72b
 
 
 
 
 
 
 
b0657a2
d97b72b
 
 
b0657a2
d97b72b
 
 
 
 
 
 
 
 
 
 
 
 
 
b0657a2
d97b72b
b0657a2
 
 
d97b72b
 
 
b0657a2
d97b72b
b0657a2
 
 
d97b72b
b0657a2
 
 
d97b72b
b0657a2
 
 
d97b72b
b0657a2
 
 
d97b72b
b0657a2
 
 
d97b72b
b0657a2
d97b72b
b0657a2
 
 
d97b72b
 
 
b0657a2
d97b72b
b0657a2
 
 
d97b72b
b0657a2
d97b72b
b0657a2
 
 
d97b72b
 
 
b0657a2
 
 
d97b72b
 
 
 
 
 
 
 
b0657a2
 
d97b72b
b0657a2
d97b72b
 
 
 
 
b0657a2
 
 
d97b72b
b0657a2
d97b72b
b0657a2
 
 
d97b72b
b0657a2
d97b72b
b0657a2
 
 
d97b72b
 
 
 
 
b0657a2
d97b72b
b0657a2
 
 
d97b72b
 
 
b0657a2
 
 
d97b72b
b0657a2
 
 
d97b72b
 
 
 
 
b0657a2
 
 
d97b72b
b0657a2
 
 
d97b72b
b0657a2
 
 
d97b72b
b0657a2
d97b72b
b0657a2
 
 
d97b72b
b0657a2
 
 
d97b72b
b0657a2
 
 
d97b72b
 
 
b0657a2
d97b72b
b0657a2
d97b72b
b0657a2
 
 
d97b72b
b0657a2
 
 
d97b72b
 
 
b0657a2
 
 
d97b72b
 
 
 
 
b0657a2
d97b72b
 
 
b0657a2
 
 
d97b72b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1978c51
d97b72b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
---
base_model:
  - Qwen/Qwen3.5-9B
  - XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
  - ornith-ai/Ornith-1.5-9B
  - Jackrong/Qwopus3.5-9B-Coder
  - OrionLLM/OxCoder-9B
  - empero-ai/Qwen3.8-9B-Distill
base_model_relation: merge
library_name: transformers
tags:
  - merge
  - ties
  - della
  - model-stock
  - geodella
  - qwen
  - qwen3.5
  - causal-lm
  - deltanet
  - linear-attention
  - agentic
  - reasoning
  - code
  - swe-bench
license: apache-2.0
language:
  - en
  - zh
pipeline_tag: text-generation
model_type: qwen3_5_text
---

![PentaCoder-9B](https://huggingface.co/pragmaticcs/PentaCoder-9B/resolve/main/assets/banner-dark.png#hf-dark-mode-only)
![PentaCoder-9B](https://huggingface.co/pragmaticcs/PentaCoder-9B/resolve/main/assets/banner-light.png#hf-light-mode-only)

<div align="center">

[![License](https://img.shields.io/badge/License-Apache%202.0-6E56CF?style=for-the-badge)](https://opensource.org/licenses/Apache-2.0)
[![Library](https://img.shields.io/badge/Library-transformers-FFD21E?style=for-the-badge&logo=huggingface&logoColor=black)](https://github.com/huggingface/transformers)
[![Merge Method](https://img.shields.io/badge/Merge%20Method-GeoDELLA--HG-27AE60?style=for-the-badge)](#merge-methodology--mathematical-formulation)
[![Architecture](https://img.shields.io/badge/Architecture-Qwen%203.5%209B%20Dense-2D9CDB?style=for-the-badge)](#architectural-specifications)
[![Attention](https://img.shields.io/badge/Hybrid-Gated%20DeltaNet%20(3:1)-EB5757?style=for-the-badge)](#architectural-specifications)
[![Context](https://img.shields.io/badge/Context%20Window-256k-F2994A?style=for-the-badge)](#architectural-specifications)

</div>

Most sub-10B coding models fail in realistic agentic environments due to a common trade-off: aggressive fine-tuning on synthetic coding instructions improves immediate benchmark pass rates but introduces brittle syntactic degradation and catastrophic repetition loops when shell commands or compiler checks fail.

**PentaCoder** addresses these limitations by uniting five specialized post-trained checkpoints of [**Qwen 3.5 9B**](https://huggingface.co/Qwen/Qwen3.5-9B) via **GeoDELLA** (Geometric Drop-and-Rescale with Task-Covariance De-Biasing and Spectral Norm Anchoring). The model synthesizes the distinct mathematical distributions of each donor:

- **Algorithmic Correctness & Architectural Decomposition** from [**Qwopus3.5-Coder**](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) (Claude 3.5 Opus distillation trajectories).
- **Multi-Turn SWE-bench Planning & Tool Protocol Integrity** from [**MiMo-V2.6-Distill**](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) (Large-scale agentic execution traces).
- **Terminal Execution Discipline & Error-Recovery Heuristics** from [**Ornith-1.5**](https://huggingface.co/ornith-ai/Ornith-1.5-9B) (Reinforcement learning for anti-looping).
- **Low-Level Systems Implementation & Runtime Robustness** from [**OxCoder**](https://huggingface.co/OrionLLM/OxCoder-9B) (Deep API, CLI, and operational coding specialization).
- **Abstract Structural Reasoning & Syntax Grounding** from [**Qwen3.8-Distill**](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) (Distilled frontier chain-of-thought representations).

> [!IMPORTANT]
> The result is a lean, blisteringly fast 9B pure-text causal engine with a native **256k context window** that runs comfortably on consumer GPUs.

---

### Contents

- [Architectural Specifications](#architectural-specifications)
- [Composition & Donor Weighting](#composition--donor-weighting)
- [Merge Methodology & Mathematical Formulation](#merge-methodology--mathematical-formulation)
- [Layer-Stratified Component Policies](#layer-stratified-component-policies)
- [Agentic Chat Template & Operational Directives](#agentic-chat-template--operational-directives)
- [Recommended Generation Parameters](#recommended-generation-parameters)
- [How to Use](#how-to-use)
- [Citation & References](#citation--references)

---

## Architectural Specifications

| Parameter | Specification |
| :--- | :---: |
| **Total Parameters** | 8.8B (Pure Text Backbone) |
| **Architecture Type** | Hybrid Recurrent-Attention Causal LM (`qwen3_5_text`) |
| **Hidden Dimension** (*d*<sub>model</sub>) | 4096 |
| **Intermediate Dimension** (*d*<sub>mlp</sub>) | 12288 (SwiGLU) |
| **Decoder Layers** | 32 |
| **Attention Layout** | 8 Blocks × (3 Gated DeltaNet Linear Layers : 1 Gated Softmax Layer) |
| **Full Attention Layers** | Layers 3, 7, 11, 15, 19, 23, 27, 31 |
| **Linear Attention Configuration** | 16 Key Heads / 32 Value Heads (*d*<sub>k</sub> = *d*<sub>v</sub> = 128) |
| **Full Attention Configuration** | 16 Query Heads / 4 Key-Value Heads (GQA, *d*<sub>h</sub> = 256) |
| **Rotary Position Embedding (RoPE)** | 1D Partial RoPE (θ = 10⁷, Factor = 0.25 → 64 dimensions) |
| **Context Window Length** | 262,144 tokens (256k) |
| **Native Precision** | `bfloat16` |
| **Vocabulary Size** | 248,320 (Padded for Fill-In-The-Middle and Tool Tokens) |

---

## Composition & Donor Weighting

The foundational weights of [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) serve as the shared topological base (*W*₀). Five specialized donor checkpoints provide non-overlapping task vectors mapped across normalized layer depth *u* ∈ [0, 1]:

| Model | Primary Focus | Depth Target |
| :--- | :--- | :---: |
| [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) | Base pre-trained manifold and state-space anchors | Global (*W*₀) |
| [empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) | Abstract token synthesis, reasoning structure, syntax | Lower & Mid Decoders |
| [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) | Typing discipline, algorithm design, functional purity | Mid Decoders (Bell Curve) |
| [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B) | Circuit-breaker recovery, environment feedback integration | Upper-Mid Decoders |
| [OrionLLM/OxCoder-9B](https://huggingface.co/OrionLLM/OxCoder-9B) | Systems engineering, runtime bug localization, CLI tooling | Deep Layers (Ascending Ramp) |
| [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) | SWE-bench multi-step planning, long-context tool invocation | Broad Central Decoders |

---

## Merge Methodology & Mathematical Formulation

The merge was executed using the **GeoDELLA-Coder** pipeline, which solves weight interference through dynamic task-correlation de-biasing, row-wise sign consensus, sub-tensor GQA disentanglement, and power-iteration spectral norm stabilization.

### 1. Task Vector Formulation

For each donor checkpoint *k* ∈ {1, ..., 5}, the parameter update delta τ<sub>*k*</sub> is computed relative to the base anchor *W*₀:

$$
\tau_k = D_k - W_0, \quad k \in \{\text{Qwen3.8}, \text{Qwopus}, \text{Ornith}, \text{OxCoder}, \text{MiMo}\}
$$

### 2. Hybrid-Aware Continuous Depth Modulation

Task vector mixing coefficients are continuously modulated over normalized depth *u* = *l* / (*L* - 1), where *l* ∈ [0, 31] and *L* = 32:

$$
u_{\text{qwen38}}(u) = 0.15 + 0.35 \cos^2\left(\frac{\pi}{2} u\right) + 0.15 \sin^2(\pi u)
$$

$$
u_{\text{qwopus}}(u) = 0.05 + 0.35 \sin^2(\pi u)
$$

$$
u_{\text{ornith}}(u) = 0.05 + 0.35 \sin^2\left(\frac{\pi}{2} u\right)
$$

$$
u_{\text{oxcoder}}(u) = 0.05 + 0.25 \sin^2\left(\frac{\pi}{2} u\right)
$$

$$
u_{\text{mimo}}(u) = 0.05 + 0.15 \sin(\pi u)
$$

The initial coefficients are normalized to form a partition of unity across all layers:

$$
w_k(l) = \frac{u_k(u)}{\sum_{j=1}^5 u_j(u)}, \quad \sum_{k=1}^5 w_k(l) = 1.0
$$

### 3. Dynamic Gram-Matrix Task De-Biasing

Fine-tuned models frequently share underlying distillation datasets, leading to collinearity that can drown out specialized task vectors. To correct for this over-representation, the empirical Gram correlation matrix *G* ∈ ℝ<sup>*K* × *K*</sup> is evaluated per tensor:

$$
G_{ij} = \frac{|\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle|}{\|\tau_i\|_2 \|\tau_j\|_2}
$$

A uniqueness coefficient *U*<sub>*k*</sub> is derived from the inverse column sum of task vector correlations:

$$
U_k = \frac{1}{\sum_{j=1}^K G_{kj}}
$$

The dynamic weights are then re-balanced and normalized:

$$
\tilde{w}_k = \frac{w_k U_k}{\sum_{j=1}^K w_j U_j}
$$

This prevents dataset overlap from suppressing distinct algorithmic representations.

### 4. Asymmetric GQA Disentangled QKV Slicing

Qwen 3.5 9B features asymmetric head ratios in both its linear attention and full attention layers. Standard fused tensor merging causes cross-head pollution by treating routing projections identically to memory projections. 

Fused projection tensors are sliced into their functional sub-matrices prior to merging:
- **Gated DeltaNet Layers (8192 × *d*<sub>model</sub>):** Sliced into Query (2048), Key (2048), and Value (4096).
- **Gated Attention Layers (6144 × *d*<sub>model</sub>):** Sliced into Query (4096), Key (1024), and Value (1024).

Query and Key slices are processed with a conservative retention density (ρ = 0.95) to preserve sharp context routing. Value matrices are processed with adaptive MLP density (ρ = 0.70) to maximize conceptual synthesis. The components are then re-concatenated along the head dimension.

### 5. Neuron-Coherent Row Gating

To eliminate destructive interference in Feed-Forward Networks (MLPs), task vectors are gated at the single-neuron (row) level:

$$
\bar{\tau} = \frac{1}{K} \sum_{k=1}^K \tau_k
$$

For each row *r* of donor delta τ<sub>*k*</sub>, the directional alignment with the consensus mean is evaluated:

$$
\cos \theta_{k, r} = \frac{\langle \tau_{k, r}, \bar{\tau}_r \rangle}{\|\tau_{k, r}\|_2 \|\bar{\tau}_r\|_2 + \epsilon}
$$

Rows exhibiting severe directional opposition (cos θ<sub>*k*, *r*</sub> < -0.10) are masked out:

$$
\hat{\tau}_{k, r} = \tau_{k, r} \cdot \mathbb{I}\left(\cos \theta_{k, r} \ge -0.10\right)
$$

This eliminates opposing gradient vectors that produce incoherent syntax generation.

### 6. Heavy-Tailed DELLA Adaptive Rescaling

Surviving parameters undergo non-linear magnitude-based sampling. Using parameter rank indices *R*<sub>*k*, *ij*</sub> ∈ [0, 1] sorted by absolute magnitude, a Pareto-style retention probability *p*<sub>*k*, *ij*</sub> is established:

$$
p_{k, ij} = p_{\min} + (p_{\max} - p_{\min}) \cdot \sqrt{R_{k, ij}}
$$

Parameters are sampled via a Bernoulli trial and rescaled by their inverse survival probability:

$$
M_{k, ij} \sim \text{Bernoulli}(p_{k, ij})
$$

$$
\tilde{\tau}_{k, ij} = \frac{\hat{\tau}_{k, ij} \odot M_{k, ij}}{p_{k, ij}}
$$

### 7. Coordinate-Wise Sign Consensus & Model Stock Scaling

Directional consensus is determined via weighted sign agreement:

$$
\Gamma = \operatorname{sgn}\left(\sum_{k=1}^K \tilde{w}_k \tilde{\tau}_k\right)
$$

$$
A_k = \mathbb{I}\left(\operatorname{sgn}(\tilde{\tau}_k) = \Gamma\right) \odot \mathbb{I}\left(\tilde{\tau}_k \neq 0\right)
$$

$$
\Delta_{\text{consensus}} = \frac{\sum_{k=1}^K \tilde{\tau}_k \odot A_k}{\sum_{k=1}^K A_k + \epsilon}
$$

The aggregated delta is projected onto the non-linear manifold using the Model Stock analytic scaling factor *t*<sup>*</sup>:

$$
\bar{\rho} = \frac{2}{K(K-1)} \sum_{i < j} \frac{\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle}{\|\tau_i\|_2 \|\tau_j\|_2}
$$

$$
t^* = \frac{K \bar{\rho}}{1 + (K - 1)\bar{\rho}}
$$

$$
\Delta_{\text{final}} = t^* \cdot \Delta_{\text{consensus}}
$$

### 8. Spectral Norm & Attention Entropy Anchoring

Attention projection matrices (*W*<sub>attn</sub>) are vulnerable to spectral explosion during merges, which contracts attention entropy and leads to repetitive generation loops. 

The dominant singular value σ(*W*) is calculated via a three-iteration deterministic power iteration:

$$
v^{(t+1)} = \frac{W^T u^{(t)}}{\|W^T u^{(t)}\|_2}, \quad u^{(t+1)} = \frac{W v^{(t+1)}}{\|W v^{(t+1)}\|_2}
$$

$$
\sigma(W) \approx {u^{(3)}}^T W v^{(3)}
$$

If the merged spectral radius grows more than 5% relative to the base model, it is scaled down:

$$
W_{\text{final}} = \begin{cases} W_{\text{merged}} \cdot \left(\frac{1.05 \cdot \sigma(W_0)}{\sigma(W_{\text{merged}})}\right) & \text{if } \sigma(W_{\text{merged}}) > 1.05 \cdot \sigma(W_0) \\ W_{\text{merged}} & \text{otherwise} \end{cases}
$$

---

## Layer-Stratified Component Policies

| Parameter Class | Target Identifiers | Applied Policy | Density (ρ) | Mathematical Constraints |
| :--- | :--- | :---: | :---: | :--- |
| **Embeddings & LM Head** | `embed_tokens`, `lm_head` | Low-Memory Streaming Blend | 1.0 | Convex iterative accumulation; vocab dimension aligned to 248,320. |
| **Linear State-Space Projections** | `linear_attn.in_proj_qkv` | Asymmetric GQA DELLA | 0.95 (QK) / 0.70 (V) | Sub-tensor slicing; separate routing and associative memory passes. |
| **Self-Attention Projections** | `self_attn.qkv_proj`, `o_proj` | Asymmetric GQA + Spectral Anchor | 0.95 (QK) / 0.70 (V) | Power-iteration clipping prevents σ > 1.05 σ₀. |
| **Feed-Forward Blocks** | `mlp.gate_proj`, `up_proj`, `down_proj` | Neuron-Gated HG-DELLA | 0.50 – 0.70 | Row-wise cosine filtering (cos θ ≥ -0.10); sign consensus. |
| **Recurrent Gates & Normalization** | `A_log`, `norm`, `conv1d` | Convex Parameter Blend | 1.0 | Preservation of *A*<sub>log</sub> ≤ 0 to guarantee BIBO stability. |

---

## Agentic Chat Template & Operational Directives

This model uses the [Improved Chat Template for Qwen 3.x by Olivia Rossi](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x) to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination.

---

## Recommended Generation Parameters

For deterministic software engineering and complex reasoning benchmarks:

| Parameter | Recommended Value | Description |
| :--- | :---: | :--- |
| **Temperature** | `0.6` | Balances strict syntactic validity with algorithmic path exploration. |
| **Top-P** | `0.95` | Eliminates low-probability token tails while preserving alternative logic paths. |
| **Top-K** | `20` | Restricts token candidate pools to prevent architectural syntax drift. |
| **Min-P** | `0.0` (Off) | Disabled in favor of explicit Top-K / Top-P governance. |
| **Repetition Penalty** | `1.0` (Off) | Disabled to prevent syntax degradation in repetitive code patterns (indentation, braces). |
| **Presence Penalty** | `0.0` | Prevents naming mutations across long-context symbol resolution. |

---

## How to Use

### Serving via vLLM

```bash
vllm serve pragmaticcs/PentaCoder-9B \
  --dtype bfloat16 \
  --max-model-len 65536 \
  --gpu-memory-utilization 0.95 \
  --enable-prefix-caching \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --enable-reasoning \
  --reasoning-parser qwen3
```

### Inference via Transformers

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "pragmaticcs/PentaCoder-9B"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

messages = [
    {
        "role": "system",
        "content": "You are an expert systems engineer. Reason step by step and output clean, robust implementations.",
    },
    {
        "role": "user",
        "content": "Write an asynchronous connection pool in Python for TCP sockets with active health-checking, backpressure control, and graceful shutdown handling.",
    },
]

inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, enable_thinking=True, return_tensors="pt"
).to(model.device)

outputs = model.generate(
    inputs,
    max_new_tokens=4096,
    temperature=0.6,
    top_p=0.95,
    top_k=20,
    do_sample=True,
)

response = tokenizer.decode(outputs[0][inputs.shape[-1] :], skip_special_tokens=True)
print(response)
```

---

## Citation & References

- [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)
- [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B)
- [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B)
- [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder)
- [OrionLLM/OxCoder-9B](https://huggingface.co/OrionLLM/OxCoder-9B)
- [empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill)
- [Improved Chat Template for Qwen 3.x](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x)

```bibtex
@inproceedings{yadav2023ties,
  title={Resolving Interference When Merging Models},
  author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  volume={36},
  pages={7093--7115},
  year={2023}
}

@article{deep2024della,
  title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling},
  author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya},
  journal={arXiv preprint arXiv:2406.11617},
  year={2024}
}

@article{jang2024modelstock,
  title={Model Stock: All We Need Is just a Few Fine-Tuned Models},
  author={Jang, Dong-Hwan and Yoon, Sang-Doo and Song, Gyeong-Moon},
  journal={arXiv preprint arXiv:2403.19522},
  year={2024}
}
```