---
base_model:
- Qwen/Qwen3.5-9B
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
- ornith-ai/Ornith-1.5-9B
- Jackrong/Qwopus3.5-9B-Coder
- OrionLLM/OxCoder-9B
- empero-ai/Qwen3.8-9B-Distill
base_model_relation: merge
library_name: transformers
tags:
- merge
- ties
- della
- model-stock
- geodella
- qwen
- qwen3.5
- causal-lm
- deltanet
- linear-attention
- agentic
- reasoning
- code
- swe-bench
license: apache-2.0
language:
- en
- zh
pipeline_tag: text-generation
model_type: qwen3_5_text
---


[](https://opensource.org/licenses/Apache-2.0)
[](https://github.com/huggingface/transformers)
[](#merge-methodology--mathematical-formulation)
[](#architectural-specifications)
[-EB5757?style=for-the-badge)](#architectural-specifications)
[](#architectural-specifications)
Most sub-10B coding models fail in realistic agentic environments due to a common trade-off: aggressive fine-tuning on synthetic coding instructions improves immediate benchmark pass rates but introduces brittle syntactic degradation and catastrophic repetition loops when shell commands or compiler checks fail.
**PentaCoder** addresses these limitations by uniting five specialized post-trained checkpoints of [**Qwen 3.5 9B**](https://huggingface.co/Qwen/Qwen3.5-9B) via **GeoDELLA** (Geometric Drop-and-Rescale with Task-Covariance De-Biasing and Spectral Norm Anchoring). The model synthesizes the distinct mathematical distributions of each donor:
- **Algorithmic Correctness & Architectural Decomposition** from [**Qwopus3.5-Coder**](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) (Claude 3.5 Opus distillation trajectories).
- **Multi-Turn SWE-bench Planning & Tool Protocol Integrity** from [**MiMo-V2.6-Distill**](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) (Large-scale agentic execution traces).
- **Terminal Execution Discipline & Error-Recovery Heuristics** from [**Ornith-1.5**](https://huggingface.co/ornith-ai/Ornith-1.5-9B) (Reinforcement learning for anti-looping).
- **Low-Level Systems Implementation & Runtime Robustness** from [**OxCoder**](https://huggingface.co/OrionLLM/OxCoder-9B) (Deep API, CLI, and operational coding specialization).
- **Abstract Structural Reasoning & Syntax Grounding** from [**Qwen3.8-Distill**](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) (Distilled frontier chain-of-thought representations).
> [!IMPORTANT]
> The result is a lean, blisteringly fast 9B pure-text causal engine with a native **256k context window** that runs comfortably on consumer GPUs.
---
### Contents
- [Architectural Specifications](#architectural-specifications)
- [Composition & Donor Weighting](#composition--donor-weighting)
- [Merge Methodology & Mathematical Formulation](#merge-methodology--mathematical-formulation)
- [Layer-Stratified Component Policies](#layer-stratified-component-policies)
- [Agentic Chat Template & Operational Directives](#agentic-chat-template--operational-directives)
- [Recommended Generation Parameters](#recommended-generation-parameters)
- [How to Use](#how-to-use)
- [Citation & References](#citation--references)
---
## Architectural Specifications
| Parameter | Specification |
| :--- | :---: |
| **Total Parameters** | 8.8B (Pure Text Backbone) |
| **Architecture Type** | Hybrid Recurrent-Attention Causal LM (`qwen3_5_text`) |
| **Hidden Dimension** (*d*model) | 4096 |
| **Intermediate Dimension** (*d*mlp) | 12288 (SwiGLU) |
| **Decoder Layers** | 32 |
| **Attention Layout** | 8 Blocks × (3 Gated DeltaNet Linear Layers : 1 Gated Softmax Layer) |
| **Full Attention Layers** | Layers 3, 7, 11, 15, 19, 23, 27, 31 |
| **Linear Attention Configuration** | 16 Key Heads / 32 Value Heads (*d*k = *d*v = 128) |
| **Full Attention Configuration** | 16 Query Heads / 4 Key-Value Heads (GQA, *d*h = 256) |
| **Rotary Position Embedding (RoPE)** | 1D Partial RoPE (θ = 10⁷, Factor = 0.25 → 64 dimensions) |
| **Context Window Length** | 262,144 tokens (256k) |
| **Native Precision** | `bfloat16` |
| **Vocabulary Size** | 248,320 (Padded for Fill-In-The-Middle and Tool Tokens) |
---
## Composition & Donor Weighting
The foundational weights of [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) serve as the shared topological base (*W*₀). Five specialized donor checkpoints provide non-overlapping task vectors mapped across normalized layer depth *u* ∈ [0, 1]:
| Model | Primary Focus | Depth Target |
| :--- | :--- | :---: |
| [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) | Base pre-trained manifold and state-space anchors | Global (*W*₀) |
| [empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) | Abstract token synthesis, reasoning structure, syntax | Lower & Mid Decoders |
| [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) | Typing discipline, algorithm design, functional purity | Mid Decoders (Bell Curve) |
| [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B) | Circuit-breaker recovery, environment feedback integration | Upper-Mid Decoders |
| [OrionLLM/OxCoder-9B](https://huggingface.co/OrionLLM/OxCoder-9B) | Systems engineering, runtime bug localization, CLI tooling | Deep Layers (Ascending Ramp) |
| [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) | SWE-bench multi-step planning, long-context tool invocation | Broad Central Decoders |
---
## Merge Methodology & Mathematical Formulation
The merge was executed using the **GeoDELLA-Coder** pipeline, which solves weight interference through dynamic task-correlation de-biasing, row-wise sign consensus, sub-tensor GQA disentanglement, and power-iteration spectral norm stabilization.
### 1. Task Vector Formulation
For each donor checkpoint *k* ∈ {1, ..., 5}, the parameter update delta τ*k* is computed relative to the base anchor *W*₀:
$$
\tau_k = D_k - W_0, \quad k \in \{\text{Qwen3.8}, \text{Qwopus}, \text{Ornith}, \text{OxCoder}, \text{MiMo}\}
$$
### 2. Hybrid-Aware Continuous Depth Modulation
Task vector mixing coefficients are continuously modulated over normalized depth *u* = *l* / (*L* - 1), where *l* ∈ [0, 31] and *L* = 32:
$$
u_{\text{qwen38}}(u) = 0.15 + 0.35 \cos^2\left(\frac{\pi}{2} u\right) + 0.15 \sin^2(\pi u)
$$
$$
u_{\text{qwopus}}(u) = 0.05 + 0.35 \sin^2(\pi u)
$$
$$
u_{\text{ornith}}(u) = 0.05 + 0.35 \sin^2\left(\frac{\pi}{2} u\right)
$$
$$
u_{\text{oxcoder}}(u) = 0.05 + 0.25 \sin^2\left(\frac{\pi}{2} u\right)
$$
$$
u_{\text{mimo}}(u) = 0.05 + 0.15 \sin(\pi u)
$$
The initial coefficients are normalized to form a partition of unity across all layers:
$$
w_k(l) = \frac{u_k(u)}{\sum_{j=1}^5 u_j(u)}, \quad \sum_{k=1}^5 w_k(l) = 1.0
$$
### 3. Dynamic Gram-Matrix Task De-Biasing
Fine-tuned models frequently share underlying distillation datasets, leading to collinearity that can drown out specialized task vectors. To correct for this over-representation, the empirical Gram correlation matrix *G* ∈ ℝ*K* × *K* is evaluated per tensor:
$$
G_{ij} = \frac{|\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle|}{\|\tau_i\|_2 \|\tau_j\|_2}
$$
A uniqueness coefficient *U**k* is derived from the inverse column sum of task vector correlations:
$$
U_k = \frac{1}{\sum_{j=1}^K G_{kj}}
$$
The dynamic weights are then re-balanced and normalized:
$$
\tilde{w}_k = \frac{w_k U_k}{\sum_{j=1}^K w_j U_j}
$$
This prevents dataset overlap from suppressing distinct algorithmic representations.
### 4. Asymmetric GQA Disentangled QKV Slicing
Qwen 3.5 9B features asymmetric head ratios in both its linear attention and full attention layers. Standard fused tensor merging causes cross-head pollution by treating routing projections identically to memory projections.
Fused projection tensors are sliced into their functional sub-matrices prior to merging:
- **Gated DeltaNet Layers (8192 × *d*model):** Sliced into Query (2048), Key (2048), and Value (4096).
- **Gated Attention Layers (6144 × *d*model):** Sliced into Query (4096), Key (1024), and Value (1024).
Query and Key slices are processed with a conservative retention density (ρ = 0.95) to preserve sharp context routing. Value matrices are processed with adaptive MLP density (ρ = 0.70) to maximize conceptual synthesis. The components are then re-concatenated along the head dimension.
### 5. Neuron-Coherent Row Gating
To eliminate destructive interference in Feed-Forward Networks (MLPs), task vectors are gated at the single-neuron (row) level:
$$
\bar{\tau} = \frac{1}{K} \sum_{k=1}^K \tau_k
$$
For each row *r* of donor delta τ*k*, the directional alignment with the consensus mean is evaluated:
$$
\cos \theta_{k, r} = \frac{\langle \tau_{k, r}, \bar{\tau}_r \rangle}{\|\tau_{k, r}\|_2 \|\bar{\tau}_r\|_2 + \epsilon}
$$
Rows exhibiting severe directional opposition (cos θ*k*, *r* < -0.10) are masked out:
$$
\hat{\tau}_{k, r} = \tau_{k, r} \cdot \mathbb{I}\left(\cos \theta_{k, r} \ge -0.10\right)
$$
This eliminates opposing gradient vectors that produce incoherent syntax generation.
### 6. Heavy-Tailed DELLA Adaptive Rescaling
Surviving parameters undergo non-linear magnitude-based sampling. Using parameter rank indices *R**k*, *ij* ∈ [0, 1] sorted by absolute magnitude, a Pareto-style retention probability *p**k*, *ij* is established:
$$
p_{k, ij} = p_{\min} + (p_{\max} - p_{\min}) \cdot \sqrt{R_{k, ij}}
$$
Parameters are sampled via a Bernoulli trial and rescaled by their inverse survival probability:
$$
M_{k, ij} \sim \text{Bernoulli}(p_{k, ij})
$$
$$
\tilde{\tau}_{k, ij} = \frac{\hat{\tau}_{k, ij} \odot M_{k, ij}}{p_{k, ij}}
$$
### 7. Coordinate-Wise Sign Consensus & Model Stock Scaling
Directional consensus is determined via weighted sign agreement:
$$
\Gamma = \operatorname{sgn}\left(\sum_{k=1}^K \tilde{w}_k \tilde{\tau}_k\right)
$$
$$
A_k = \mathbb{I}\left(\operatorname{sgn}(\tilde{\tau}_k) = \Gamma\right) \odot \mathbb{I}\left(\tilde{\tau}_k \neq 0\right)
$$
$$
\Delta_{\text{consensus}} = \frac{\sum_{k=1}^K \tilde{\tau}_k \odot A_k}{\sum_{k=1}^K A_k + \epsilon}
$$
The aggregated delta is projected onto the non-linear manifold using the Model Stock analytic scaling factor *t**:
$$
\bar{\rho} = \frac{2}{K(K-1)} \sum_{i < j} \frac{\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle}{\|\tau_i\|_2 \|\tau_j\|_2}
$$
$$
t^* = \frac{K \bar{\rho}}{1 + (K - 1)\bar{\rho}}
$$
$$
\Delta_{\text{final}} = t^* \cdot \Delta_{\text{consensus}}
$$
### 8. Spectral Norm & Attention Entropy Anchoring
Attention projection matrices (*W*attn) are vulnerable to spectral explosion during merges, which contracts attention entropy and leads to repetitive generation loops.
The dominant singular value σ(*W*) is calculated via a three-iteration deterministic power iteration:
$$
v^{(t+1)} = \frac{W^T u^{(t)}}{\|W^T u^{(t)}\|_2}, \quad u^{(t+1)} = \frac{W v^{(t+1)}}{\|W v^{(t+1)}\|_2}
$$
$$
\sigma(W) \approx {u^{(3)}}^T W v^{(3)}
$$
If the merged spectral radius grows more than 5% relative to the base model, it is scaled down:
$$
W_{\text{final}} = \begin{cases} W_{\text{merged}} \cdot \left(\frac{1.05 \cdot \sigma(W_0)}{\sigma(W_{\text{merged}})}\right) & \text{if } \sigma(W_{\text{merged}}) > 1.05 \cdot \sigma(W_0) \\ W_{\text{merged}} & \text{otherwise} \end{cases}
$$
---
## Layer-Stratified Component Policies
| Parameter Class | Target Identifiers | Applied Policy | Density (ρ) | Mathematical Constraints |
| :--- | :--- | :---: | :---: | :--- |
| **Embeddings & LM Head** | `embed_tokens`, `lm_head` | Low-Memory Streaming Blend | 1.0 | Convex iterative accumulation; vocab dimension aligned to 248,320. |
| **Linear State-Space Projections** | `linear_attn.in_proj_qkv` | Asymmetric GQA DELLA | 0.95 (QK) / 0.70 (V) | Sub-tensor slicing; separate routing and associative memory passes. |
| **Self-Attention Projections** | `self_attn.qkv_proj`, `o_proj` | Asymmetric GQA + Spectral Anchor | 0.95 (QK) / 0.70 (V) | Power-iteration clipping prevents σ > 1.05 σ₀. |
| **Feed-Forward Blocks** | `mlp.gate_proj`, `up_proj`, `down_proj` | Neuron-Gated HG-DELLA | 0.50 – 0.70 | Row-wise cosine filtering (cos θ ≥ -0.10); sign consensus. |
| **Recurrent Gates & Normalization** | `A_log`, `norm`, `conv1d` | Convex Parameter Blend | 1.0 | Preservation of *A*log ≤ 0 to guarantee BIBO stability. |
---
## Agentic Chat Template & Operational Directives
This model uses the [Improved Chat Template for Qwen 3.x by Olivia Rossi](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x) to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination.
---
## Recommended Generation Parameters
For deterministic software engineering and complex reasoning benchmarks:
| Parameter | Recommended Value | Description |
| :--- | :---: | :--- |
| **Temperature** | `0.6` | Balances strict syntactic validity with algorithmic path exploration. |
| **Top-P** | `0.95` | Eliminates low-probability token tails while preserving alternative logic paths. |
| **Top-K** | `20` | Restricts token candidate pools to prevent architectural syntax drift. |
| **Min-P** | `0.0` (Off) | Disabled in favor of explicit Top-K / Top-P governance. |
| **Repetition Penalty** | `1.0` (Off) | Disabled to prevent syntax degradation in repetitive code patterns (indentation, braces). |
| **Presence Penalty** | `0.0` | Prevents naming mutations across long-context symbol resolution. |
---
## How to Use
### Serving via vLLM
```bash
vllm serve pragmaticcs/PentaCoder-9B \
--dtype bfloat16 \
--max-model-len 65536 \
--gpu-memory-utilization 0.95 \
--enable-prefix-caching \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--enable-reasoning \
--reasoning-parser qwen3
```
### Inference via Transformers
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "pragmaticcs/PentaCoder-9B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto"
)
messages = [
{
"role": "system",
"content": "You are an expert systems engineer. Reason step by step and output clean, robust implementations.",
},
{
"role": "user",
"content": "Write an asynchronous connection pool in Python for TCP sockets with active health-checking, backpressure control, and graceful shutdown handling.",
},
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, enable_thinking=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=4096,
temperature=0.6,
top_p=0.95,
top_k=20,
do_sample=True,
)
response = tokenizer.decode(outputs[0][inputs.shape[-1] :], skip_special_tokens=True)
print(response)
```
---
## Citation & References
- [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)
- [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B)
- [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B)
- [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder)
- [OrionLLM/OxCoder-9B](https://huggingface.co/OrionLLM/OxCoder-9B)
- [empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill)
- [Improved Chat Template for Qwen 3.x](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x)
```bibtex
@inproceedings{yadav2023ties,
title={Resolving Interference When Merging Models},
author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
volume={36},
pages={7093--7115},
year={2023}
}
@article{deep2024della,
title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling},
author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya},
journal={arXiv preprint arXiv:2406.11617},
year={2024}
}
@article{jang2024modelstock,
title={Model Stock: All We Need Is just a Few Fine-Tuned Models},
author={Jang, Dong-Hwan and Yoon, Sang-Doo and Song, Gyeong-Moon},
journal={arXiv preprint arXiv:2403.19522},
year={2024}
}
```