--- base_model: - Qwen/Qwen3.5-9B - XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B - ornith-ai/Ornith-1.5-9B - Jackrong/Qwopus3.5-9B-Coder - OrionLLM/OxCoder-9B - empero-ai/Qwen3.8-9B-Distill base_model_relation: merge library_name: transformers tags: - merge - ties - della - model-stock - geodella - qwen - qwen3.5 - causal-lm - deltanet - linear-attention - agentic - reasoning - code - swe-bench license: apache-2.0 language: - en - zh pipeline_tag: text-generation model_type: qwen3_5_text --- ![PentaCoder-9B](https://huggingface.co/pragmaticcs/PentaCoder-9B/resolve/main/assets/banner-dark.png#hf-dark-mode-only) ![PentaCoder-9B](https://huggingface.co/pragmaticcs/PentaCoder-9B/resolve/main/assets/banner-light.png#hf-light-mode-only)
[![License](https://img.shields.io/badge/License-Apache%202.0-6E56CF?style=for-the-badge)](https://opensource.org/licenses/Apache-2.0) [![Library](https://img.shields.io/badge/Library-transformers-FFD21E?style=for-the-badge&logo=huggingface&logoColor=black)](https://github.com/huggingface/transformers) [![Merge Method](https://img.shields.io/badge/Merge%20Method-GeoDELLA--HG-27AE60?style=for-the-badge)](#merge-methodology--mathematical-formulation) [![Architecture](https://img.shields.io/badge/Architecture-Qwen%203.5%209B%20Dense-2D9CDB?style=for-the-badge)](#architectural-specifications) [![Attention](https://img.shields.io/badge/Hybrid-Gated%20DeltaNet%20(3:1)-EB5757?style=for-the-badge)](#architectural-specifications) [![Context](https://img.shields.io/badge/Context%20Window-256k-F2994A?style=for-the-badge)](#architectural-specifications)
Most sub-10B coding models fail in realistic agentic environments due to a common trade-off: aggressive fine-tuning on synthetic coding instructions improves immediate benchmark pass rates but introduces brittle syntactic degradation and catastrophic repetition loops when shell commands or compiler checks fail. **PentaCoder** addresses these limitations by uniting five specialized post-trained checkpoints of [**Qwen 3.5 9B**](https://huggingface.co/Qwen/Qwen3.5-9B) via **GeoDELLA** (Geometric Drop-and-Rescale with Task-Covariance De-Biasing and Spectral Norm Anchoring). The model synthesizes the distinct mathematical distributions of each donor: - **Algorithmic Correctness & Architectural Decomposition** from [**Qwopus3.5-Coder**](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) (Claude 3.5 Opus distillation trajectories). - **Multi-Turn SWE-bench Planning & Tool Protocol Integrity** from [**MiMo-V2.6-Distill**](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) (Large-scale agentic execution traces). - **Terminal Execution Discipline & Error-Recovery Heuristics** from [**Ornith-1.5**](https://huggingface.co/ornith-ai/Ornith-1.5-9B) (Reinforcement learning for anti-looping). - **Low-Level Systems Implementation & Runtime Robustness** from [**OxCoder**](https://huggingface.co/OrionLLM/OxCoder-9B) (Deep API, CLI, and operational coding specialization). - **Abstract Structural Reasoning & Syntax Grounding** from [**Qwen3.8-Distill**](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) (Distilled frontier chain-of-thought representations). > [!IMPORTANT] > The result is a lean, blisteringly fast 9B pure-text causal engine with a native **256k context window** that runs comfortably on consumer GPUs. --- ### Contents - [Architectural Specifications](#architectural-specifications) - [Composition & Donor Weighting](#composition--donor-weighting) - [Merge Methodology & Mathematical Formulation](#merge-methodology--mathematical-formulation) - [Layer-Stratified Component Policies](#layer-stratified-component-policies) - [Agentic Chat Template & Operational Directives](#agentic-chat-template--operational-directives) - [Recommended Generation Parameters](#recommended-generation-parameters) - [How to Use](#how-to-use) - [Citation & References](#citation--references) --- ## Architectural Specifications | Parameter | Specification | | :--- | :---: | | **Total Parameters** | 8.8B (Pure Text Backbone) | | **Architecture Type** | Hybrid Recurrent-Attention Causal LM (`qwen3_5_text`) | | **Hidden Dimension** (*d*model) | 4096 | | **Intermediate Dimension** (*d*mlp) | 12288 (SwiGLU) | | **Decoder Layers** | 32 | | **Attention Layout** | 8 Blocks × (3 Gated DeltaNet Linear Layers : 1 Gated Softmax Layer) | | **Full Attention Layers** | Layers 3, 7, 11, 15, 19, 23, 27, 31 | | **Linear Attention Configuration** | 16 Key Heads / 32 Value Heads (*d*k = *d*v = 128) | | **Full Attention Configuration** | 16 Query Heads / 4 Key-Value Heads (GQA, *d*h = 256) | | **Rotary Position Embedding (RoPE)** | 1D Partial RoPE (θ = 10⁷, Factor = 0.25 → 64 dimensions) | | **Context Window Length** | 262,144 tokens (256k) | | **Native Precision** | `bfloat16` | | **Vocabulary Size** | 248,320 (Padded for Fill-In-The-Middle and Tool Tokens) | --- ## Composition & Donor Weighting The foundational weights of [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) serve as the shared topological base (*W*₀). Five specialized donor checkpoints provide non-overlapping task vectors mapped across normalized layer depth *u* ∈ [0, 1]: | Model | Primary Focus | Depth Target | | :--- | :--- | :---: | | [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) | Base pre-trained manifold and state-space anchors | Global (*W*₀) | | [empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) | Abstract token synthesis, reasoning structure, syntax | Lower & Mid Decoders | | [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) | Typing discipline, algorithm design, functional purity | Mid Decoders (Bell Curve) | | [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B) | Circuit-breaker recovery, environment feedback integration | Upper-Mid Decoders | | [OrionLLM/OxCoder-9B](https://huggingface.co/OrionLLM/OxCoder-9B) | Systems engineering, runtime bug localization, CLI tooling | Deep Layers (Ascending Ramp) | | [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) | SWE-bench multi-step planning, long-context tool invocation | Broad Central Decoders | --- ## Merge Methodology & Mathematical Formulation The merge was executed using the **GeoDELLA-Coder** pipeline, which solves weight interference through dynamic task-correlation de-biasing, row-wise sign consensus, sub-tensor GQA disentanglement, and power-iteration spectral norm stabilization. ### 1. Task Vector Formulation For each donor checkpoint *k* ∈ {1, ..., 5}, the parameter update delta τ*k* is computed relative to the base anchor *W*₀: $$ \tau_k = D_k - W_0, \quad k \in \{\text{Qwen3.8}, \text{Qwopus}, \text{Ornith}, \text{OxCoder}, \text{MiMo}\} $$ ### 2. Hybrid-Aware Continuous Depth Modulation Task vector mixing coefficients are continuously modulated over normalized depth *u* = *l* / (*L* - 1), where *l* ∈ [0, 31] and *L* = 32: $$ u_{\text{qwen38}}(u) = 0.15 + 0.35 \cos^2\left(\frac{\pi}{2} u\right) + 0.15 \sin^2(\pi u) $$ $$ u_{\text{qwopus}}(u) = 0.05 + 0.35 \sin^2(\pi u) $$ $$ u_{\text{ornith}}(u) = 0.05 + 0.35 \sin^2\left(\frac{\pi}{2} u\right) $$ $$ u_{\text{oxcoder}}(u) = 0.05 + 0.25 \sin^2\left(\frac{\pi}{2} u\right) $$ $$ u_{\text{mimo}}(u) = 0.05 + 0.15 \sin(\pi u) $$ The initial coefficients are normalized to form a partition of unity across all layers: $$ w_k(l) = \frac{u_k(u)}{\sum_{j=1}^5 u_j(u)}, \quad \sum_{k=1}^5 w_k(l) = 1.0 $$ ### 3. Dynamic Gram-Matrix Task De-Biasing Fine-tuned models frequently share underlying distillation datasets, leading to collinearity that can drown out specialized task vectors. To correct for this over-representation, the empirical Gram correlation matrix *G* ∈ ℝ*K* × *K* is evaluated per tensor: $$ G_{ij} = \frac{|\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle|}{\|\tau_i\|_2 \|\tau_j\|_2} $$ A uniqueness coefficient *U**k* is derived from the inverse column sum of task vector correlations: $$ U_k = \frac{1}{\sum_{j=1}^K G_{kj}} $$ The dynamic weights are then re-balanced and normalized: $$ \tilde{w}_k = \frac{w_k U_k}{\sum_{j=1}^K w_j U_j} $$ This prevents dataset overlap from suppressing distinct algorithmic representations. ### 4. Asymmetric GQA Disentangled QKV Slicing Qwen 3.5 9B features asymmetric head ratios in both its linear attention and full attention layers. Standard fused tensor merging causes cross-head pollution by treating routing projections identically to memory projections. Fused projection tensors are sliced into their functional sub-matrices prior to merging: - **Gated DeltaNet Layers (8192 × *d*model):** Sliced into Query (2048), Key (2048), and Value (4096). - **Gated Attention Layers (6144 × *d*model):** Sliced into Query (4096), Key (1024), and Value (1024). Query and Key slices are processed with a conservative retention density (ρ = 0.95) to preserve sharp context routing. Value matrices are processed with adaptive MLP density (ρ = 0.70) to maximize conceptual synthesis. The components are then re-concatenated along the head dimension. ### 5. Neuron-Coherent Row Gating To eliminate destructive interference in Feed-Forward Networks (MLPs), task vectors are gated at the single-neuron (row) level: $$ \bar{\tau} = \frac{1}{K} \sum_{k=1}^K \tau_k $$ For each row *r* of donor delta τ*k*, the directional alignment with the consensus mean is evaluated: $$ \cos \theta_{k, r} = \frac{\langle \tau_{k, r}, \bar{\tau}_r \rangle}{\|\tau_{k, r}\|_2 \|\bar{\tau}_r\|_2 + \epsilon} $$ Rows exhibiting severe directional opposition (cos θ*k*, *r* < -0.10) are masked out: $$ \hat{\tau}_{k, r} = \tau_{k, r} \cdot \mathbb{I}\left(\cos \theta_{k, r} \ge -0.10\right) $$ This eliminates opposing gradient vectors that produce incoherent syntax generation. ### 6. Heavy-Tailed DELLA Adaptive Rescaling Surviving parameters undergo non-linear magnitude-based sampling. Using parameter rank indices *R**k*, *ij* ∈ [0, 1] sorted by absolute magnitude, a Pareto-style retention probability *p**k*, *ij* is established: $$ p_{k, ij} = p_{\min} + (p_{\max} - p_{\min}) \cdot \sqrt{R_{k, ij}} $$ Parameters are sampled via a Bernoulli trial and rescaled by their inverse survival probability: $$ M_{k, ij} \sim \text{Bernoulli}(p_{k, ij}) $$ $$ \tilde{\tau}_{k, ij} = \frac{\hat{\tau}_{k, ij} \odot M_{k, ij}}{p_{k, ij}} $$ ### 7. Coordinate-Wise Sign Consensus & Model Stock Scaling Directional consensus is determined via weighted sign agreement: $$ \Gamma = \operatorname{sgn}\left(\sum_{k=1}^K \tilde{w}_k \tilde{\tau}_k\right) $$ $$ A_k = \mathbb{I}\left(\operatorname{sgn}(\tilde{\tau}_k) = \Gamma\right) \odot \mathbb{I}\left(\tilde{\tau}_k \neq 0\right) $$ $$ \Delta_{\text{consensus}} = \frac{\sum_{k=1}^K \tilde{\tau}_k \odot A_k}{\sum_{k=1}^K A_k + \epsilon} $$ The aggregated delta is projected onto the non-linear manifold using the Model Stock analytic scaling factor *t**: $$ \bar{\rho} = \frac{2}{K(K-1)} \sum_{i < j} \frac{\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle}{\|\tau_i\|_2 \|\tau_j\|_2} $$ $$ t^* = \frac{K \bar{\rho}}{1 + (K - 1)\bar{\rho}} $$ $$ \Delta_{\text{final}} = t^* \cdot \Delta_{\text{consensus}} $$ ### 8. Spectral Norm & Attention Entropy Anchoring Attention projection matrices (*W*attn) are vulnerable to spectral explosion during merges, which contracts attention entropy and leads to repetitive generation loops. The dominant singular value σ(*W*) is calculated via a three-iteration deterministic power iteration: $$ v^{(t+1)} = \frac{W^T u^{(t)}}{\|W^T u^{(t)}\|_2}, \quad u^{(t+1)} = \frac{W v^{(t+1)}}{\|W v^{(t+1)}\|_2} $$ $$ \sigma(W) \approx {u^{(3)}}^T W v^{(3)} $$ If the merged spectral radius grows more than 5% relative to the base model, it is scaled down: $$ W_{\text{final}} = \begin{cases} W_{\text{merged}} \cdot \left(\frac{1.05 \cdot \sigma(W_0)}{\sigma(W_{\text{merged}})}\right) & \text{if } \sigma(W_{\text{merged}}) > 1.05 \cdot \sigma(W_0) \\ W_{\text{merged}} & \text{otherwise} \end{cases} $$ --- ## Layer-Stratified Component Policies | Parameter Class | Target Identifiers | Applied Policy | Density (ρ) | Mathematical Constraints | | :--- | :--- | :---: | :---: | :--- | | **Embeddings & LM Head** | `embed_tokens`, `lm_head` | Low-Memory Streaming Blend | 1.0 | Convex iterative accumulation; vocab dimension aligned to 248,320. | | **Linear State-Space Projections** | `linear_attn.in_proj_qkv` | Asymmetric GQA DELLA | 0.95 (QK) / 0.70 (V) | Sub-tensor slicing; separate routing and associative memory passes. | | **Self-Attention Projections** | `self_attn.qkv_proj`, `o_proj` | Asymmetric GQA + Spectral Anchor | 0.95 (QK) / 0.70 (V) | Power-iteration clipping prevents σ > 1.05 σ₀. | | **Feed-Forward Blocks** | `mlp.gate_proj`, `up_proj`, `down_proj` | Neuron-Gated HG-DELLA | 0.50 – 0.70 | Row-wise cosine filtering (cos θ ≥ -0.10); sign consensus. | | **Recurrent Gates & Normalization** | `A_log`, `norm`, `conv1d` | Convex Parameter Blend | 1.0 | Preservation of *A*log ≤ 0 to guarantee BIBO stability. | --- ## Agentic Chat Template & Operational Directives This model uses the [Improved Chat Template for Qwen 3.x by Olivia Rossi](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x) to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination. --- ## Recommended Generation Parameters For deterministic software engineering and complex reasoning benchmarks: | Parameter | Recommended Value | Description | | :--- | :---: | :--- | | **Temperature** | `0.6` | Balances strict syntactic validity with algorithmic path exploration. | | **Top-P** | `0.95` | Eliminates low-probability token tails while preserving alternative logic paths. | | **Top-K** | `20` | Restricts token candidate pools to prevent architectural syntax drift. | | **Min-P** | `0.0` (Off) | Disabled in favor of explicit Top-K / Top-P governance. | | **Repetition Penalty** | `1.0` (Off) | Disabled to prevent syntax degradation in repetitive code patterns (indentation, braces). | | **Presence Penalty** | `0.0` | Prevents naming mutations across long-context symbol resolution. | --- ## How to Use ### Serving via vLLM ```bash vllm serve pragmaticcs/PentaCoder-9B \ --dtype bfloat16 \ --max-model-len 65536 \ --gpu-memory-utilization 0.95 \ --enable-prefix-caching \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coder \ --enable-reasoning \ --reasoning-parser qwen3 ``` ### Inference via Transformers ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "pragmaticcs/PentaCoder-9B" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto" ) messages = [ { "role": "system", "content": "You are an expert systems engineer. Reason step by step and output clean, robust implementations.", }, { "role": "user", "content": "Write an asynchronous connection pool in Python for TCP sockets with active health-checking, backpressure control, and graceful shutdown handling.", }, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, enable_thinking=True, return_tensors="pt" ).to(model.device) outputs = model.generate( inputs, max_new_tokens=4096, temperature=0.6, top_p=0.95, top_k=20, do_sample=True, ) response = tokenizer.decode(outputs[0][inputs.shape[-1] :], skip_special_tokens=True) print(response) ``` --- ## Citation & References - [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) - [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) - [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B) - [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) - [OrionLLM/OxCoder-9B](https://huggingface.co/OrionLLM/OxCoder-9B) - [empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) - [Improved Chat Template for Qwen 3.x](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x) ```bibtex @inproceedings{yadav2023ties, title={Resolving Interference When Merging Models}, author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit}, booktitle={Advances in Neural Information Processing Systems (NeurIPS)}, volume={36}, pages={7093--7115}, year={2023} } @article{deep2024della, title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling}, author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya}, journal={arXiv preprint arXiv:2406.11617}, year={2024} } @article{jang2024modelstock, title={Model Stock: All We Need Is just a Few Fine-Tuned Models}, author={Jang, Dong-Hwan and Yoon, Sang-Doo and Song, Gyeong-Moon}, journal={arXiv preprint arXiv:2403.19522}, year={2024} } ```