Text Generation
Transformers
Safetensors
English
Chinese
qwen3_5_text
Merge
ties
della
model-stock
geodella
qwen
qwen3.5
causal-lm
deltanet
linear-attention
agentic
reasoning
code
swe-bench
conversational
Instructions to use pragmaticcs/PentaCoder-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pragmaticcs/PentaCoder-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="pragmaticcs/PentaCoder-9B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("pragmaticcs/PentaCoder-9B") model = AutoModelForCausalLM.from_pretrained("pragmaticcs/PentaCoder-9B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use pragmaticcs/PentaCoder-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pragmaticcs/PentaCoder-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/PentaCoder-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/pragmaticcs/PentaCoder-9B
- SGLang
How to use pragmaticcs/PentaCoder-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "pragmaticcs/PentaCoder-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/PentaCoder-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "pragmaticcs/PentaCoder-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/PentaCoder-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use pragmaticcs/PentaCoder-9B with Docker Model Runner:
docker model run hf.co/pragmaticcs/PentaCoder-9B
File size: 17,412 Bytes
d97b72b d3ca140 149fb12 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b b0657a2 d97b72b 1978c51 d97b72b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 | ---
base_model:
- Qwen/Qwen3.5-9B
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
- ornith-ai/Ornith-1.5-9B
- Jackrong/Qwopus3.5-9B-Coder
- OrionLLM/OxCoder-9B
- empero-ai/Qwen3.8-9B-Distill
base_model_relation: merge
library_name: transformers
tags:
- merge
- ties
- della
- model-stock
- geodella
- qwen
- qwen3.5
- causal-lm
- deltanet
- linear-attention
- agentic
- reasoning
- code
- swe-bench
license: apache-2.0
language:
- en
- zh
pipeline_tag: text-generation
model_type: qwen3_5_text
---


<div align="center">
[](https://opensource.org/licenses/Apache-2.0)
[](https://github.com/huggingface/transformers)
[](#merge-methodology--mathematical-formulation)
[](#architectural-specifications)
[-EB5757?style=for-the-badge)](#architectural-specifications)
[](#architectural-specifications)
</div>
Most sub-10B coding models fail in realistic agentic environments due to a common trade-off: aggressive fine-tuning on synthetic coding instructions improves immediate benchmark pass rates but introduces brittle syntactic degradation and catastrophic repetition loops when shell commands or compiler checks fail.
**PentaCoder** addresses these limitations by uniting five specialized post-trained checkpoints of [**Qwen 3.5 9B**](https://huggingface.co/Qwen/Qwen3.5-9B) via **GeoDELLA** (Geometric Drop-and-Rescale with Task-Covariance De-Biasing and Spectral Norm Anchoring). The model synthesizes the distinct mathematical distributions of each donor:
- **Algorithmic Correctness & Architectural Decomposition** from [**Qwopus3.5-Coder**](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) (Claude 3.5 Opus distillation trajectories).
- **Multi-Turn SWE-bench Planning & Tool Protocol Integrity** from [**MiMo-V2.6-Distill**](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) (Large-scale agentic execution traces).
- **Terminal Execution Discipline & Error-Recovery Heuristics** from [**Ornith-1.5**](https://huggingface.co/ornith-ai/Ornith-1.5-9B) (Reinforcement learning for anti-looping).
- **Low-Level Systems Implementation & Runtime Robustness** from [**OxCoder**](https://huggingface.co/OrionLLM/OxCoder-9B) (Deep API, CLI, and operational coding specialization).
- **Abstract Structural Reasoning & Syntax Grounding** from [**Qwen3.8-Distill**](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) (Distilled frontier chain-of-thought representations).
> [!IMPORTANT]
> The result is a lean, blisteringly fast 9B pure-text causal engine with a native **256k context window** that runs comfortably on consumer GPUs.
---
### Contents
- [Architectural Specifications](#architectural-specifications)
- [Composition & Donor Weighting](#composition--donor-weighting)
- [Merge Methodology & Mathematical Formulation](#merge-methodology--mathematical-formulation)
- [Layer-Stratified Component Policies](#layer-stratified-component-policies)
- [Agentic Chat Template & Operational Directives](#agentic-chat-template--operational-directives)
- [Recommended Generation Parameters](#recommended-generation-parameters)
- [How to Use](#how-to-use)
- [Citation & References](#citation--references)
---
## Architectural Specifications
| Parameter | Specification |
| :--- | :---: |
| **Total Parameters** | 8.8B (Pure Text Backbone) |
| **Architecture Type** | Hybrid Recurrent-Attention Causal LM (`qwen3_5_text`) |
| **Hidden Dimension** (*d*<sub>model</sub>) | 4096 |
| **Intermediate Dimension** (*d*<sub>mlp</sub>) | 12288 (SwiGLU) |
| **Decoder Layers** | 32 |
| **Attention Layout** | 8 Blocks × (3 Gated DeltaNet Linear Layers : 1 Gated Softmax Layer) |
| **Full Attention Layers** | Layers 3, 7, 11, 15, 19, 23, 27, 31 |
| **Linear Attention Configuration** | 16 Key Heads / 32 Value Heads (*d*<sub>k</sub> = *d*<sub>v</sub> = 128) |
| **Full Attention Configuration** | 16 Query Heads / 4 Key-Value Heads (GQA, *d*<sub>h</sub> = 256) |
| **Rotary Position Embedding (RoPE)** | 1D Partial RoPE (θ = 10⁷, Factor = 0.25 → 64 dimensions) |
| **Context Window Length** | 262,144 tokens (256k) |
| **Native Precision** | `bfloat16` |
| **Vocabulary Size** | 248,320 (Padded for Fill-In-The-Middle and Tool Tokens) |
---
## Composition & Donor Weighting
The foundational weights of [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) serve as the shared topological base (*W*₀). Five specialized donor checkpoints provide non-overlapping task vectors mapped across normalized layer depth *u* ∈ [0, 1]:
| Model | Primary Focus | Depth Target |
| :--- | :--- | :---: |
| [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) | Base pre-trained manifold and state-space anchors | Global (*W*₀) |
| [empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) | Abstract token synthesis, reasoning structure, syntax | Lower & Mid Decoders |
| [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) | Typing discipline, algorithm design, functional purity | Mid Decoders (Bell Curve) |
| [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B) | Circuit-breaker recovery, environment feedback integration | Upper-Mid Decoders |
| [OrionLLM/OxCoder-9B](https://huggingface.co/OrionLLM/OxCoder-9B) | Systems engineering, runtime bug localization, CLI tooling | Deep Layers (Ascending Ramp) |
| [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) | SWE-bench multi-step planning, long-context tool invocation | Broad Central Decoders |
---
## Merge Methodology & Mathematical Formulation
The merge was executed using the **GeoDELLA-Coder** pipeline, which solves weight interference through dynamic task-correlation de-biasing, row-wise sign consensus, sub-tensor GQA disentanglement, and power-iteration spectral norm stabilization.
### 1. Task Vector Formulation
For each donor checkpoint *k* ∈ {1, ..., 5}, the parameter update delta τ<sub>*k*</sub> is computed relative to the base anchor *W*₀:
$$
\tau_k = D_k - W_0, \quad k \in \{\text{Qwen3.8}, \text{Qwopus}, \text{Ornith}, \text{OxCoder}, \text{MiMo}\}
$$
### 2. Hybrid-Aware Continuous Depth Modulation
Task vector mixing coefficients are continuously modulated over normalized depth *u* = *l* / (*L* - 1), where *l* ∈ [0, 31] and *L* = 32:
$$
u_{\text{qwen38}}(u) = 0.15 + 0.35 \cos^2\left(\frac{\pi}{2} u\right) + 0.15 \sin^2(\pi u)
$$
$$
u_{\text{qwopus}}(u) = 0.05 + 0.35 \sin^2(\pi u)
$$
$$
u_{\text{ornith}}(u) = 0.05 + 0.35 \sin^2\left(\frac{\pi}{2} u\right)
$$
$$
u_{\text{oxcoder}}(u) = 0.05 + 0.25 \sin^2\left(\frac{\pi}{2} u\right)
$$
$$
u_{\text{mimo}}(u) = 0.05 + 0.15 \sin(\pi u)
$$
The initial coefficients are normalized to form a partition of unity across all layers:
$$
w_k(l) = \frac{u_k(u)}{\sum_{j=1}^5 u_j(u)}, \quad \sum_{k=1}^5 w_k(l) = 1.0
$$
### 3. Dynamic Gram-Matrix Task De-Biasing
Fine-tuned models frequently share underlying distillation datasets, leading to collinearity that can drown out specialized task vectors. To correct for this over-representation, the empirical Gram correlation matrix *G* ∈ ℝ<sup>*K* × *K*</sup> is evaluated per tensor:
$$
G_{ij} = \frac{|\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle|}{\|\tau_i\|_2 \|\tau_j\|_2}
$$
A uniqueness coefficient *U*<sub>*k*</sub> is derived from the inverse column sum of task vector correlations:
$$
U_k = \frac{1}{\sum_{j=1}^K G_{kj}}
$$
The dynamic weights are then re-balanced and normalized:
$$
\tilde{w}_k = \frac{w_k U_k}{\sum_{j=1}^K w_j U_j}
$$
This prevents dataset overlap from suppressing distinct algorithmic representations.
### 4. Asymmetric GQA Disentangled QKV Slicing
Qwen 3.5 9B features asymmetric head ratios in both its linear attention and full attention layers. Standard fused tensor merging causes cross-head pollution by treating routing projections identically to memory projections.
Fused projection tensors are sliced into their functional sub-matrices prior to merging:
- **Gated DeltaNet Layers (8192 × *d*<sub>model</sub>):** Sliced into Query (2048), Key (2048), and Value (4096).
- **Gated Attention Layers (6144 × *d*<sub>model</sub>):** Sliced into Query (4096), Key (1024), and Value (1024).
Query and Key slices are processed with a conservative retention density (ρ = 0.95) to preserve sharp context routing. Value matrices are processed with adaptive MLP density (ρ = 0.70) to maximize conceptual synthesis. The components are then re-concatenated along the head dimension.
### 5. Neuron-Coherent Row Gating
To eliminate destructive interference in Feed-Forward Networks (MLPs), task vectors are gated at the single-neuron (row) level:
$$
\bar{\tau} = \frac{1}{K} \sum_{k=1}^K \tau_k
$$
For each row *r* of donor delta τ<sub>*k*</sub>, the directional alignment with the consensus mean is evaluated:
$$
\cos \theta_{k, r} = \frac{\langle \tau_{k, r}, \bar{\tau}_r \rangle}{\|\tau_{k, r}\|_2 \|\bar{\tau}_r\|_2 + \epsilon}
$$
Rows exhibiting severe directional opposition (cos θ<sub>*k*, *r*</sub> < -0.10) are masked out:
$$
\hat{\tau}_{k, r} = \tau_{k, r} \cdot \mathbb{I}\left(\cos \theta_{k, r} \ge -0.10\right)
$$
This eliminates opposing gradient vectors that produce incoherent syntax generation.
### 6. Heavy-Tailed DELLA Adaptive Rescaling
Surviving parameters undergo non-linear magnitude-based sampling. Using parameter rank indices *R*<sub>*k*, *ij*</sub> ∈ [0, 1] sorted by absolute magnitude, a Pareto-style retention probability *p*<sub>*k*, *ij*</sub> is established:
$$
p_{k, ij} = p_{\min} + (p_{\max} - p_{\min}) \cdot \sqrt{R_{k, ij}}
$$
Parameters are sampled via a Bernoulli trial and rescaled by their inverse survival probability:
$$
M_{k, ij} \sim \text{Bernoulli}(p_{k, ij})
$$
$$
\tilde{\tau}_{k, ij} = \frac{\hat{\tau}_{k, ij} \odot M_{k, ij}}{p_{k, ij}}
$$
### 7. Coordinate-Wise Sign Consensus & Model Stock Scaling
Directional consensus is determined via weighted sign agreement:
$$
\Gamma = \operatorname{sgn}\left(\sum_{k=1}^K \tilde{w}_k \tilde{\tau}_k\right)
$$
$$
A_k = \mathbb{I}\left(\operatorname{sgn}(\tilde{\tau}_k) = \Gamma\right) \odot \mathbb{I}\left(\tilde{\tau}_k \neq 0\right)
$$
$$
\Delta_{\text{consensus}} = \frac{\sum_{k=1}^K \tilde{\tau}_k \odot A_k}{\sum_{k=1}^K A_k + \epsilon}
$$
The aggregated delta is projected onto the non-linear manifold using the Model Stock analytic scaling factor *t*<sup>*</sup>:
$$
\bar{\rho} = \frac{2}{K(K-1)} \sum_{i < j} \frac{\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle}{\|\tau_i\|_2 \|\tau_j\|_2}
$$
$$
t^* = \frac{K \bar{\rho}}{1 + (K - 1)\bar{\rho}}
$$
$$
\Delta_{\text{final}} = t^* \cdot \Delta_{\text{consensus}}
$$
### 8. Spectral Norm & Attention Entropy Anchoring
Attention projection matrices (*W*<sub>attn</sub>) are vulnerable to spectral explosion during merges, which contracts attention entropy and leads to repetitive generation loops.
The dominant singular value σ(*W*) is calculated via a three-iteration deterministic power iteration:
$$
v^{(t+1)} = \frac{W^T u^{(t)}}{\|W^T u^{(t)}\|_2}, \quad u^{(t+1)} = \frac{W v^{(t+1)}}{\|W v^{(t+1)}\|_2}
$$
$$
\sigma(W) \approx {u^{(3)}}^T W v^{(3)}
$$
If the merged spectral radius grows more than 5% relative to the base model, it is scaled down:
$$
W_{\text{final}} = \begin{cases} W_{\text{merged}} \cdot \left(\frac{1.05 \cdot \sigma(W_0)}{\sigma(W_{\text{merged}})}\right) & \text{if } \sigma(W_{\text{merged}}) > 1.05 \cdot \sigma(W_0) \\ W_{\text{merged}} & \text{otherwise} \end{cases}
$$
---
## Layer-Stratified Component Policies
| Parameter Class | Target Identifiers | Applied Policy | Density (ρ) | Mathematical Constraints |
| :--- | :--- | :---: | :---: | :--- |
| **Embeddings & LM Head** | `embed_tokens`, `lm_head` | Low-Memory Streaming Blend | 1.0 | Convex iterative accumulation; vocab dimension aligned to 248,320. |
| **Linear State-Space Projections** | `linear_attn.in_proj_qkv` | Asymmetric GQA DELLA | 0.95 (QK) / 0.70 (V) | Sub-tensor slicing; separate routing and associative memory passes. |
| **Self-Attention Projections** | `self_attn.qkv_proj`, `o_proj` | Asymmetric GQA + Spectral Anchor | 0.95 (QK) / 0.70 (V) | Power-iteration clipping prevents σ > 1.05 σ₀. |
| **Feed-Forward Blocks** | `mlp.gate_proj`, `up_proj`, `down_proj` | Neuron-Gated HG-DELLA | 0.50 – 0.70 | Row-wise cosine filtering (cos θ ≥ -0.10); sign consensus. |
| **Recurrent Gates & Normalization** | `A_log`, `norm`, `conv1d` | Convex Parameter Blend | 1.0 | Preservation of *A*<sub>log</sub> ≤ 0 to guarantee BIBO stability. |
---
## Agentic Chat Template & Operational Directives
This model uses the [Improved Chat Template for Qwen 3.x by Olivia Rossi](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x) to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination.
---
## Recommended Generation Parameters
For deterministic software engineering and complex reasoning benchmarks:
| Parameter | Recommended Value | Description |
| :--- | :---: | :--- |
| **Temperature** | `0.6` | Balances strict syntactic validity with algorithmic path exploration. |
| **Top-P** | `0.95` | Eliminates low-probability token tails while preserving alternative logic paths. |
| **Top-K** | `20` | Restricts token candidate pools to prevent architectural syntax drift. |
| **Min-P** | `0.0` (Off) | Disabled in favor of explicit Top-K / Top-P governance. |
| **Repetition Penalty** | `1.0` (Off) | Disabled to prevent syntax degradation in repetitive code patterns (indentation, braces). |
| **Presence Penalty** | `0.0` | Prevents naming mutations across long-context symbol resolution. |
---
## How to Use
### Serving via vLLM
```bash
vllm serve pragmaticcs/PentaCoder-9B \
--dtype bfloat16 \
--max-model-len 65536 \
--gpu-memory-utilization 0.95 \
--enable-prefix-caching \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--enable-reasoning \
--reasoning-parser qwen3
```
### Inference via Transformers
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "pragmaticcs/PentaCoder-9B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto"
)
messages = [
{
"role": "system",
"content": "You are an expert systems engineer. Reason step by step and output clean, robust implementations.",
},
{
"role": "user",
"content": "Write an asynchronous connection pool in Python for TCP sockets with active health-checking, backpressure control, and graceful shutdown handling.",
},
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, enable_thinking=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=4096,
temperature=0.6,
top_p=0.95,
top_k=20,
do_sample=True,
)
response = tokenizer.decode(outputs[0][inputs.shape[-1] :], skip_special_tokens=True)
print(response)
```
---
## Citation & References
- [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)
- [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B)
- [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B)
- [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder)
- [OrionLLM/OxCoder-9B](https://huggingface.co/OrionLLM/OxCoder-9B)
- [empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill)
- [Improved Chat Template for Qwen 3.x](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x)
```bibtex
@inproceedings{yadav2023ties,
title={Resolving Interference When Merging Models},
author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
volume={36},
pages={7093--7115},
year={2023}
}
@article{deep2024della,
title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling},
author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya},
journal={arXiv preprint arXiv:2406.11617},
year={2024}
}
@article{jang2024modelstock,
title={Model Stock: All We Need Is just a Few Fine-Tuned Models},
author={Jang, Dong-Hwan and Yoon, Sang-Doo and Song, Gyeong-Moon},
journal={arXiv preprint arXiv:2403.19522},
year={2024}
}
``` |