Text Generation
Transformers
Safetensors
English
Chinese
qwen3_5_text
Merge
ties
della
model-stock
geodella
qwen
qwen3.5
causal-lm
deltanet
linear-attention
agentic
reasoning
code
swe-bench
conversational
Instructions to use pragmaticcs/PentaCoder-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pragmaticcs/PentaCoder-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="pragmaticcs/PentaCoder-9B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("pragmaticcs/PentaCoder-9B") model = AutoModelForCausalLM.from_pretrained("pragmaticcs/PentaCoder-9B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use pragmaticcs/PentaCoder-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pragmaticcs/PentaCoder-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/PentaCoder-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/pragmaticcs/PentaCoder-9B
- SGLang
How to use pragmaticcs/PentaCoder-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "pragmaticcs/PentaCoder-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/PentaCoder-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "pragmaticcs/PentaCoder-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/PentaCoder-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use pragmaticcs/PentaCoder-9B with Docker Model Runner:
docker model run hf.co/pragmaticcs/PentaCoder-9B
|
Download README.md from pragmaticcs/PentaCoder-9B: direct link, hf CLI and curl.
- Browser
- Download file 17.4 kB
-
https://huggingface.co/pragmaticcs/PentaCoder-9B/resolve/main/README.md
- Command line
-
hf download hf://pragmaticcs/PentaCoder-9B/README.md
-
curl -L -o README.md https://huggingface.co/pragmaticcs/PentaCoder-9B/resolve/main/README.md
17.4 kB
| base_model: | |
| - Qwen/Qwen3.5-9B | |
| - XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B | |
| - ornith-ai/Ornith-1.5-9B | |
| - Jackrong/Qwopus3.5-9B-Coder | |
| - OrionLLM/OxCoder-9B | |
| - empero-ai/Qwen3.8-9B-Distill | |
| base_model_relation: merge | |
| library_name: transformers | |
| tags: | |
| - merge | |
| - ties | |
| - della | |
| - model-stock | |
| - geodella | |
| - qwen | |
| - qwen3.5 | |
| - causal-lm | |
| - deltanet | |
| - linear-attention | |
| - agentic | |
| - reasoning | |
| - code | |
| - swe-bench | |
| license: apache-2.0 | |
| language: | |
| - en | |
| - zh | |
| pipeline_tag: text-generation | |
| model_type: qwen3_5_text | |
|  | |
|  | |
| <div align="center"> | |
| [](https://opensource.org/licenses/Apache-2.0) | |
| [](https://github.com/huggingface/transformers) | |
| [](#merge-methodology--mathematical-formulation) | |
| [](#architectural-specifications) | |
| [-EB5757?style=for-the-badge)](#architectural-specifications) | |
| [](#architectural-specifications) | |
| </div> | |
| Most sub-10B coding models fail in realistic agentic environments due to a common trade-off: aggressive fine-tuning on synthetic coding instructions improves immediate benchmark pass rates but introduces brittle syntactic degradation and catastrophic repetition loops when shell commands or compiler checks fail. | |
| **PentaCoder** addresses these limitations by uniting five specialized post-trained checkpoints of [**Qwen 3.5 9B**](https://huggingface.co/Qwen/Qwen3.5-9B) via **GeoDELLA** (Geometric Drop-and-Rescale with Task-Covariance De-Biasing and Spectral Norm Anchoring). The model synthesizes the distinct mathematical distributions of each donor: | |
| - **Algorithmic Correctness & Architectural Decomposition** from [**Qwopus3.5-Coder**](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) (Claude 3.5 Opus distillation trajectories). | |
| - **Multi-Turn SWE-bench Planning & Tool Protocol Integrity** from [**MiMo-V2.6-Distill**](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) (Large-scale agentic execution traces). | |
| - **Terminal Execution Discipline & Error-Recovery Heuristics** from [**Ornith-1.5**](https://huggingface.co/ornith-ai/Ornith-1.5-9B) (Reinforcement learning for anti-looping). | |
| - **Low-Level Systems Implementation & Runtime Robustness** from [**OxCoder**](https://huggingface.co/OrionLLM/OxCoder-9B) (Deep API, CLI, and operational coding specialization). | |
| - **Abstract Structural Reasoning & Syntax Grounding** from [**Qwen3.8-Distill**](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) (Distilled frontier chain-of-thought representations). | |
| > [!IMPORTANT] | |
| > The result is a lean, blisteringly fast 9B pure-text causal engine with a native **256k context window** that runs comfortably on consumer GPUs. | |
| --- | |
| ### Contents | |
| - [Architectural Specifications](#architectural-specifications) | |
| - [Composition & Donor Weighting](#composition--donor-weighting) | |
| - [Merge Methodology & Mathematical Formulation](#merge-methodology--mathematical-formulation) | |
| - [Layer-Stratified Component Policies](#layer-stratified-component-policies) | |
| - [Agentic Chat Template & Operational Directives](#agentic-chat-template--operational-directives) | |
| - [Recommended Generation Parameters](#recommended-generation-parameters) | |
| - [How to Use](#how-to-use) | |
| - [Citation & References](#citation--references) | |
| --- | |
| ## Architectural Specifications | |
| | Parameter | Specification | | |
| | :--- | :---: | | |
| | **Total Parameters** | 8.8B (Pure Text Backbone) | | |
| | **Architecture Type** | Hybrid Recurrent-Attention Causal LM (`qwen3_5_text`) | | |
| | **Hidden Dimension** (*d*<sub>model</sub>) | 4096 | | |
| | **Intermediate Dimension** (*d*<sub>mlp</sub>) | 12288 (SwiGLU) | | |
| | **Decoder Layers** | 32 | | |
| | **Attention Layout** | 8 Blocks × (3 Gated DeltaNet Linear Layers : 1 Gated Softmax Layer) | | |
| | **Full Attention Layers** | Layers 3, 7, 11, 15, 19, 23, 27, 31 | | |
| | **Linear Attention Configuration** | 16 Key Heads / 32 Value Heads (*d*<sub>k</sub> = *d*<sub>v</sub> = 128) | | |
| | **Full Attention Configuration** | 16 Query Heads / 4 Key-Value Heads (GQA, *d*<sub>h</sub> = 256) | | |
| | **Rotary Position Embedding (RoPE)** | 1D Partial RoPE (θ = 10⁷, Factor = 0.25 → 64 dimensions) | | |
| | **Context Window Length** | 262,144 tokens (256k) | | |
| | **Native Precision** | `bfloat16` | | |
| | **Vocabulary Size** | 248,320 (Padded for Fill-In-The-Middle and Tool Tokens) | | |
| --- | |
| ## Composition & Donor Weighting | |
| The foundational weights of [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) serve as the shared topological base (*W*₀). Five specialized donor checkpoints provide non-overlapping task vectors mapped across normalized layer depth *u* ∈ [0, 1]: | |
| | Model | Primary Focus | Depth Target | | |
| | :--- | :--- | :---: | | |
| | [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) | Base pre-trained manifold and state-space anchors | Global (*W*₀) | | |
| | [empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) | Abstract token synthesis, reasoning structure, syntax | Lower & Mid Decoders | | |
| | [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) | Typing discipline, algorithm design, functional purity | Mid Decoders (Bell Curve) | | |
| | [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B) | Circuit-breaker recovery, environment feedback integration | Upper-Mid Decoders | | |
| | [OrionLLM/OxCoder-9B](https://huggingface.co/OrionLLM/OxCoder-9B) | Systems engineering, runtime bug localization, CLI tooling | Deep Layers (Ascending Ramp) | | |
| | [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) | SWE-bench multi-step planning, long-context tool invocation | Broad Central Decoders | | |
| --- | |
| ## Merge Methodology & Mathematical Formulation | |
| The merge was executed using the **GeoDELLA-Coder** pipeline, which solves weight interference through dynamic task-correlation de-biasing, row-wise sign consensus, sub-tensor GQA disentanglement, and power-iteration spectral norm stabilization. | |
| ### 1. Task Vector Formulation | |
| For each donor checkpoint *k* ∈ {1, ..., 5}, the parameter update delta τ<sub>*k*</sub> is computed relative to the base anchor *W*₀: | |
| $$ | |
| \tau_k = D_k - W_0, \quad k \in \{\text{Qwen3.8}, \text{Qwopus}, \text{Ornith}, \text{OxCoder}, \text{MiMo}\} | |
| $$ | |
| ### 2. Hybrid-Aware Continuous Depth Modulation | |
| Task vector mixing coefficients are continuously modulated over normalized depth *u* = *l* / (*L* - 1), where *l* ∈ [0, 31] and *L* = 32: | |
| $$ | |
| u_{\text{qwen38}}(u) = 0.15 + 0.35 \cos^2\left(\frac{\pi}{2} u\right) + 0.15 \sin^2(\pi u) | |
| $$ | |
| $$ | |
| u_{\text{qwopus}}(u) = 0.05 + 0.35 \sin^2(\pi u) | |
| $$ | |
| $$ | |
| u_{\text{ornith}}(u) = 0.05 + 0.35 \sin^2\left(\frac{\pi}{2} u\right) | |
| $$ | |
| $$ | |
| u_{\text{oxcoder}}(u) = 0.05 + 0.25 \sin^2\left(\frac{\pi}{2} u\right) | |
| $$ | |
| $$ | |
| u_{\text{mimo}}(u) = 0.05 + 0.15 \sin(\pi u) | |
| $$ | |
| The initial coefficients are normalized to form a partition of unity across all layers: | |
| $$ | |
| w_k(l) = \frac{u_k(u)}{\sum_{j=1}^5 u_j(u)}, \quad \sum_{k=1}^5 w_k(l) = 1.0 | |
| $$ | |
| ### 3. Dynamic Gram-Matrix Task De-Biasing | |
| Fine-tuned models frequently share underlying distillation datasets, leading to collinearity that can drown out specialized task vectors. To correct for this over-representation, the empirical Gram correlation matrix *G* ∈ ℝ<sup>*K* × *K*</sup> is evaluated per tensor: | |
| $$ | |
| G_{ij} = \frac{|\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle|}{\|\tau_i\|_2 \|\tau_j\|_2} | |
| $$ | |
| A uniqueness coefficient *U*<sub>*k*</sub> is derived from the inverse column sum of task vector correlations: | |
| $$ | |
| U_k = \frac{1}{\sum_{j=1}^K G_{kj}} | |
| $$ | |
| The dynamic weights are then re-balanced and normalized: | |
| $$ | |
| \tilde{w}_k = \frac{w_k U_k}{\sum_{j=1}^K w_j U_j} | |
| $$ | |
| This prevents dataset overlap from suppressing distinct algorithmic representations. | |
| ### 4. Asymmetric GQA Disentangled QKV Slicing | |
| Qwen 3.5 9B features asymmetric head ratios in both its linear attention and full attention layers. Standard fused tensor merging causes cross-head pollution by treating routing projections identically to memory projections. | |
| Fused projection tensors are sliced into their functional sub-matrices prior to merging: | |
| - **Gated DeltaNet Layers (8192 × *d*<sub>model</sub>):** Sliced into Query (2048), Key (2048), and Value (4096). | |
| - **Gated Attention Layers (6144 × *d*<sub>model</sub>):** Sliced into Query (4096), Key (1024), and Value (1024). | |
| Query and Key slices are processed with a conservative retention density (ρ = 0.95) to preserve sharp context routing. Value matrices are processed with adaptive MLP density (ρ = 0.70) to maximize conceptual synthesis. The components are then re-concatenated along the head dimension. | |
| ### 5. Neuron-Coherent Row Gating | |
| To eliminate destructive interference in Feed-Forward Networks (MLPs), task vectors are gated at the single-neuron (row) level: | |
| $$ | |
| \bar{\tau} = \frac{1}{K} \sum_{k=1}^K \tau_k | |
| $$ | |
| For each row *r* of donor delta τ<sub>*k*</sub>, the directional alignment with the consensus mean is evaluated: | |
| $$ | |
| \cos \theta_{k, r} = \frac{\langle \tau_{k, r}, \bar{\tau}_r \rangle}{\|\tau_{k, r}\|_2 \|\bar{\tau}_r\|_2 + \epsilon} | |
| $$ | |
| Rows exhibiting severe directional opposition (cos θ<sub>*k*, *r*</sub> < -0.10) are masked out: | |
| $$ | |
| \hat{\tau}_{k, r} = \tau_{k, r} \cdot \mathbb{I}\left(\cos \theta_{k, r} \ge -0.10\right) | |
| $$ | |
| This eliminates opposing gradient vectors that produce incoherent syntax generation. | |
| ### 6. Heavy-Tailed DELLA Adaptive Rescaling | |
| Surviving parameters undergo non-linear magnitude-based sampling. Using parameter rank indices *R*<sub>*k*, *ij*</sub> ∈ [0, 1] sorted by absolute magnitude, a Pareto-style retention probability *p*<sub>*k*, *ij*</sub> is established: | |
| $$ | |
| p_{k, ij} = p_{\min} + (p_{\max} - p_{\min}) \cdot \sqrt{R_{k, ij}} | |
| $$ | |
| Parameters are sampled via a Bernoulli trial and rescaled by their inverse survival probability: | |
| $$ | |
| M_{k, ij} \sim \text{Bernoulli}(p_{k, ij}) | |
| $$ | |
| $$ | |
| \tilde{\tau}_{k, ij} = \frac{\hat{\tau}_{k, ij} \odot M_{k, ij}}{p_{k, ij}} | |
| $$ | |
| ### 7. Coordinate-Wise Sign Consensus & Model Stock Scaling | |
| Directional consensus is determined via weighted sign agreement: | |
| $$ | |
| \Gamma = \operatorname{sgn}\left(\sum_{k=1}^K \tilde{w}_k \tilde{\tau}_k\right) | |
| $$ | |
| $$ | |
| A_k = \mathbb{I}\left(\operatorname{sgn}(\tilde{\tau}_k) = \Gamma\right) \odot \mathbb{I}\left(\tilde{\tau}_k \neq 0\right) | |
| $$ | |
| $$ | |
| \Delta_{\text{consensus}} = \frac{\sum_{k=1}^K \tilde{\tau}_k \odot A_k}{\sum_{k=1}^K A_k + \epsilon} | |
| $$ | |
| The aggregated delta is projected onto the non-linear manifold using the Model Stock analytic scaling factor *t*<sup>*</sup>: | |
| $$ | |
| \bar{\rho} = \frac{2}{K(K-1)} \sum_{i < j} \frac{\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle}{\|\tau_i\|_2 \|\tau_j\|_2} | |
| $$ | |
| $$ | |
| t^* = \frac{K \bar{\rho}}{1 + (K - 1)\bar{\rho}} | |
| $$ | |
| $$ | |
| \Delta_{\text{final}} = t^* \cdot \Delta_{\text{consensus}} | |
| $$ | |
| ### 8. Spectral Norm & Attention Entropy Anchoring | |
| Attention projection matrices (*W*<sub>attn</sub>) are vulnerable to spectral explosion during merges, which contracts attention entropy and leads to repetitive generation loops. | |
| The dominant singular value σ(*W*) is calculated via a three-iteration deterministic power iteration: | |
| $$ | |
| v^{(t+1)} = \frac{W^T u^{(t)}}{\|W^T u^{(t)}\|_2}, \quad u^{(t+1)} = \frac{W v^{(t+1)}}{\|W v^{(t+1)}\|_2} | |
| $$ | |
| $$ | |
| \sigma(W) \approx {u^{(3)}}^T W v^{(3)} | |
| $$ | |
| If the merged spectral radius grows more than 5% relative to the base model, it is scaled down: | |
| $$ | |
| W_{\text{final}} = \begin{cases} W_{\text{merged}} \cdot \left(\frac{1.05 \cdot \sigma(W_0)}{\sigma(W_{\text{merged}})}\right) & \text{if } \sigma(W_{\text{merged}}) > 1.05 \cdot \sigma(W_0) \\ W_{\text{merged}} & \text{otherwise} \end{cases} | |
| $$ | |
| --- | |
| ## Layer-Stratified Component Policies | |
| | Parameter Class | Target Identifiers | Applied Policy | Density (ρ) | Mathematical Constraints | | |
| | :--- | :--- | :---: | :---: | :--- | | |
| | **Embeddings & LM Head** | `embed_tokens`, `lm_head` | Low-Memory Streaming Blend | 1.0 | Convex iterative accumulation; vocab dimension aligned to 248,320. | | |
| | **Linear State-Space Projections** | `linear_attn.in_proj_qkv` | Asymmetric GQA DELLA | 0.95 (QK) / 0.70 (V) | Sub-tensor slicing; separate routing and associative memory passes. | | |
| | **Self-Attention Projections** | `self_attn.qkv_proj`, `o_proj` | Asymmetric GQA + Spectral Anchor | 0.95 (QK) / 0.70 (V) | Power-iteration clipping prevents σ > 1.05 σ₀. | | |
| | **Feed-Forward Blocks** | `mlp.gate_proj`, `up_proj`, `down_proj` | Neuron-Gated HG-DELLA | 0.50 – 0.70 | Row-wise cosine filtering (cos θ ≥ -0.10); sign consensus. | | |
| | **Recurrent Gates & Normalization** | `A_log`, `norm`, `conv1d` | Convex Parameter Blend | 1.0 | Preservation of *A*<sub>log</sub> ≤ 0 to guarantee BIBO stability. | | |
| --- | |
| ## Agentic Chat Template & Operational Directives | |
| This model uses the [Improved Chat Template for Qwen 3.x by Olivia Rossi](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x) to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination. | |
| --- | |
| ## Recommended Generation Parameters | |
| For deterministic software engineering and complex reasoning benchmarks: | |
| | Parameter | Recommended Value | Description | | |
| | :--- | :---: | :--- | | |
| | **Temperature** | `0.6` | Balances strict syntactic validity with algorithmic path exploration. | | |
| | **Top-P** | `0.95` | Eliminates low-probability token tails while preserving alternative logic paths. | | |
| | **Top-K** | `20` | Restricts token candidate pools to prevent architectural syntax drift. | | |
| | **Min-P** | `0.0` (Off) | Disabled in favor of explicit Top-K / Top-P governance. | | |
| | **Repetition Penalty** | `1.0` (Off) | Disabled to prevent syntax degradation in repetitive code patterns (indentation, braces). | | |
| | **Presence Penalty** | `0.0` | Prevents naming mutations across long-context symbol resolution. | | |
| --- | |
| ## How to Use | |
| ### Serving via vLLM | |
| ```bash | |
| vllm serve pragmaticcs/PentaCoder-9B \ | |
| --dtype bfloat16 \ | |
| --max-model-len 65536 \ | |
| --gpu-memory-utilization 0.95 \ | |
| --enable-prefix-caching \ | |
| --enable-auto-tool-choice \ | |
| --tool-call-parser qwen3_coder \ | |
| --enable-reasoning \ | |
| --reasoning-parser qwen3 | |
| ``` | |
| ### Inference via Transformers | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "pragmaticcs/PentaCoder-9B" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, torch_dtype=torch.bfloat16, device_map="auto" | |
| ) | |
| messages = [ | |
| { | |
| "role": "system", | |
| "content": "You are an expert systems engineer. Reason step by step and output clean, robust implementations.", | |
| }, | |
| { | |
| "role": "user", | |
| "content": "Write an asynchronous connection pool in Python for TCP sockets with active health-checking, backpressure control, and graceful shutdown handling.", | |
| }, | |
| ] | |
| inputs = tokenizer.apply_chat_template( | |
| messages, add_generation_prompt=True, enable_thinking=True, return_tensors="pt" | |
| ).to(model.device) | |
| outputs = model.generate( | |
| inputs, | |
| max_new_tokens=4096, | |
| temperature=0.6, | |
| top_p=0.95, | |
| top_k=20, | |
| do_sample=True, | |
| ) | |
| response = tokenizer.decode(outputs[0][inputs.shape[-1] :], skip_special_tokens=True) | |
| print(response) | |
| ``` | |
| --- | |
| ## Citation & References | |
| - [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) | |
| - [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) | |
| - [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B) | |
| - [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) | |
| - [OrionLLM/OxCoder-9B](https://huggingface.co/OrionLLM/OxCoder-9B) | |
| - [empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill) | |
| - [Improved Chat Template for Qwen 3.x](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x) | |
| ```bibtex | |
| @inproceedings{yadav2023ties, | |
| title={Resolving Interference When Merging Models}, | |
| author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit}, | |
| booktitle={Advances in Neural Information Processing Systems (NeurIPS)}, | |
| volume={36}, | |
| pages={7093--7115}, | |
| year={2023} | |
| } | |
| @article{deep2024della, | |
| title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling}, | |
| author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya}, | |
| journal={arXiv preprint arXiv:2406.11617}, | |
| year={2024} | |
| } | |
| @article{jang2024modelstock, | |
| title={Model Stock: All We Need Is just a Few Fine-Tuned Models}, | |
| author={Jang, Dong-Hwan and Yoon, Sang-Doo and Song, Gyeong-Moon}, | |
| journal={arXiv preprint arXiv:2403.19522}, | |
| year={2024} | |
| } | |
| ``` |