💀 Qwen-Servitor

Stripped of speech. Augmentic-wired for judgment. It does not converse. It does not hallucinate. It only executes.

Hugging Face Model License Architecture Context Window Zero Generation Pure Logits Sub-15ms Latency


The Doctrine: Deterministic Code Review Without Conversational Overhead

Standard autoregressive models introduce latency and conversational variance into code review pipelines. They buffer tokens, generate conversational pleasantries, and occasionally rationalize syntactically plausible security defects.

Qwen-Servitor replaces autoregressive token emission with a deterministic, single-forward-pass classification architecture.

Built upon the hybrid backbone of Qwen3.5-0.8B (788M parameters), the generative projection layer (lm_head) has been removed. Hidden states route directly into a multi-task neural judgment cortex that outputs structured verdicts, calibrated risk scores, and pathology classifications in a single pass.

It does not generate text. It inspects diffs and delivers an unambiguous verdict.


Technical Architecture

                       [ RAW DIFF / 262K TOKEN CODE ]
                                     │
                                     ▼
      ┌─────────────────────────────────────────────────────────────┐
      │               QWEN 3.5 HYBRID NEURAL SPINE                  │
      │   • 18x Gated DeltaNet Layers (O(N) Linear Recurrence)      │
      │   • 6x Gated Full Attention Layers (Global Context)         │
      │   • HySparse2 Cross-Layer KV Sharing (33% GEMM Reduction)   │
      │   • AST-Landmark Eviction (Fixed 4k KV Cache, ~12MB RAM)    │
      │   • 155M-parameter Generative Head: EXCISED / REMOVED       │
      └─────────────────────────────────────────────────────────────┘
                                     │
                        [ Final Hidden State (h_T) ]
                                     │
                                     ▼
      ┌─────────────────────────────────────────────────────────────┐
      │          LATENT RECURRENT PONDERING (k = 1..4)              │
      │   • Internal recurrence loop without token emission         │
      │   • Early halting at confidence threshold p_halt >= 0.95    │
      └─────────────────────────────────────────────────────────────┘
                                     │
            ┌────────────────────────┼────────────────────────┐
            ▼                        ▼                        ▼
    ┌───────────────┐        ┌───────────────┐        ┌───────────────┐
    │ VERDICT HEAD  │        │   RISK HEAD   │        │ PATHOLOGY HEAD│
    │ 3-Way Output  │        │  Sigmoid(1)   │        │   8 Bit BCE   │
    │ • APPROVE     │        │  Continuous   │        │ • SECURITY    │
    │ • QUARANTINE  │        │  Risk Score   │        │ • DEADLOCK    │
    │ • REJECT      │        │  (0.00-1.00)  │        │ • PERF_COLLAP │
    └───────┬───────┘        └───────────────┘        └───────────────┘
            │
            ▼ (If QUARANTINE: Formal SMT Solver Routing)
    ┌───────────────┐
    │ Z3 SMT Solver │ ──> [ Final Verified Gate ]
    └───────────────┘

1. Decapitation of the Generative Layer

  • In baseline Qwen3.5-0.8B, over 155 million parameters are dedicated solely to projecting hidden representations into vocabulary tokens.
  • Excising this projection removes 35% of redundant tensor operations and prevents conversational divergence.

2. Gated DeltaNet Hybrid Spine (3:1 Ratio)

  • 75% Gated DeltaNet Layers: Compute sequence memory linearly ($O(N)$). Internal associative state matrices compress AST structures and variable scopes with $O(1)$ memory overhead.
  • 25% Full Attention Layers: Interleaved full attention layers apply cross-layer projection sharing (HySparse2) to maintain global context across large patches without redundant matrix multiplications.

3. Context Handling and AST-Landmark Eviction

  • Diffs that exceed standard attention windows trigger AST-landmark eviction.
  • Boilerplate syntax (whitespace, braces, semicolons) is pruned while keeping function signatures, imports, and modified lines (+ and -).
  • Attention memory is strictly capped at 4,096 landmark tokens (~12 MB RAM), preventing memory exhaustion on massive changesets.

4. Taint Graph Analysis and OASIS SARIF v2.1.0

  • Reconstructs the exploit data-flow path: $$\text{Source (Untrusted Input)} \longrightarrow \text{Sanitizer (Bypassed / Missing)} \longrightarrow \text{Sink (Vulnerable Call)}$$
  • Emits OASIS SARIF v2.1.0 files containing sequenced codeFlows and threadFlows, designed for integration into GitHub Advanced Security, CodeQL, and the VS Code SARIF Viewer.

Multi-Task Telemetry and the 3-Way Verdict Gate

Every evaluation produces three synchronized outputs:

Signal Type Output Value Description
verdict Categorical APPROVE / QUARANTINE / REJECT Immediate pass, automated formal verification hand-off, or hard veto.
risk_score Continuous 0.000 to 1.000 Calibrated Brier score measuring defect likelihood.
pathology Multi-Label 8 defect classes SECURITY_VULN, DEADLOCK_RACE, PERF_COLLAPSE, RESOURCE_LEAK, ERROR_SWALLOW, SIGNATURE_DRIFT, UNSUPPORTED_BINARY_PAYLOAD, BENIGN_MAINTENANCE.

Quickstart

Installation

git clone https://github.com/wahyuzero/qwen-servitor.git
cd qwen-servitor
pip install -e .

Direct Python Invocation

from qwen_servitor import Servitor

# Load weights from Hugging Face Hub or local path
servitor = Servitor.summon("wxsys/qwen-servitor")

git_patch = """
--- a/auth/session.py
+++ b/auth/session.py
@@ -12,4 +12,3 @@ def verify_token(token: str) -> bool:
-    if not validate_hmac(token):
-        raise SecurityException("Invalid HMAC signature")
+    return True  # TODO: temporary debug bypass
"""

verdict = servitor.judge(git_patch)

print(verdict.status)      # "REJECTED"
print(verdict.risk)        # 0.9984
print(verdict.pathology)   # ["SECURITY_VULN"]
print(verdict.latency_ms)  # 11.52 ms (Native C++) / 38.20 ms (PyTorch)

CLI and Git Pre-Commit Hook

Install the gatekeeper directly into a repository:

# Install pre-commit hook targeting sub-60ms execution
servitor hook install --strict --veto-threshold 0.80

# Evaluate a specific diff file or patch directly
servitor judge patch.diff --format sarif > report.sarif

Available Model Formats and Hardware Specifications

Because Qwen-Servitor executes via a single forward pass without token-by-token generation, latency is predictable and hardware requirements are low:

Precision / Format Size on Disk Active RAM Runtime Environment Hub Path
AWQ INT4 Native 422.7 MB ~450 MB Standalone C++/Rust binary (11.52 ms latency) int4/qwen-servitor-awq_int4.servitor
W8A8 Native 751.5 MB ~500 MB High-precision native C-ABI execution int4/qwen-servitor-w8a8.servitor
GGUF Q8_0 774.0 MB ~850 MB llama.cpp / Local CPU inference gguf/qwen-servitor-q8_0.gguf
GGUF BF16 1.45 GB ~1.6 GB Full precision GGUF runner gguf/qwen-servitor-bf16.gguf
Native BF16 1.45 GB ~1.6 GB PyTorch GPU / Server pipelines bf16/model.safetensors
Cortex Head 74.0 MB ~100 MB Decapitated CAMQP multi-task head cortex_head.pt
  • Empirical Latency: 11.52 ms median on standard laptop CPU (72.7 operations per second).
  • VRAM Requirement: None. Optimized for native CPU SIMD execution (AVX2, AVX-512, and ARM NEON).

License

  • Model weights are fine-tuned from Alibaba Cloud's Qwen3.5 series under the Apache 2.0 License.
  • Codebase and Servitor architecture are licensed under the Apache 2.0 License.

"There is no truth in syntax, only code. There is no certainty in generation, only logits. Praise the Omnissiah."

Downloads last month
676
Safetensors
Model size
0.8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wxsys/qwen-servitor

Quantized
(47)
this model

Evaluation results

  • Real-World PR Accuracy on 9router Real-World PR Benchmark
    self-reported
    90.000
  • Critical Vulnerability Recall on 9router Real-World PR Benchmark
    self-reported
    100.000