Instructions to use wxsys/qwen-servitor with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use wxsys/qwen-servitor with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="wxsys/qwen-servitor") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("wxsys/qwen-servitor") model = AutoModelForMultimodalLM.from_pretrained("wxsys/qwen-servitor", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use wxsys/qwen-servitor with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf wxsys/qwen-servitor:BF16 # Run inference directly in the terminal: llama cli -hf wxsys/qwen-servitor:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf wxsys/qwen-servitor:BF16 # Run inference directly in the terminal: llama cli -hf wxsys/qwen-servitor:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf wxsys/qwen-servitor:BF16 # Run inference directly in the terminal: ./llama-cli -hf wxsys/qwen-servitor:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf wxsys/qwen-servitor:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf wxsys/qwen-servitor:BF16
Use Docker
docker model run hf.co/wxsys/qwen-servitor:BF16
- LM Studio
- Jan
- Ollama
How to use wxsys/qwen-servitor with Ollama:
ollama run hf.co/wxsys/qwen-servitor:BF16
- Unsloth Desktop
- Pi
How to use wxsys/qwen-servitor with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wxsys/qwen-servitor:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "wxsys/qwen-servitor:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use wxsys/qwen-servitor with Docker Model Runner:
docker model run hf.co/wxsys/qwen-servitor:BF16
- Lemonade
How to use wxsys/qwen-servitor with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull wxsys/qwen-servitor:BF16
Run and chat with the model
lemonade run user.qwen-servitor-BF16
List all available models
lemonade list
- Hermes Agent
How to use wxsys/qwen-servitor with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wxsys/qwen-servitor:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default wxsys/qwen-servitor:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use wxsys/qwen-servitor with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wxsys/qwen-servitor:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "wxsys/qwen-servitor:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
💀 Qwen-Servitor
Stripped of speech. Augmentic-wired for judgment. It does not converse. It does not hallucinate. It only executes.
The Doctrine: Deterministic Code Review Without Conversational Overhead
Standard autoregressive models introduce latency and conversational variance into code review pipelines. They buffer tokens, generate conversational pleasantries, and occasionally rationalize syntactically plausible security defects.
Qwen-Servitor replaces autoregressive token emission with a deterministic, single-forward-pass classification architecture.
Built upon the hybrid backbone of Qwen3.5-0.8B (788M parameters), the generative projection layer (lm_head) has been removed. Hidden states route directly into a multi-task neural judgment cortex that outputs structured verdicts, calibrated risk scores, and pathology classifications in a single pass.
It does not generate text. It inspects diffs and delivers an unambiguous verdict.
Technical Architecture
[ RAW DIFF / 262K TOKEN CODE ]
│
▼
┌─────────────────────────────────────────────────────────────┐
│ QWEN 3.5 HYBRID NEURAL SPINE │
│ • 18x Gated DeltaNet Layers (O(N) Linear Recurrence) │
│ • 6x Gated Full Attention Layers (Global Context) │
│ • HySparse2 Cross-Layer KV Sharing (33% GEMM Reduction) │
│ • AST-Landmark Eviction (Fixed 4k KV Cache, ~12MB RAM) │
│ • 155M-parameter Generative Head: EXCISED / REMOVED │
└─────────────────────────────────────────────────────────────┘
│
[ Final Hidden State (h_T) ]
│
▼
┌─────────────────────────────────────────────────────────────┐
│ LATENT RECURRENT PONDERING (k = 1..4) │
│ • Internal recurrence loop without token emission │
│ • Early halting at confidence threshold p_halt >= 0.95 │
└─────────────────────────────────────────────────────────────┘
│
┌────────────────────────┼────────────────────────┐
▼ ▼ ▼
┌───────────────┐ ┌───────────────┐ ┌───────────────┐
│ VERDICT HEAD │ │ RISK HEAD │ │ PATHOLOGY HEAD│
│ 3-Way Output │ │ Sigmoid(1) │ │ 8 Bit BCE │
│ • APPROVE │ │ Continuous │ │ • SECURITY │
│ • QUARANTINE │ │ Risk Score │ │ • DEADLOCK │
│ • REJECT │ │ (0.00-1.00) │ │ • PERF_COLLAP │
└───────┬───────┘ └───────────────┘ └───────────────┘
│
▼ (If QUARANTINE: Formal SMT Solver Routing)
┌───────────────┐
│ Z3 SMT Solver │ ──> [ Final Verified Gate ]
└───────────────┘
1. Decapitation of the Generative Layer
- In baseline Qwen3.5-0.8B, over 155 million parameters are dedicated solely to projecting hidden representations into vocabulary tokens.
- Excising this projection removes 35% of redundant tensor operations and prevents conversational divergence.
2. Gated DeltaNet Hybrid Spine (3:1 Ratio)
- 75% Gated DeltaNet Layers: Compute sequence memory linearly ($O(N)$). Internal associative state matrices compress AST structures and variable scopes with $O(1)$ memory overhead.
- 25% Full Attention Layers: Interleaved full attention layers apply cross-layer projection sharing (HySparse2) to maintain global context across large patches without redundant matrix multiplications.
3. Context Handling and AST-Landmark Eviction
- Diffs that exceed standard attention windows trigger AST-landmark eviction.
- Boilerplate syntax (whitespace, braces, semicolons) is pruned while keeping function signatures, imports, and modified lines (
+and-). - Attention memory is strictly capped at 4,096 landmark tokens (~12 MB RAM), preventing memory exhaustion on massive changesets.
4. Taint Graph Analysis and OASIS SARIF v2.1.0
- Reconstructs the exploit data-flow path: $$\text{Source (Untrusted Input)} \longrightarrow \text{Sanitizer (Bypassed / Missing)} \longrightarrow \text{Sink (Vulnerable Call)}$$
- Emits OASIS SARIF v2.1.0 files containing sequenced
codeFlowsandthreadFlows, designed for integration into GitHub Advanced Security, CodeQL, and the VS Code SARIF Viewer.
Multi-Task Telemetry and the 3-Way Verdict Gate
Every evaluation produces three synchronized outputs:
| Signal | Type | Output Value | Description |
|---|---|---|---|
verdict |
Categorical | APPROVE / QUARANTINE / REJECT |
Immediate pass, automated formal verification hand-off, or hard veto. |
risk_score |
Continuous | 0.000 to 1.000 |
Calibrated Brier score measuring defect likelihood. |
pathology |
Multi-Label | 8 defect classes | SECURITY_VULN, DEADLOCK_RACE, PERF_COLLAPSE, RESOURCE_LEAK, ERROR_SWALLOW, SIGNATURE_DRIFT, UNSUPPORTED_BINARY_PAYLOAD, BENIGN_MAINTENANCE. |
Quickstart
Installation
git clone https://github.com/wahyuzero/qwen-servitor.git
cd qwen-servitor
pip install -e .
Direct Python Invocation
from qwen_servitor import Servitor
# Load weights from Hugging Face Hub or local path
servitor = Servitor.summon("wxsys/qwen-servitor")
git_patch = """
--- a/auth/session.py
+++ b/auth/session.py
@@ -12,4 +12,3 @@ def verify_token(token: str) -> bool:
- if not validate_hmac(token):
- raise SecurityException("Invalid HMAC signature")
+ return True # TODO: temporary debug bypass
"""
verdict = servitor.judge(git_patch)
print(verdict.status) # "REJECTED"
print(verdict.risk) # 0.9984
print(verdict.pathology) # ["SECURITY_VULN"]
print(verdict.latency_ms) # 11.52 ms (Native C++) / 38.20 ms (PyTorch)
CLI and Git Pre-Commit Hook
Install the gatekeeper directly into a repository:
# Install pre-commit hook targeting sub-60ms execution
servitor hook install --strict --veto-threshold 0.80
# Evaluate a specific diff file or patch directly
servitor judge patch.diff --format sarif > report.sarif
Available Model Formats and Hardware Specifications
Because Qwen-Servitor executes via a single forward pass without token-by-token generation, latency is predictable and hardware requirements are low:
| Precision / Format | Size on Disk | Active RAM | Runtime Environment | Hub Path |
|---|---|---|---|---|
| AWQ INT4 Native | 422.7 MB | ~450 MB | Standalone C++/Rust binary (11.52 ms latency) | int4/qwen-servitor-awq_int4.servitor |
| W8A8 Native | 751.5 MB | ~500 MB | High-precision native C-ABI execution | int4/qwen-servitor-w8a8.servitor |
| GGUF Q8_0 | 774.0 MB | ~850 MB | llama.cpp / Local CPU inference | gguf/qwen-servitor-q8_0.gguf |
| GGUF BF16 | 1.45 GB | ~1.6 GB | Full precision GGUF runner | gguf/qwen-servitor-bf16.gguf |
| Native BF16 | 1.45 GB | ~1.6 GB | PyTorch GPU / Server pipelines | bf16/model.safetensors |
| Cortex Head | 74.0 MB | ~100 MB | Decapitated CAMQP multi-task head | cortex_head.pt |
- Empirical Latency: 11.52 ms median on standard laptop CPU (72.7 operations per second).
- VRAM Requirement: None. Optimized for native CPU SIMD execution (AVX2, AVX-512, and ARM NEON).
License
- Model weights are fine-tuned from Alibaba Cloud's Qwen3.5 series under the Apache 2.0 License.
- Codebase and Servitor architecture are licensed under the Apache 2.0 License.
"There is no truth in syntax, only code. There is no certainty in generation, only logits. Praise the Omnissiah."
- Downloads last month
- 676
Model tree for wxsys/qwen-servitor
Base model
Qwen/Qwen3.5-0.8B-BaseEvaluation results
- Real-World PR Accuracy on 9router Real-World PR Benchmarkself-reported90.000
- Critical Vulnerability Recall on 9router Real-World PR Benchmarkself-reported100.000