Instructions to use poisonxa/PXA-Supporter-Models with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use poisonxa/PXA-Supporter-Models with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf poisonxa/PXA-Supporter-Models # Run inference directly in the terminal: llama cli -hf poisonxa/PXA-Supporter-Models
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf poisonxa/PXA-Supporter-Models # Run inference directly in the terminal: llama cli -hf poisonxa/PXA-Supporter-Models
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf poisonxa/PXA-Supporter-Models # Run inference directly in the terminal: ./llama-cli -hf poisonxa/PXA-Supporter-Models
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf poisonxa/PXA-Supporter-Models # Run inference directly in the terminal: ./build/bin/llama-cli -hf poisonxa/PXA-Supporter-Models
Use Docker
docker model run hf.co/poisonxa/PXA-Supporter-Models
- LM Studio
- Jan
- Ollama
How to use poisonxa/PXA-Supporter-Models with Ollama:
ollama run hf.co/poisonxa/PXA-Supporter-Models
- Unsloth Desktop
- Pi
How to use poisonxa/PXA-Supporter-Models with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf poisonxa/PXA-Supporter-Models
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "poisonxa/PXA-Supporter-Models" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use poisonxa/PXA-Supporter-Models with Docker Model Runner:
docker model run hf.co/poisonxa/PXA-Supporter-Models
- Lemonade
How to use poisonxa/PXA-Supporter-Models with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull poisonxa/PXA-Supporter-Models
Run and chat with the model
lemonade run user.PXA-Supporter-Models-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use poisonxa/PXA-Supporter-Models with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf poisonxa/PXA-Supporter-Models
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default poisonxa/PXA-Supporter-Models
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use poisonxa/PXA-Supporter-Models with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf poisonxa/PXA-Supporter-Models
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "poisonxa/PXA-Supporter-Models" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
PXA Network supporters only
Access is approved automatically for Supporters and Valued Supporters on the PXA Network Discord. Put your Discord username in the form below and you're approved within 5 minutes (not in the Discord yet? https://discord.gg/EqazvV9tf). Support on Ko-fi: https://ko-fi.com/shatteredrealms1
Log in or Sign Up to review the conditions and access this model content.
PXA-Supporter-Models
Supporter-only model downloads from PXA Labs for the PXA Network community. This repo is for Supporters and Valued Supporters. (The Supporter and Valued Supporter Discord roles both open it.)
How to get access
- Join the PXA Network Discord: https://discord.gg/EqazvV9tf
- Support on Ko-fi: https://ko-fi.com/shatteredrealms1. Ko-fi gives you the supporter role in the Discord automatically.
- In the Discord, run
/hf-link <your-hugging-face-username>(only you see the reply)./hf-statusshows what you have access to. - Click Request supporter access on this page. The PXA Network bot approves it within a few minutes.
If you requested before linking, the request is declined with a note. You do not need to request again: once you are linked and hold the role, access is granted automatically.
Access follows the Discord role. If the supporter role ends, access is removed in the nightly check. It comes back automatically when the role returns.
Community
- Discord: https://discord.gg/EqazvV9tf
- Ko-fi: https://ko-fi.com/shatteredrealms1
Models
All files here are PXQN (PXQ-Next) quants from PXA Labs. They need pxa v2026.10.2 or later (https://github.com/poisonxa16/pxa/releases/latest). v2026.09.20 and older cannot load PXQN files. Quality is reported as KL divergence (KLD) on assistant-role tokens of multi-turn chat: that is the text the model writes, which is what you read. Lower is better, and 0 is identical to the reference.
| Folder | Base model | Size | bpw | Assistant KLD | Cards |
|---|---|---|---|---|---|
Qwen3.8-27B-PXQN-onecard/ |
Qwen3.8-27B | 13.57 GB (12.63 GiB) | 3.97 | 0.0213 | one 16 GB card, 128k context |
Qwen3.8-27B-PXQN4/ |
Qwen3.8-27B | 15.72 GB (14.63 GiB) | 4.60 | 0.0079 | two 16 GB cards |
Qwen3.8-27B-PXQN5/ |
Qwen3.8-27B | 18.79 GB (17.49 GiB) | 5.50 | 0.0022 | two 16 GB cards |
Qwen3.8-Flash-Next-PXQN/ |
Qwen3.8-Flash-Next (Uncensored) | 98.66 GB (91.87 GiB), 3 parts | 4.46 | 0.0705 (vs a Q8_0-dense reference) | four 16 GB cards + 64 GB system RAM |
Swift-1.5-Qwen3.8-Flash-Next-PXQN-48GB/ |
Swift 1.5 Qwen3.8-Flash-Next | 98.84 GB (92.06 GiB) | - | pending | three 16 GB cards |
Swift-1.5-Qwen3.8-Flash-Next-PXQN-64GB/ |
Swift 1.5 Qwen3.8-Flash-Next | 110.90 GB (103.28 GiB) | - | pending | four 16 GB cards |
Swift-1.5-Qwen3.8-Flash-Next-PXQN-80GB/ |
Swift 1.5 Qwen3.8-Flash-Next | 123.48 GB (115.00 GiB) | - | pending | five 16 GB cards |
Qwen3.8-27B-PXQN-onecard
Qwen3.8-27B, dense, with its MTP head, built to run 128k context on one 16 GB card (tested on a Tesla V100 16 GB and a Tesla P100 16 GB). The mix gives more bits to the layers that matter most, so it has 14% lower KLD than our standard balanced mix of the same size.
- File:
Qwen3.8-27B-PXQN-onecard/Qwen3.8-27B-PXQN-onecard.gguf, 13,566,509,376 bytes, 3.97 bpw. - Quality: assistant-token KLD 0.0213 against Q8_0 of the original; the top token matches it 95% of the time (PPL ratio 1.002).
- Command (one 16 GB card, 131072 context, q4_0 KV cache, MTP speculative decoding at depth 1):
llama-server -m Qwen3.8-27B-PXQN-onecard.gguf -ngl 99 -np 1 -c 131072 -ctk q4_0 -ctv q4_0 -fa on -ub 512 --spec-type mtp:n_max=1
- Measured on our box: V100 about 30 t/s plain and 45-48 t/s with MTP (acceptance about 0.8); P100 about 23 t/s plain and 23-26 t/s with MTP. A 128k needle test passes on both cards. Prefill on V100: 912 t/s at 4k, 341 t/s at 128k.
- On a P100 at very long context (64k+), plain decoding (drop
--spec-type) is slightly faster than MTP. - sha256:
8fccd7aa8f671a2f78cbd364564e9423a8c8f35c9f8a9227cdee4cf59cd6ba30
Qwen3.8-27B-PXQN4
Qwen3.8-27B, dense, at the PXQN4 tier: near-lossless quality at 4.6 bpw.
- File:
Qwen3.8-27B-PXQN4/Qwen3.8-27B-PXQN4.gguf, 15,720,262,976 bytes, 4.60 bpw. - Quality: assistant-token KLD 0.0079 against Q8_0 of the original; the top token matches 97.1% of the time.
- Cards: two 16 GB cards (tested on 2x V100 and 2x P100 with tensor split), and it should also fit one 24 GB card (not tested).
- Command (two cards):
llama-server -m Qwen3.8-27B-PXQN4.gguf -ngl 99 -sm tensor -fa on -np 1 -c 32768
- sha256:
04dcba334f1a9613652b148663e6198fe996a7110cdefb2f9e320b25b5a47979
Qwen3.8-27B-PXQN5
Qwen3.8-27B, dense, at the PXQN5 tier: the highest-fidelity file here at 5.5 bpw.
- File:
Qwen3.8-27B-PXQN5/Qwen3.8-27B-PXQN5.gguf, 18,785,381,696 bytes, 5.50 bpw. - Quality: assistant-token KLD 0.0022 against Q8_0 of the original; the top token matches 98.4% of the time.
- Cards: two 16 GB cards (tested on 2x V100), and it should also fit one 24 GB card (not tested).
- Command (two cards):
llama-server -m Qwen3.8-27B-PXQN5.gguf -ngl 99 -sm tensor -fa on -np 1 -c 32768
- sha256:
2ac462d1bbbaa422d18e72a7b8c8940f8677333a23e5837a4849140db3bdd8ab
Qwen3.8-Flash-Next-PXQN
Qwen3.8 Flash-Next Uncensored (hybrid MoE, 177B total parameters, 512 experts with 10 active per token) in PXQN (PXQ-Next), our best Flash-Next file so far: every layer, experts included, is PXQN.
- Files (one model in three parts; point the engine at part 1 and it loads the rest from the same folder):
Qwen3.8-Flash-Next-PXQN/Qwen3.8-Flash-Next-Uncensored-PXQN-00001-of-00003.gguf, 1,067,038,400 bytesQwen3.8-Flash-Next-PXQN/Qwen3.8-Flash-Next-Uncensored-PXQN-00002-of-00003.gguf, 54,400,261,312 bytesQwen3.8-Flash-Next-PXQN/Qwen3.8-Flash-Next-Uncensored-PXQN-00003-of-00003.gguf, 43,192,937,632 bytes- total 98,660,237,024 bytes (91.87 GiB), 4.46 bpw.
- Part 2 is larger than 50 GB (it holds the 51 GiB per-layer embedding table, which cannot be split). A plain browser or
wgetdownload of files over 50 GB fails on Hugging Face. Download with the Hugging Face CLI, which handles it:hf download poisonxa/PXA-Supporter-Models --include "Qwen3.8-Flash-Next-PXQN/*" --local-dir .(log in first withhf auth login). - Quality: assistant-token KLD 0.0705 against our high-precision reference build (Q8_0 dense layers; the BF16 original is 306 GB, too large for the instrument). Our previous Flash-Next file scores 0.0879 on the same test, so this one has 20% lower KLD. The top token matches the reference 91.2% of the time.
- Cards: four 16 GB cards (tested on 4x Tesla P100 16 GB) plus at least 64 GB of system RAM: the 51 GiB per-layer embedding table stays in host memory by design.
- Measured on 4x P100 with tensor split, 32k context: decode about 35 t/s (7-8% faster than our previous Flash-Next file), prefill about 520 t/s at 3k tokens and 477 t/s at 16k.
- Command (4 cards):
llama-server -m Qwen3.8-Flash-Next-Uncensored-PXQN-00001-of-00003.gguf -ngl 99 -sm tensor -fa on -c 32768 -ub 512 -ot "per_layer_token_embd\.weight=CPU"
The pxa launcher picks these per-card settings for you.
- sha256:
- part 1:
8fdaa1fcc8c38f4ee95b63239cb34326155941fe7e37ba6e4384e076e3fce6f4 - part 2:
52817eb26b5e1997d27e7ed8adfe7421a4acd9f7f6267509e07bf1362cd1e775 - part 3:
18469f3f1bbf6a371db11f62c152ddec77b9d135ad289ed4c6eed9247df0dd37
- part 1:
Swift-1.5-Qwen3.8-Flash-Next-PXQN-48GB
Swift 1.5 Qwen3.8-Flash-Next (UkisAI's reasoning-efficient fine-tune of Qwen3.8-Flash-Next), in PXQN, sized for three 16 GB cards with long context. The 48 GB file is the three-card size; the 64 GB file below is the same model at a higher tier.
- File:
Swift-1.5-Qwen3.8-Flash-Next-PXQN-48GB/Swift-1.5-Qwen3.8-Flash-Next-PXQN-48GB.gguf, 98,839,819,168 bytes. - Cards: three 16 GB cards (48 GB of VRAM), plus about 51 GB of host memory for the built-in n-gram table.
- The file is 98.84 GB but only about 48 GB has to be on the GPU: the n-gram table is served from a read-only file
mapping and stays in host memory. Keep
PXA_PLE_MMAPat its default and keep the-otline below โ the mapping only engages on a tensor the engine has already placed on a host buffer. - Long context: needle recall found at 72,681 and at 115,583 tokens.
- Quality (KLD) and the speed table are being measured and land with the next update.
- sha256:
76f4ff926825672d8ed15c41968138e0293e4d6c71d8229e4238d898b587313c - Command (three cards):
llama-server -m Swift-1.5-Qwen3.8-Flash-Next-PXQN-48GB.gguf -ngl 99 -c 120000 -fa on -sm layer -ts 1,1,1 \
-ot "per_layer_token_embd\.weight=CPU"
Swift-1.5-Qwen3.8-Flash-Next-PXQN-64GB
The same Swift 1.5 Qwen3.8-Flash-Next, at a higher tier and sized for four 16 GB cards. More of the weights are held at the higher PXQN tier, so it is the closer-to-source of the two.
- File:
Swift-1.5-Qwen3.8-Flash-Next-PXQN-64GB/Swift-1.5-Qwen3.8-Flash-Next-PXQN-64GB.gguf, 110,898,443,168 bytes. - Cards: four 16 GB cards (64 GB of VRAM), plus about 51 GB of host memory for the n-gram table.
- Measured on four Tesla P100 16 GB cards, 131072 context,
-sm layer: decode 23.8 t/s at low context fill and 25.0 t/s at ~7.2k fill, prefill 453-466 t/s. Decode holds steady as context grows.-sm tensoris refused on this architecture (hyper-connections), so the split is by layer. - Long context: needle recall found at 72,681 and at 115,583 tokens.
- Quality (KLD) is being measured and lands with the next update.
- sha256:
2a8ef3d7e163553424e02ad4cfb157de374ea79eae4d10289fd671fcb7502a91 - Command (four cards):
llama-server -m Swift-1.5-Qwen3.8-Flash-Next-PXQN-64GB.gguf -ngl 99 -c 131072 -fa on -sm layer -ts 1,1,1,1 \
-ot "per_layer_token_embd\.weight=CPU"
Both Swift 1.5 files need pxa v2026.10.2 or later, and both keep their 51 GB n-gram table off the GPU the same way.
Licence (Flash-Next)
Base model: Qwen/Qwen3.8-Flash-Next, licensed under the Qwen Community
License 1.0 (licence text). It is not Apache-2.0.
This file is quantized from the Uncensored fine-tune
orcarouter/Qwen3.8-Flash-Next-Uncensored. The Qwen
Community License 1.0 applies to this derivative. Its text, with Qwen's copyright notice, is included as
Qwen3.8-Flash-Next-PXQN/LICENSE. Read it before any commercial use: it has conditions for very large products and for
model-as-a-service or AI work-assistant businesses.
Licence (the two Swift 1.5 Qwen3.8-Flash-Next files)
Base model: ukisai/Swift-1.5-Qwen3.8-Flash-Next, UkisAI's
fine-tune of Qwen3.8-Flash-Next. UkisAI's contribution, including the adapted weights, is under the Swift Open License
v1.0 (LICENSE in each folder): personal, research, educational, evaluation and commercial use are free for
individuals and organizations with gross annual revenue, including affiliates, of up to US$1,000,000. Above that
threshold, commercial use needs a separate Swift Enterprise License from UkisAI. The underlying Qwen Community License
1.0 still applies to the derivative and is included as LICENSE-QWEN. See NOTICE and CHANGES.md in each folder.
Licence (the three Qwen3.8-27B files)
Base model: Qwen/Qwen3.8-27B, licensed Apache-2.0. These quantized derivatives are distributed under the same Apache-2.0 licence.
vllm/ holds pre-converted checkpoints (PXQN4, PXQN5) for the PXA vLLM v2026.10 images; see vllm/README.md.
- Downloads last month
- 107
We're not able to determine the quantization variants.