Instructions to use juspay/Xor-26B-A4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use juspay/Xor-26B-A4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="juspay/Xor-26B-A4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("juspay/Xor-26B-A4B") model = AutoModelForMultimodalLM.from_pretrained("juspay/Xor-26B-A4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use juspay/Xor-26B-A4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "juspay/Xor-26B-A4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juspay/Xor-26B-A4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/juspay/Xor-26B-A4B
- SGLang
How to use juspay/Xor-26B-A4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "juspay/Xor-26B-A4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juspay/Xor-26B-A4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "juspay/Xor-26B-A4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juspay/Xor-26B-A4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use juspay/Xor-26B-A4B with Docker Model Runner:
docker model run hf.co/juspay/Xor-26B-A4B
Xor 26B-A4B 1.0
Xor 26B-A4B 1.0 (xor-26b-a4b-1.0) is a post-trained build of google/gemma-4-26B-A4B-it for typed decision tasks, released as merged BF16 weights. It is served through a TypeSafe-compatible /v1/systemone API that answers noul, choice and score questions with full probability distributions.
A 4-bit build of the same weights, juspay/Xor-26B-A4B-NVFP4, is about one third of the size and scores within one question of this model on JEVBench (see Evaluation).
At a glance
- Base:
google/gemma-4-26B-A4B-it(mixture of experts: 30 layers, 128 routed experts with 8 active per token, about 4B active of 26B parameters), revision4d7ae4984b7db7de8f8457170b3f1a419ee76d52 - Weights: BF16 throughout, 51.6 GB
- Runtime: SGLang image
prakhar1611/xor-sglang@sha256:94c48d2a…(SGLang with a logprob fix), tensor parallelism 1, one replica per GPU - Validated on 2 x NVIDIA RTX PRO 6000 Blackwell (96 GB), data parallelism 2
- Serving: the bundled compatibility server (
xn_base_server.py) with achoicetemperature of 1.95; it is required for reproducible results - Public JEVBench (self-run): 203 of 231
Revisions
| Revision | Release |
|---|---|
main |
Latest stable release (currently Xor 26B-A4B 1.0) |
v1 |
Xor 26B-A4B 1.0 (xor-26b-a4b-1.0), released 2026-10-01, immutable |
Pin a revision for reproducible results:
hf download juspay/Xor-26B-A4B --revision v1 --local-dir xor-26b-a4b-1.0
Composition
Two post-training lineages were trained separately on the same base, each merged into BF16, then weight-averaged and pulled back towards the base:
| Lineage | Modules adapted | LoRA rank / alpha | Merge scale |
|---|---|---|---|
| A | attention q/k/v/o_proj and dense MLP gate/up/down_proj, all 30 layers |
32 / 64 | 1.0 |
| B | the same tables plus router.proj |
32 / 64 | 0.75 (lora_B x 0.75) |
| B, routed experts | all 128 experts of all 30 layers (experts.gate_up_proj, experts.down_proj); tiered rank by routing frequency: 16 for the 32 most-routed experts of a layer, 8 for the next 32, 4 for the remaining 64 |
16 / 8 / 4 | 1.0 |
W_A = merge(W_base, lineage A)
W_B = merge(W_base, 0.75 x lineage B tables), then the expert deltas folded in place
W_soup = 0.5 x W_A + 0.5 x W_B
W = 0.85 x W_soup + 0.15 x W_base
Each step is stored in BF16; the two weighted averages are computed in FP32. In total W ≈ W_base + 0.425 ΔA + 0.425 (0.75 ΔB_tables + ΔB_experts). Changed tensors: 235 attention, dense-MLP and router tables and the 60 grouped expert tensors. Embeddings, norms, layer scalars and the vision tower are the base's.
recipe/ contains the merge scripts; the adapters themselves are not published. RELEASE_PROVENANCE.json records the hashes of every input.
Interface
noul: binary probabilitychoice: categorical decision and full probability distribution (2 to 255 candidates)score: expected ordinal score and full probability distribution (2 to 10 levels)
Each question is scored as an independent prompt in the model's chat template (thinking disabled): the state, the question and the lettered options, ending in Answer: (. The answer distribution is a softmax over the logprobs of the option-label tokens at that position, at temperature 1.95 for choice and 1.0 for noul and score. Every answer also carries x_lp, the raw label logprobs in option order.
Requests are text only: a request with images is rejected with 400. A request whose prompt exceeds the context length (131,072 tokens) is rejected with 422.
Quick start
Download the release, verify and extract the serving bundle, and start Xor 26B-A4B on two GPUs:
hf download juspay/Xor-26B-A4B --revision v1 --local-dir xor-26b-a4b-1.0
(cd xor-26b-a4b-1.0/serving && sha256sum -c xor-26b-a4b-1.0-serving.tar.gz.sha256)
mkdir -p xor-26b-a4b-1.0-runtime
tar -xzf xor-26b-a4b-1.0/serving/xor-26b-a4b-1.0-serving.tar.gz -C xor-26b-a4b-1.0-runtime --strip-components=1
cd xor-26b-a4b-1.0-runtime
cp .env.example .env
sed -i "s|^MODEL_DIR=.*|MODEL_DIR=$(cd ../xor-26b-a4b-1.0 && pwd)|" .env
./run.sh
run.sh verifies every model file against checksums.sha256 before starting. When the smoke test succeeds, the API is available at http://127.0.0.1:30002/v1/systemone.
curl -sS -X POST http://127.0.0.1:30002/v1/systemone \
-H 'Content-Type: application/json' \
--data @examples/request.json
.env.example uses GPUs 0,1 with one replica each. The model needs about 50 GB per replica, so one GPU with enough memory also works: set CUDA_VISIBLE_DEVICES=0 and DP_SIZE=1 (not separately benchmarked). The setup requires Linux x86-64, the Hugging Face CLI, Docker Engine with Docker Compose v2, the NVIDIA Container Toolkit, and about 60 GB of free disk space.
Manual start (without Docker Compose)
IMG=prakhar1611/xor-sglang@sha256:94c48d2a6cc98dc456cf93f723707ea7dd81dddfe1061e823b348d68bbe8158f
docker run -d --name xor-26b-a4b-1.0-sglang --gpus '"device=0,1"' --network host --ipc host -e HF_HUB_OFFLINE=1 \
-v "$PWD/xor-26b-a4b-1.0:/models/xor:ro" $IMG \
python3 -m sglang.launch_server --model-path /models/xor --trust-remote-code \
--tp-size 1 --dp-size 2 --port 30000 --host 127.0.0.1 \
--context-length 131072 --max-prefill-tokens 131072 --chunked-prefill-size 16384 \
--mem-fraction-static 0.80 --load-balance-method round_robin
docker run -d --name xor-26b-a4b-1.0-api --network host -e HF_HUB_OFFLINE=1 -e HOME=/tmp \
-e XN_SGLANG_URL=http://127.0.0.1:30000 -e XN_PORT=30002 -e XN_MODEL_ID=xor-26b-a4b-1.0 \
-e XN_MODEL_PATH=/models/xor -e XN_TEMP_JSON='{"choice": 1.95, "noul": 1.0, "score": 1.0}' \
-v "$PWD/xor-26b-a4b-1.0:/models/xor:ro" -v "$PWD/xor-26b-a4b-1.0-runtime/xn_base_server.py:/app/xn_base_server.py:ro" \
--entrypoint python3 $IMG -u /app/xn_base_server.py
The compatibility server needs only the Python standard library and transformers (for the tokenizer and chat template); it runs in the same image as SGLang so that both versions match the evaluated setup.
Evaluation
Both builds were evaluated with identical setups on the same machine: 2 x NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), the SGLang image above, data parallelism 2, the bundled compatibility server, one request at a time, and the KV cache flushed before every run. Each figure is the mean of three runs.
Public JEVBench self-run
JEVBench public tiers, harness 1bcc55e, typesafe adapter.
| Tier | Items | Xor 26B-A4B correct | Xor 26B-A4B-NVFP4 correct | Xor 26B-A4B p50 / p95 | Xor 26B-A4B-NVFP4 p50 / p95 | Xor 26B-A4B ECE | Xor 26B-A4B-NVFP4 ECE |
|---|---|---|---|---|---|---|---|
| Easy | 48 | 48 | 48 | 0.047 / 0.059 s | 0.047 / 0.058 s | 0.044 | 0.046 |
| Original | 72 | 70 | 69.7 | 0.047 / 0.054 s | 0.047 / 0.048 s | 0.083 | 0.071 |
| Hard public | 111 | 85 | 86.7 | 0.071 / 0.170 s | 0.049 / 0.139 s | 0.074 | 0.090 |
| All public | 231 | 203 | 204.3 | 0.054 / 0.145 s | 0.047 / 0.114 s | 0.062 | 0.065 |
Every request returned a valid answer in every run. Xor 26B-A4B answered 203 of 231 correctly in each of the three runs; the NVFP4 build answered 204, 204 and 205. Brier score over all 231 items: 0.174 (BF16) and 0.184 (NVFP4).
SGLang reuses cached prompt prefixes across requests, and a warm cache changes the computed probabilities slightly. The published runs flush the cache before each run; with that, this release reproduces them exactly (checked through the serving bundle). Without flushing, after other traffic, individual answers can flip: in one bundle test the BF16 build answered 2 of 231 items differently, with the same total.
These are self-run results, not an official JEVBench rank. Latency includes the compatibility server, is hardware-specific and was measured locally without network overhead.
Operational notes
- Model files occupy 51.6 GB; each data-parallel worker loads a complete replica (about 49 GB of GPU memory before the KV cache).
- Alternative hardware and parallelism settings must be validated independently before publishing performance results.
checksums.sha256covers every model file;RELEASE_PROVENANCE.jsonrecords the base revision, the input hashes, the merge, and the runtime.- The compatibility server has no authentication. Remote deployments must add authentication, TLS, rate limits, and request-size limits at the ingress layer.
License
Apache License 2.0. Derived from Gemma 4 26B-A4B-it (Apache 2.0). See THIRD_PARTY_NOTICES.md.
- Downloads last month
- 423