Instructions to use juspay/xor-lite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use juspay/xor-lite with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="juspay/xor-lite") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("juspay/xor-lite") model = AutoModelForCausalLM.from_pretrained("juspay/xor-lite", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use juspay/xor-lite with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "juspay/xor-lite" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juspay/xor-lite", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/juspay/xor-lite
- SGLang
How to use juspay/xor-lite with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "juspay/xor-lite" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juspay/xor-lite", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "juspay/xor-lite" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juspay/xor-lite", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use juspay/xor-lite with Docker Model Runner:
docker model run hf.co/juspay/xor-lite
XOR Lite
XOR Lite (xor-lite) is a post-trained version of openbmb/MiniCPM5-2B for typed decision tasks. It is served through a TypeSafe-compatible /v1/systemone API
Release
| Field | Value |
|---|---|
| Base checkpoint | openbmb/MiniCPM5-2B |
| Base revision | 12a3808a956f869c767195e9266b59c4d21d92e2 |
| Architecture | Dense LlamaForCausalLM |
| Parameters | Approximately 2.52 billion |
| Released precision | BF16 |
| License | Apache License 2.0 |
| Packaging | Fully merged weights; no adapter loading is required |
Pin the published revision for reproducible use
hf download juspay/xor-lite --revision REPLACE_WITH_RELEASE_REVISION --local-dir xor-lite
Interface
The server accepts a state and a map of typed questions
noul: binary probabilitychoice: categorical decision and full probability distributionscore: expected ordinal score and full probability distribution
The serving layer performs constrained single-token candidate readout, forward and reverse option-order evaluation, probability calibration, and SystemOne schema conversion. It supports up to 255 candidates using the included MiniCPM-specific marker table. The released serving layer is part of the inference configuration required for reproducible results
Quick start
hf download juspay/xor-lite --revision REPLACE_WITH_RELEASE_REVISION --local-dir xor-lite
(cd xor-lite/serving && sha256sum -c xor-lite-serving.tar.gz.sha256)
mkdir -p xor-lite-runtime
tar -xzf xor-lite/serving/xor-lite-serving.tar.gz -C xor-lite-runtime --strip-components=1
cd xor-lite-runtime
cp .env.example .env
sed -i "s|^MODEL_DIR=.*|MODEL_DIR=$(cd ../xor-lite && pwd)|" .env
./run.sh
When the smoke test succeeds, the API is available at http://127.0.0.1:30002/v1/systemone
Public JEVBench self-run
XOR Lite was evaluated locally on all 231 public JEVBench decisions using harness commit fd54ea7dc02bbe29c6ac8f6e015a54cdcff26805, the existing typesafe adapter, and one request at a time
The run used the released configuration on 1 x NVIDIA H200 NVL (143 GB), tensor parallelism 1, data parallelism 1, with request caching disabled
| Tier | Attempted | Valid | Correct | Accuracy | Brier | ECE | p50 | p95 |
|---|---|---|---|---|---|---|---|---|
| Easy | 48 | 48 | 48 | 1.0000 | 0.1467 | 0.2294 | 0.0250 s | 0.0279 s |
| Original | 72 | 72 | 66 | 0.9167 | 0.2777 | 0.2986 | 0.0251 s | 0.0291 s |
| Hard public | 111 | 111 | 55 | 0.4955 | 0.6153 | 0.0902 | 0.0308 s | 0.0659 s |
| All public | 231 | 231 | 169 | 0.7316 |
Operational success, coverage, schema validity, and strict schema validity were 1.0000 for each tier. These are self-run public-tier results, not an official JEVBench rank. Latency is hardware-specific and was measured locally without network overhead
Validated runtime
| Setting | Value |
|---|---|
| SGLang image | lmsysorg/sglang@sha256:6bcaa47db52f78ce0d67863b8b2431221b79bc23204a80cad757fa819d00e921 |
| Tensor parallelism | 1 |
| Data parallelism | 1 |
| Validated GPU | 1 x NVIDIA H200 NVL, 143 GB |
| Maximum prefill tokens | 250,000 |
| Static memory fraction | 0.50 |
| Probability temperature | 100.0 for noul, choice, and score |
NVIDIA L4 validation
The immutable v1.0 release was also validated without modification on one NVIDIA L4 with 24 GB of memory using tensor parallelism 1 and data parallelism 1. The complete public JEVBench run reproduced the H200 result exactly at 169/231 correct, with no operational, coverage, or schema failures
| Tier | Correct | Accuracy | p50 | p95 |
|---|---|---|---|---|
| Easy | 48/48 | 1.0000 | 0.0458 s | 0.0481 s |
| Original | 66/72 | 0.9167 | 0.0465 s | 0.0488 s |
| Hard public | 55/111 | 0.4955 | 0.1175 s | 0.5868 s |
| All public | 169/231 | 0.7316 |
The loaded service used 13,803 MiB of the L4's 23,034 MiB. The run used NVIDIA driver 595.91.07 and the same pinned SGLang image shown above. The evidence archive has SHA-256 6389b280c9abfffd608a3864a1b8ed7570974ca415ad5f079d0224f429694ed9
The L4 result demonstrates that the H200 used for the original release measurement is not a minimum hardware requirement. It does not define or announce a hosted-service tariff
The service binds to localhost by default and does not require an API key. Remote deployments must add authentication, TLS, rate limits, and request-size limits at the ingress layer
- Downloads last month
- 464
Model tree for juspay/xor-lite
Base model
openbmb/MiniCPM5-2B