nyx / README.md
Devansh Batham
Add required tokenizer assets
47845a3
|
Raw History Blame Contribute Delete
2.74 kB
---
license: apache-2.0
pipeline_tag: text-generation
base_model:
- Qwen/Qwen3.5-4B
tags:
- nyx
- qwen3.5
- gguf
- llama-cpp
- quantized
- classification
---
# nyx
nyx is a compact 4B decision model for Choice, Score, and Noul workloads. The release is one 3.29 GB GGUF for llama.cpp on CPU, Apple Metal, NVIDIA CUDA, and AMD HIP.
![Animated green-and-white chart comparing nyx and Jev across six classification benchmarks](assets/nyx-benchmarks.svg)
## Model
| Property | Value |
|---|---|
| Base | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) |
| File | `nyx.gguf` |
| Size | 3.29 GB / 3.06 GiB |
| Format | GGUF, importance-matrix IQ4_XS with selected Q5_K/Q6_K tensors |
| SHA-256 | `39d64fb19863cfc067e2411fcb2a0444113dda13798b15de5617736d91914359` |
## Run
```bash
git clone https://github.com/devanshbatham/nyx.git /opt/nyx
python3 -m venv /opt/nyx/.venv
/opt/nyx/.venv/bin/pip install '/opt/nyx[server]'
/opt/nyx/.venv/bin/hf download devanshbatham/nyx --local-dir /opt/nyx-model
/opt/nyx/scripts/build-llama-cpp.sh /opt/llama.cpp
export LLAMA_SERVER_BIN=/opt/llama.cpp/build/bin/llama-server
export NYX_GGUF_PATH=/opt/nyx-model/nyx.gguf
/opt/nyx/scripts/serve-model.sh
```
Allow 8 GB of free RAM, VRAM, or unified memory and 6 GB of disk for one 4,096-token slot. Python 3.11+, Git, CMake, and a recent C++ compiler are required. Larger contexts and concurrent slots need more memory.
The production API and typed Python/TypeScript clients are in [devanshbatham/nyx](https://github.com/devanshbatham/nyx).
## Benchmarks
Accuracy on 2,277 frozen classification requests:
| Benchmark | Qwen3.5-4B base | nyx reference | Jev 1.13 |
|---|---:|---:|---:|
| AG News | **87.75%** | 85.75% | 86.50% |
| TREC coarse | 86.60% | 84.20% | **93.00%** |
| CoLA | 77.75% | **81.50%** | 77.75% |
| RTE | 84.12% | 88.09% | **91.34%** |
| Emotion | **47.67%** | 46.33% | 47.00% |
| IMDb | 95.75% | 95.75% | **97.00%** |
| Dataset-macro accuracy | 79.94% | 80.27% | **82.10%** |
| Pooled accuracy | 81.42% | 81.51% | **83.62%** |
| Round trip, median / p95 | 54.85 / 67.96 ms | 55.56 / 72.77 ms | 182.33 / 245.12 ms |
| Throughput | 17.44 req/s | 17.15 req/s | 42.43 req/s |
Local Qwen and nyx latency: MI325X, concurrency 1. Jev latency: remote HTTPS, concurrency 8. The published GGUF matched the nyx reference on 98.73% of a separate 1,024-case fidelity set. [Reproduce the nyx/Jev benchmark](https://github.com/devanshbatham/nyx/tree/main/benchmarks).
## License
Apache-2.0. nyx is derived from [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B); Qwen and Alibaba Cloud are credited as the original model authors. This project is not affiliated with Qwen, Alibaba Cloud, TypeSafe, or Jev.