Shingi 27B — 審議

Shingi — 審議

Shingi 27B is a local decision model. You give it context (text, and optionally images), a question and the possible answers. It returns a choice, calibrated probabilities or an ordinal score. It reads the logit of every candidate answer directly and generates no text.

The model is a single 7.2 GB ternary GGUF file (PQ2_0). It is Bonsai 2 27B with retrained block scales: the ternary weights are unchanged, only their scales were trained. It needs no adapter.

Run it

On Linux with an NVIDIA GPU and the CUDA toolkit, or on a Mac with Apple Silicon:

curl -fsSL https://raw.githubusercontent.com/kortexa-ai/shingi-27b/main/run.sh | bash

This builds the pinned Prism runtime, downloads this model into your Hugging Face cache and serves the API on http://127.0.0.1:8765. Use --port to change the port. The code, API examples and options are in kortexa-ai/shingi-27b.

Requirements

Resource Need
GPU NVIDIA with at least 20 GB: RTX 4090, RTX PRO 6000 or DGX Spark (tested). The model uses about 8 GiB at the 16K context, about 9 GiB with image input.
System RAM about 1 GB for the server process, plus page cache for the 7.2 GB file
Disk 7.2 GB for the weights, plus the runtime build
Mac Apple Silicon with 24 GB or more of unified memory (16 GB minimum), using Metal
Software Linux (x86-64 or aarch64) with the CUDA toolkit, or macOS with the Xcode command line tools; CMake and a C++17 compiler

Speed

Median latency per decision, one request at a time, loading excluded:

Request RTX PRO 6000 RTX 4090 DGX Spark
Short yes/no (~120 tokens) 66 ms 85 ms 148 ms
Short choice, 3 options 73 ms 95 ms 181 ms
Choice, 10 options (~1,300 tokens) 246 ms 306 ms 802 ms
Long choice, 3 options (~3,000 tokens) 881 ms 1,045 ms 2,984 ms
GPU memory in use 8.3 GB 8.2 GB 7.9 GB

With image input on the RTX 4090, one image and one question take about 0.66 s (median) and the model uses about 9.2 GB.

Scores

External suites, canonical choice order, 16K context:

Suite Items Accuracy
JevBench public 231 84.8%
JevBench hard 111 70.3%
DecisionBench 586 72.7%
DecisionBench hard 293 62.8%
This/That 7,305 67.6%

These suites also informed the choice of training data, so treat them as development results rather than a clean held-out test.

Images

Shingi reads images through the Bonsai 2 27B vision projector (mmproj.gguf, included here). Send up to 8 images per request: /v1/decisions follows SGLang's decision endpoint, and /v1/systemone takes an optional images list.

Zero-shot, on an RTX 4090:

Suite Items Accuracy
VSR (spatial yes/no) 300 79.3%
A-OKVQA (4-way) 300 87.7%
VQAv2 yes/no 300 88.7%

The model was not trained on images; these come from the base model's vision with Shingi's decision training on top.

Training data

Public datasets, used under their stated licenses: Banking77, CLINC150, MMLU, HelpSteer2, HelpSteer, Measuring Hate Speech, Civil Comments, GoEmotions, LEDGAR (LexGLUE), MASSIVE, CommonsenseQA, WinoGrande and GSM8K. Several of them came through the jev-bench repackaging.

About a third of the training tokens come from a private synthetic dataset of decision tasks.

Limitations

  • English only.
  • Much slower on a Mac than on an NVIDIA GPU: on an M4 Pro, a short decision takes about 1.2 s and a 1,100-token one about 12.6 s. MLX was measured alongside Metal and was about 10% slower on the M4, so the Mac build uses Metal.

License

Apache-2.0. Derived from Bonsai 2 27B by Prism ML (Apache-2.0), which descends from Qwen3.8-27B; see NOTICE. The training datasets keep their own licenses.

Downloads last month
540
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kortexa-ai/shingi-27b

Base model

Qwen/Qwen3.8-27B
Finetuned
(5)
this model

Datasets used to train kortexa-ai/shingi-27b