Instructions to use kortexa-ai/shingi-27b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use kortexa-ai/shingi-27b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf kortexa-ai/shingi-27b # Run inference directly in the terminal: llama cli -hf kortexa-ai/shingi-27b
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf kortexa-ai/shingi-27b # Run inference directly in the terminal: llama cli -hf kortexa-ai/shingi-27b
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf kortexa-ai/shingi-27b # Run inference directly in the terminal: ./llama-cli -hf kortexa-ai/shingi-27b
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf kortexa-ai/shingi-27b # Run inference directly in the terminal: ./build/bin/llama-cli -hf kortexa-ai/shingi-27b
Use Docker
docker model run hf.co/kortexa-ai/shingi-27b
- LM Studio
- Jan
- Ollama
How to use kortexa-ai/shingi-27b with Ollama:
ollama run hf.co/kortexa-ai/shingi-27b
- Unsloth Desktop
- Pi
How to use kortexa-ai/shingi-27b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kortexa-ai/shingi-27b
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "kortexa-ai/shingi-27b" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use kortexa-ai/shingi-27b with Docker Model Runner:
docker model run hf.co/kortexa-ai/shingi-27b
- Lemonade
How to use kortexa-ai/shingi-27b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull kortexa-ai/shingi-27b
Run and chat with the model
lemonade run user.shingi-27b-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use kortexa-ai/shingi-27b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kortexa-ai/shingi-27b
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default kortexa-ai/shingi-27b
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use kortexa-ai/shingi-27b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kortexa-ai/shingi-27b
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "kortexa-ai/shingi-27b" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Shingi 27B — 審議
Shingi 27B is a local decision model. You give it context (text, and optionally images), a question and the possible answers. It returns a choice, calibrated probabilities or an ordinal score. It reads the logit of every candidate answer directly and generates no text.
The model is a single 7.2 GB ternary GGUF file (PQ2_0). It is Bonsai 2 27B with retrained block scales: the ternary weights are unchanged, only their scales were trained. It needs no adapter.
Run it
On Linux with an NVIDIA GPU and the CUDA toolkit, or on a Mac with Apple Silicon:
curl -fsSL https://raw.githubusercontent.com/kortexa-ai/shingi-27b/main/run.sh | bash
This builds the pinned Prism runtime,
downloads this model into your Hugging Face cache and serves the API on
http://127.0.0.1:8765. Use --port to change the port. The code, API examples
and options are in kortexa-ai/shingi-27b.
Requirements
| Resource | Need |
|---|---|
| GPU | NVIDIA with at least 20 GB: RTX 4090, RTX PRO 6000 or DGX Spark (tested). The model uses about 8 GiB at the 16K context, about 9 GiB with image input. |
| System RAM | about 1 GB for the server process, plus page cache for the 7.2 GB file |
| Disk | 7.2 GB for the weights, plus the runtime build |
| Mac | Apple Silicon with 24 GB or more of unified memory (16 GB minimum), using Metal |
| Software | Linux (x86-64 or aarch64) with the CUDA toolkit, or macOS with the Xcode command line tools; CMake and a C++17 compiler |
Speed
Median latency per decision, one request at a time, loading excluded:
| Request | RTX PRO 6000 | RTX 4090 | DGX Spark |
|---|---|---|---|
| Short yes/no (~120 tokens) | 66 ms | 85 ms | 148 ms |
| Short choice, 3 options | 73 ms | 95 ms | 181 ms |
| Choice, 10 options (~1,300 tokens) | 246 ms | 306 ms | 802 ms |
| Long choice, 3 options (~3,000 tokens) | 881 ms | 1,045 ms | 2,984 ms |
| GPU memory in use | 8.3 GB | 8.2 GB | 7.9 GB |
With image input on the RTX 4090, one image and one question take about 0.66 s (median) and the model uses about 9.2 GB.
Scores
External suites, canonical choice order, 16K context:
| Suite | Items | Accuracy |
|---|---|---|
| JevBench public | 231 | 84.8% |
| JevBench hard | 111 | 70.3% |
| DecisionBench | 586 | 72.7% |
| DecisionBench hard | 293 | 62.8% |
| This/That | 7,305 | 67.6% |
These suites also informed the choice of training data, so treat them as development results rather than a clean held-out test.
Images
Shingi reads images through the Bonsai 2 27B vision projector (mmproj.gguf, included
here). Send up to 8 images per request: /v1/decisions follows SGLang's decision
endpoint, and /v1/systemone takes an optional images list.
Zero-shot, on an RTX 4090:
| Suite | Items | Accuracy |
|---|---|---|
| VSR (spatial yes/no) | 300 | 79.3% |
| A-OKVQA (4-way) | 300 | 87.7% |
| VQAv2 yes/no | 300 | 88.7% |
The model was not trained on images; these come from the base model's vision with Shingi's decision training on top.
Training data
Public datasets, used under their stated licenses: Banking77, CLINC150, MMLU, HelpSteer2, HelpSteer, Measuring Hate Speech, Civil Comments, GoEmotions, LEDGAR (LexGLUE), MASSIVE, CommonsenseQA, WinoGrande and GSM8K. Several of them came through the jev-bench repackaging.
About a third of the training tokens come from a private synthetic dataset of decision tasks.
Limitations
- English only.
- Much slower on a Mac than on an NVIDIA GPU: on an M4 Pro, a short decision takes about 1.2 s and a 1,100-token one about 12.6 s. MLX was measured alongside Metal and was about 10% slower on the M4, so the Mac build uses Metal.
License
Apache-2.0. Derived from Bonsai 2 27B by Prism ML (Apache-2.0), which descends from Qwen3.8-27B; see NOTICE. The training datasets keep their own licenses.
- Downloads last month
- 540
We're not able to determine the quantization variants.
