How to use from
Pi
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "Vontra/Solar-Open2-250B-MLX-8bit"
Configure the model in Pi
# Install Pi:
npm install -g @mariozechner/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "mlx-lm": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "Vontra/Solar-Open2-250B-MLX-8bit"
        }
      ]
    }
  }
}
Run Pi
# Start Pi in your project directory:
pi
Quick Links

Solar-Open2-250B-MLX-8bit

Built with Solar. This is an MLX 8-bit affine quantization of upstage/Solar-Open2-250B, converted for Apple Silicon / MLX workflows.

Details

  • Source model: upstage/Solar-Open2-250B
  • Quantization: 8-bit affine, group size 64
  • Local size: 248G
  • Weight shards: 62
  • Architecture: Solar Open 2 hybrid-attention MoE, 250B total / ~15B active parameters
  • Context: source model advertises 1M-token context; practical MLX context depends on memory and runtime settings

Important runtime notes

Solar Open2 is not yet a stock mlx-lm architecture in many installs. This repo includes solar_open2.py; launch with --trust-remote-code when serving or loading from Hugging Face.

mlx_lm.server \
  --model Vontra/Solar-Open2-250B-MLX-8bit \
  --host 0.0.0.0 \
  --port 8021 \
  --trust-remote-code \
  --temp 0.2 \
  --top-p 0.9 \
  --max-tokens 32768

You may see a transformers warning that mentions loading model_type=solar_open2 into a blank model type. With the included custom MLX loader this warning is expected; the important check is that the model actually loads.

The tokenizer template uses Solar/Whale-style tool markers such as <|tool_call:start|> and <|tool_arg:start|>. For OpenAI-compatible tool calling, your serving runtime must parse those markers into structured tool_calls. Plain text generation does not need this parser.

Use with MLX

This repo includes a small solar_open2.py MLX loader because upstream mlx-lm does not yet ship native Solar Open 2 support.

pip install -U mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("Vontra/Solar-Open2-250B-MLX-8bit")
prompt = "Write a short Python function that validates an IPv4 CIDR string."
print(generate(model, tokenizer, prompt=prompt, max_tokens=256, verbose=True))

Notes

This is an independent community conversion under the Vontra organization. It is not an official Upstage release.

License

The source model is released under the Upstage Solar License. A copy is included in LICENSE. Please review the upstream model card and license before use or redistribution.

Downloads last month
114
Safetensors
Model size
250B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vontra/Solar-Open2-250B-MLX-8bit

Quantized
(13)
this model