Instructions to use lewismoten/palace-9 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lewismoten/palace-9 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="lewismoten/palace-9")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("lewismoten/palace-9") model = AutoModelForCausalLM.from_pretrained("lewismoten/palace-9", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use lewismoten/palace-9 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf lewismoten/palace-9:Q4_K_M # Run inference directly in the terminal: llama cli -hf lewismoten/palace-9:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf lewismoten/palace-9:Q4_K_M # Run inference directly in the terminal: llama cli -hf lewismoten/palace-9:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf lewismoten/palace-9:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf lewismoten/palace-9:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf lewismoten/palace-9:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf lewismoten/palace-9:Q4_K_M
Use Docker
docker model run hf.co/lewismoten/palace-9:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use lewismoten/palace-9 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lewismoten/palace-9" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lewismoten/palace-9", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/lewismoten/palace-9:Q4_K_M
- SGLang
How to use lewismoten/palace-9 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "lewismoten/palace-9" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lewismoten/palace-9", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "lewismoten/palace-9" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lewismoten/palace-9", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Ollama
How to use lewismoten/palace-9 with Ollama:
ollama run hf.co/lewismoten/palace-9:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use lewismoten/palace-9 with Docker Model Runner:
docker model run hf.co/lewismoten/palace-9:Q4_K_M
- Lemonade
How to use lewismoten/palace-9 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull lewismoten/palace-9:Q4_K_M
Run and chat with the model
lemonade run user.palace-9-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Palace-9 raw-history model
A compact, from-scratch Qwen2MoeForCausalLM model for a 3×3 tic-tac-toe move-history task. It is a raw, one-token state-completion model—not a general conversational model.
Live browser demonstration: lewismoten.github.io/palace-9
Source repository: lewismoten/palace-9
Browser presentation label: PALACE — Predictive Autonomous Learning And Nuclear Contingency Evaluation. The struck word is intentional fictional presentation; the release is a 3×3 game-state model with no operational capability.
Browser weight inspector. The interactive site decodes the selected Palace-9 release artifact locally and visualizes its causal-model forward path. The playable board, map, DEFCON display, trajectories, and cipher are browser presentation layers—not Hugging Face inference or an operational system.
Model contract
Submit only a raw move history made from a through i—one letter for each occupied square, in chronological order. The deployment template is literal <bos>{{ .Prompt }}. Use temperature 0, a 16-token context, and generate exactly one token.
- A legal history returns an optimal unoccupied square.
- A malformed, repeated-square, post-terminal, or otherwise out-of-protocol history returns
!. - This model does not implement a multi-turn chat protocol.
Example raw history: a
Checked F16 fixture result: e
Architecture
- Architecture:
Qwen2MoeForCausalLM, trained from scratch - Vocabulary: custom 261-token byte-level GPT-2-compatible vocabulary
- Context: 16 tokens
- Decoder layers: 1
- Hidden size: 36
- Attention: 9 query heads × 4 dimensions; 3 KV heads × 4 dimensions
- Routed experts: 9; top-2 routing
- Shared expert: 1
Release artifacts
| Artifact | Runtime evidence |
|---|---|
| F16 GGUF | 978,003 raw-history cases; 0 failures; local Ollama fixture verified |
| Q6_K GGUF | 978,003 raw-history cases; 0 failures |
| Q4_K_M GGUF | 978,003 raw-history cases; 0 failures |
| Source checkpoint | Hugging Face-compatible config, tokenizer, and safetensors weights |
The Q4_K_M release is truthfully mixed storage: narrow tensors that cannot use a particular block layout remain F16/F32 or Q6_K where required. No Q8_0 artifact is provided because the model's 36- and 18-wide tensors do not meet Q8_0 block-size requirements.
Use with Ollama
Use the F16 GGUF with a raw-completion template:
FROM ./palace9-qwen2moe-raw-history-f16.gguf
PARAMETER num_ctx 16
PARAMETER num_predict 1
PARAMETER temperature 0
TEMPLATE """<bos>{{ .Prompt }}"""
Then import with ollama create palace-9:f16 -f Modelfile.
The Ollama package contains only the GGUF and Modelfile configuration. The playable board, weight inspector, and fictional visual overlay run independently in browser JavaScript and are available through the source repository and live demo.
Safety and scope
This is a small, fictional 3×3 tic-tac-toe state policy. It has no external command authority, no real-world data, and no operational-system capability. The browser demo's map, DEFCON display, trajectories, and cipher graphics are fictional presentation layers.
Provenance and validation
The release includes source checkpoint files, GGUF artifacts, checksums, a model card, and machine-readable validation evidence. The exhaustive runtime gates cover 294,777 legal histories and 683,226 invalid histories (978,003 total) per F16, Q6_K, and Q4_K_M artifact, with zero failures.
Acknowledgments
PALACE-9 was designed and directed by Lewis Moten. Its code and documentation were developed with assistance from GPT-5.6-terra Med, accessed through Hermes and using Honcho for context and project-memory support. Lewis Moten remains the project designer, maintainer, and publisher.
License
Copyright 2026 Lewis Moten. Released under Apache-2.0.
- Downloads last month
- 494
