Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Upload FINDINGS.md with huggingface_hub
Browse files- FINDINGS.md +14 -6
FINDINGS.md
CHANGED
|
@@ -25,12 +25,20 @@ that seems obvious.
|
|
| 25 |
|
| 26 |
Throughput is **invariant** across configurations that differ enormously:
|
| 27 |
|
| 28 |
-
| Configuration | VRAM used | CPU-side bytes/token | decode |
|
| 29 |
-
|---|---|---|---|
|
| 30 |
-
| autofit, 16 experts | 87.7 GB | ~33 GB | 2.46 / 3.00 |
|
| 31 |
-
| `--cpu-moe` + moe-cache | 60.2 GB | ~4 GB | 2.21 / 2.68 |
|
| 32 |
-
| `--n-cpu-moe 86` | 69.9 GB | ~4 GB | 2.51 / 3.06 |
|
| 33 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
|
| 35 |
Cutting host traffic by 8× moved nothing. Halving expert compute moved +28%.
|
| 36 |
|
|
|
|
| 25 |
|
| 26 |
Throughput is **invariant** across configurations that differ enormously:
|
| 27 |
|
| 28 |
+
| Configuration | VRAM used | attention | CPU-side bytes/token | decode |
|
| 29 |
+
|---|---|---|---|---|
|
| 30 |
+
| autofit, 16 experts | 87.7 GB | 48 layers on host | ~33 GB | 2.46 / 3.00 |
|
| 31 |
+
| `--cpu-moe` + moe-cache | 60.2 GB | all in VRAM | ~4 GB | 2.21 / 2.68 |
|
| 32 |
+
| `--n-cpu-moe 86` | 69.9 GB | all in VRAM | ~4 GB | 2.51 / 3.06 |
|
| 33 |
+
| manual balanced `-ot` | 87.8 GB | all in VRAM, 44.3/43.5 split | ~4 GB | 2.31 / 2.98 |
|
| 34 |
+
| autofit, 8 experts | 87.4 GB | 48 layers on host | ~17 GB | 3.15 |
|
| 35 |
+
| `-ngl 0` (no GPU) | 1.5 GB | all on host | all | **0.65** |
|
| 36 |
+
|
| 37 |
+
The manual placement is the cleanest of these: explicit `-ot ...=CUDA0/CUDA1`
|
| 38 |
+
layer ranges fill both cards to 96%/94% with every layer's attention resident
|
| 39 |
+
and only 72 layers' routed experts on the host. It took three calibration
|
| 40 |
+
rounds, each guided by the shortfall rather than a guess — 45256 MiB over, then
|
| 41 |
+
1038 MiB over, then fitting. It performs exactly like everything else.
|
| 42 |
|
| 43 |
Cutting host traffic by 8× moved nothing. Halving expert compute moved +28%.
|
| 44 |
|