Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Upload k3_bootstrap.sh with huggingface_hub
Browse files- k3_bootstrap.sh +12 -2
k3_bootstrap.sh
CHANGED
|
@@ -525,8 +525,18 @@ fi
|
|
| 525 |
CPU_EXPERTS='blk\.([1-9]|[12][0-9]|3[0-6]|4[7-9]|[5-7][0-9]|8[0-2])\.ffn_.*_exps'
|
| 526 |
GPU0_LAYERS='blk\.([0-9]|[1-3][0-9]|4[0-6])\.'
|
| 527 |
GPU1_LAYERS='blk\.(4[7-9]|[5-8][0-9]|9[0-2])\.'
|
| 528 |
-
|
| 529 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 530 |
|
| 531 |
# Only 86 is kept: 62, 68, 74 and 80 were each measured OOM-ing on device 1,
|
| 532 |
# and re-proving that costs a minute of paid GPU apiece on every boot.
|
|
|
|
| 525 |
CPU_EXPERTS='blk\.([1-9]|[12][0-9]|3[0-6]|4[7-9]|[5-7][0-9]|8[0-2])\.ffn_.*_exps'
|
| 526 |
GPU0_LAYERS='blk\.([0-9]|[1-3][0-9]|4[0-6])\.'
|
| 527 |
GPU1_LAYERS='blk\.(4[7-9]|[5-8][0-9]|9[0-2])\.'
|
| 528 |
+
# Two-GPU only: the regexes name CUDA0/CUDA1 explicitly and push 72 layers'
|
| 529 |
+
# experts to the host. On a machine with enough VRAM to hold everything, that
|
| 530 |
+
# would deliberately recreate the bottleneck we are trying to remove.
|
| 531 |
+
if [ "$NGPU" = 2 ]; then
|
| 532 |
+
try "manual balanced placement" "${COMMON[@]}" -ngl 99 --tensor-split 1,1 \
|
| 533 |
+
-ot "$CPU_EXPERTS=CPU" -ot "$GPU0_LAYERS=CUDA0" -ot "$GPU1_LAYERS=CUDA1"
|
| 534 |
+
else
|
| 535 |
+
log "skipping manual balanced placement (built for 2 GPUs, this host has $NGPU)"
|
| 536 |
+
# Everything resident, nothing on the host. 194 GB of weights against
|
| 537 |
+
# NGPU x 45 GB; the ggml microbenchmark puts VRAM-resident MoE at 5.5 ms/token.
|
| 538 |
+
try "all-VRAM (-ngl 99, no expert offload)" "${COMMON[@]}" -ngl 99
|
| 539 |
+
fi
|
| 540 |
|
| 541 |
# Only 86 is kept: 62, 68, 74 and 80 were each measured OOM-ing on device 1,
|
| 542 |
# and re-proving that costs a minute of paid GPU apiece on every boot.
|