Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Upload vllm_bootstrap.sh with huggingface_hub
Browse files- vllm_bootstrap.sh +13 -2
vllm_bootstrap.sh
CHANGED
|
@@ -16,7 +16,7 @@ PROXY_PORT=8080 # pont Anthropic, seul port expose par le pod
|
|
| 16 |
# Banniere de version : sans elle, impossible de distinguer "la correction n'a
|
| 17 |
# pas marche" de "la correction n'est pas encore arrivee", et on debogue le
|
| 18 |
# mauvais probleme.
|
| 19 |
-
SCRIPT_VERSION="
|
| 20 |
log() { echo "[VL] $(date -u +%H:%M:%S) $*"; }
|
| 21 |
|
| 22 |
fatal() {
|
|
@@ -145,14 +145,25 @@ fi
|
|
| 145 |
# vLLM rabaisse max_num_batched_tokens a 2048 quand la speculation est active et
|
| 146 |
# previent lui-meme que c'est sous-optimal : les brouillons devorent le budget de
|
| 147 |
# batch. Mesure sans le corriger : 32 flux tombaient de 1166 a 296 tok/s.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 148 |
COMMON=(
|
| 149 |
"$MODEL"
|
| 150 |
--served-model-name "$NAME"
|
| 151 |
--host 0.0.0.0 --port "$PORT"
|
| 152 |
--gpu-memory-utilization "${VL_UTIL:-0.90}"
|
| 153 |
--max-num-batched-tokens "${VL_BATCHED:-16384}"
|
| 154 |
-
--max-num-seqs "${VL_SEQS:-
|
| 155 |
--enable-prefix-caching
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 156 |
--enable-auto-tool-choice
|
| 157 |
"${EXTRA[@]}"
|
| 158 |
"${SPEC[@]}"
|
|
|
|
| 16 |
# Banniere de version : sans elle, impossible de distinguer "la correction n'a
|
| 17 |
# pas marche" de "la correction n'est pas encore arrivee", et on debogue le
|
| 18 |
# mauvais probleme.
|
| 19 |
+
SCRIPT_VERSION="v8-async"
|
| 20 |
log() { echo "[VL] $(date -u +%H:%M:%S) $*"; }
|
| 21 |
|
| 22 |
fatal() {
|
|
|
|
| 145 |
# vLLM rabaisse max_num_batched_tokens a 2048 quand la speculation est active et
|
| 146 |
# previent lui-meme que c'est sous-optimal : les brouillons devorent le budget de
|
| 147 |
# batch. Mesure sans le corriger : 32 flux tombaient de 1166 a 296 tok/s.
|
| 148 |
+
# L'ordonnancement asynchrone est incompatible avec la speculation : on ne
|
| 149 |
+
# l'active donc que lorsqu'elle est coupee.
|
| 150 |
+
ASYNC_FLAG=""
|
| 151 |
+
[ "${#SPEC[@]}" -eq 0 ] && [ "${VL_ASYNC:-on}" = "on" ] && ASYNC_FLAG="--async-scheduling"
|
| 152 |
+
[ -n "$ASYNC_FLAG" ] && log "ordonnancement asynchrone active"
|
| 153 |
+
|
| 154 |
COMMON=(
|
| 155 |
"$MODEL"
|
| 156 |
--served-model-name "$NAME"
|
| 157 |
--host 0.0.0.0 --port "$PORT"
|
| 158 |
--gpu-memory-utilization "${VL_UTIL:-0.90}"
|
| 159 |
--max-num-batched-tokens "${VL_BATCHED:-16384}"
|
| 160 |
+
--max-num-seqs "${VL_SEQS:-64}"
|
| 161 |
--enable-prefix-caching
|
| 162 |
+
# vLLM refusait l'ordonnancement asynchrone tant que la speculation ngram
|
| 163 |
+
# etait active ("Async scheduling not supported with ngram-based speculative
|
| 164 |
+
# decoding and will be disabled") : il recouvre l'ordonnancement CPU avec le
|
| 165 |
+
# calcul GPU, donc il ne rapporte que si on ne le desactive pas soi-meme.
|
| 166 |
+
${ASYNC_FLAG}
|
| 167 |
--enable-auto-tool-choice
|
| 168 |
"${EXTRA[@]}"
|
| 169 |
"${SPEC[@]}"
|