Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
journal de demarrage
Browse files- etat/grefggayunyaj9.log +51 -1
etat/grefggayunyaj9.log
CHANGED
|
@@ -1,4 +1,4 @@
|
|
| 1 |
-
=== bootstrap v84-nom-nemotron | 20:49:
|
| 2 |
[VL] 20:39:20 bootstrap v84-nom-nemotron
|
| 3 |
[VL] 20:39:20 pilote 595.71.05, CUDA runtime 13.2
|
| 4 |
[VL] 20:39:20 pilote 595.71.05 : compat CUDA non necessaire
|
|
@@ -20,6 +20,7 @@
|
|
| 20 |
[VL] 20:42:36 cache de compilation, cle db5e9bed4beefa32 (NVIDIA_RTX_PRO_5000_Blackwell, vllm 0.27.1)
|
| 21 |
[VL] 20:42:40 cache de compilation RESTAURE (demarrage court attendu)
|
| 22 |
[VL] 20:42:40 journal distant : https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/etat/grefggayunyaj9.log
|
|
|
|
| 23 |
|
| 24 |
=== nvidia-smi ===
|
| 25 |
41365 MiB, 48935 MiB
|
|
@@ -232,5 +233,54 @@
|
|
| 232 |
(APIServer pid=5739) INFO 08-29 20:48:37 [api_server.py:678] Supported tasks: ['generate']
|
| 233 |
(APIServer pid=5739) INFO 08-29 20:48:37 [parser_manager.py:37] "auto" tool choice has been enabled.
|
| 234 |
(APIServer pid=5739) INFO 08-29 20:48:44 [hf.py:540] Detected the chat template content format to be 'openai'. You can set `--chat-template-content-format` to override this.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 235 |
|
| 236 |
=== proxy.log (fin) ===
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
=== bootstrap v84-nom-nemotron | 20:49:47 UTC ===
|
| 2 |
[VL] 20:39:20 bootstrap v84-nom-nemotron
|
| 3 |
[VL] 20:39:20 pilote 595.71.05, CUDA runtime 13.2
|
| 4 |
[VL] 20:39:20 pilote 595.71.05 : compat CUDA non necessaire
|
|
|
|
| 20 |
[VL] 20:42:36 cache de compilation, cle db5e9bed4beefa32 (NVIDIA_RTX_PRO_5000_Blackwell, vllm 0.27.1)
|
| 21 |
[VL] 20:42:40 cache de compilation RESTAURE (demarrage court attendu)
|
| 22 |
[VL] 20:42:40 journal distant : https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/etat/grefggayunyaj9.log
|
| 23 |
+
[VL] 20:49:20 READY qwen38nvfp4
|
| 24 |
|
| 25 |
=== nvidia-smi ===
|
| 26 |
41365 MiB, 48935 MiB
|
|
|
|
| 233 |
(APIServer pid=5739) INFO 08-29 20:48:37 [api_server.py:678] Supported tasks: ['generate']
|
| 234 |
(APIServer pid=5739) INFO 08-29 20:48:37 [parser_manager.py:37] "auto" tool choice has been enabled.
|
| 235 |
(APIServer pid=5739) INFO 08-29 20:48:44 [hf.py:540] Detected the chat template content format to be 'openai'. You can set `--chat-template-content-format` to override this.
|
| 236 |
+
(APIServer pid=5739) INFO 08-29 20:49:13 [base.py:235] Multi-modal warmup completed in 29.053s
|
| 237 |
+
(APIServer pid=5739) INFO 08-29 20:49:14 [base.py:235] Readonly multi-modal warmup completed in 0.692s
|
| 238 |
+
(APIServer pid=5739) WARNING 08-29 20:49:14 [model.py:1637] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 1.0, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
| 239 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [api_server.py:682] Starting vLLM server on http://0.0.0.0:18081
|
| 240 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:37] Available routes are:
|
| 241 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 242 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 243 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 244 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 245 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /load, Methods: GET
|
| 246 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /version, Methods: GET
|
| 247 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /health, Methods: GET
|
| 248 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /metrics, Methods: GET
|
| 249 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 250 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 251 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 252 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /ping, Methods: GET
|
| 253 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /ping, Methods: POST
|
| 254 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /invocations, Methods: POST
|
| 255 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 256 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 257 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 258 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 259 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 260 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 261 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 262 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 263 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 264 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 265 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 266 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 267 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 268 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 269 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 270 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 271 |
+
(APIServer pid=5739) INFO 08-29 20:49:15 [launcher.py:99] API server: waiting for HTTP server to start
|
| 272 |
+
(APIServer pid=5739) INFO: Started server process [5739]
|
| 273 |
+
(APIServer pid=5739) INFO: Waiting for application startup.
|
| 274 |
+
(APIServer pid=5739) INFO: Application startup complete.
|
| 275 |
+
(APIServer pid=5739) INFO 08-29 20:49:16 [launcher.py:105] API server: HTTP server started
|
| 276 |
+
(APIServer pid=5739) INFO: 127.0.0.1:60916 - "GET /health HTTP/1.1" 200 OK
|
| 277 |
+
(APIServer pid=5739) INFO: 127.0.0.1:60932 - "GET /health HTTP/1.1" 200 OK
|
| 278 |
+
(APIServer pid=5739) INFO: 127.0.0.1:57534 - "GET /health HTTP/1.1" 200 OK
|
| 279 |
|
| 280 |
=== proxy.log (fin) ===
|
| 281 |
+
INFO: Started server process [7200]
|
| 282 |
+
INFO: Waiting for application startup.
|
| 283 |
+
INFO: Application startup complete.
|
| 284 |
+
ERROR: [Errno 98] error while attempting to bind on address ('0.0.0.0', 8080): address already in use
|
| 285 |
+
INFO: Waiting for application shutdown.
|
| 286 |
+
INFO: Application shutdown complete.
|