Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
journal de demarrage
Browse files- etat/6tqc2otcxvsry3.log +54 -1
etat/6tqc2otcxvsry3.log
CHANGED
|
@@ -1,4 +1,4 @@
|
|
| 1 |
-
=== bootstrap v82-humming-extras-cu13 | 15:56:
|
| 2 |
[VL] 15:45:50 bootstrap v82-humming-extras-cu13
|
| 3 |
[VL] 15:45:51 pilote 550.54.15, CUDA runtime 13.0
|
| 4 |
[VL] 15:45:51 compat CUDA 13 active (LD_LIBRARY_PATH)
|
|
@@ -16,6 +16,10 @@
|
|
| 16 |
[VL] 15:46:48 cache de compilation, cle e32bc8688250c69d (NVIDIA_A100_80GB_PCIe, vllm 0.27.1)
|
| 17 |
[VL] 15:46:49 pas de cache pour cette cle : premiere compilation, il sera publie ensuite
|
| 18 |
[VL] 15:46:49 journal distant : https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/etat/6tqc2otcxvsry3.log
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
|
| 20 |
=== nvidia-smi ===
|
| 21 |
75441 MiB, 81920 MiB
|
|
@@ -162,3 +166,52 @@
|
|
| 162 |
(EngineCore pid=1604) INFO 08-23 15:56:10 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 163 |
(EngineCore pid=1604) INFO 08-23 15:56:11 [core.py:348] init engine (profile, create kv cache, warmup model) took 273.96 s (compilation: 23.71 s)
|
| 164 |
(EngineCore pid=1604) [fastokens] patch_transformers: successfully patched transformers v5.15.1
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
=== bootstrap v82-humming-extras-cu13 | 15:56:48 UTC ===
|
| 2 |
[VL] 15:45:50 bootstrap v82-humming-extras-cu13
|
| 3 |
[VL] 15:45:51 pilote 550.54.15, CUDA runtime 13.0
|
| 4 |
[VL] 15:45:51 compat CUDA 13 active (LD_LIBRARY_PATH)
|
|
|
|
| 16 |
[VL] 15:46:48 cache de compilation, cle e32bc8688250c69d (NVIDIA_A100_80GB_PCIe, vllm 0.27.1)
|
| 17 |
[VL] 15:46:49 pas de cache pour cette cle : premiere compilation, il sera publie ensuite
|
| 18 |
[VL] 15:46:49 journal distant : https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/etat/6tqc2otcxvsry3.log
|
| 19 |
+
[VL] 15:56:40 READY nemotron
|
| 20 |
+
[VL] 15:56:44 pont Anthropic en ligne sur 8080 (/v1/messages)
|
| 21 |
+
[VL] 15:56:44 serveur en ligne sur 18081 (API OpenAI, modele 'ornith')
|
| 22 |
+
[VL] 15:56:44 surveillance de https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/vllm_bootstrap.sh et du pont (rechargement en place)
|
| 23 |
|
| 24 |
=== nvidia-smi ===
|
| 25 |
75441 MiB, 81920 MiB
|
|
|
|
| 166 |
(EngineCore pid=1604) INFO 08-23 15:56:10 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 167 |
(EngineCore pid=1604) INFO 08-23 15:56:11 [core.py:348] init engine (profile, create kv cache, warmup model) took 273.96 s (compilation: 23.71 s)
|
| 168 |
(EngineCore pid=1604) [fastokens] patch_transformers: successfully patched transformers v5.15.1
|
| 169 |
+
(EngineCore pid=1604) INFO 08-23 15:56:20 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
|
| 170 |
+
(APIServer pid=872) INFO 08-23 15:56:20 [api_server.py:678] Supported tasks: ['generate']
|
| 171 |
+
(APIServer pid=872) INFO 08-23 15:56:21 [parser_manager.py:37] "auto" tool choice has been enabled.
|
| 172 |
+
(APIServer pid=872) INFO 08-23 15:56:31 [hf.py:540] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 173 |
+
(APIServer pid=872) WARNING 08-23 15:56:32 [model.py:1637] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 1.0, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
| 174 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [api_server.py:682] Starting vLLM server on http://0.0.0.0:18081
|
| 175 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:37] Available routes are:
|
| 176 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 177 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 178 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 179 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 180 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /load, Methods: GET
|
| 181 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /version, Methods: GET
|
| 182 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /health, Methods: GET
|
| 183 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /metrics, Methods: GET
|
| 184 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 185 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 186 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 187 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /ping, Methods: GET
|
| 188 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /ping, Methods: POST
|
| 189 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /invocations, Methods: POST
|
| 190 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 191 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 192 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 193 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 194 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 195 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 196 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 197 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 198 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 199 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 200 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 201 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 202 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 203 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 204 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 205 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 206 |
+
(APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:99] API server: waiting for HTTP server to start
|
| 207 |
+
(APIServer pid=872) INFO: Started server process [872]
|
| 208 |
+
(APIServer pid=872) INFO: Waiting for application startup.
|
| 209 |
+
(APIServer pid=872) INFO: Application startup complete.
|
| 210 |
+
(APIServer pid=872) INFO 08-23 15:56:36 [launcher.py:105] API server: HTTP server started
|
| 211 |
+
(APIServer pid=872) INFO: 127.0.0.1:35050 - "GET /health HTTP/1.1" 200 OK
|
| 212 |
+
(APIServer pid=872) INFO: 127.0.0.1:35060 - "GET /health HTTP/1.1" 200 OK
|
| 213 |
+
(APIServer pid=872) INFO: 127.0.0.1:56960 - "GET /health HTTP/1.1" 200 OK
|
| 214 |
+
(APIServer pid=872) INFO: 127.0.0.1:56960 - "GET /health HTTP/1.1" 200 OK
|
| 215 |
+
(APIServer pid=872) INFO: 127.0.0.1:56960 - "GET /v1/models HTTP/1.1" 200 OK
|
| 216 |
+
(APIServer pid=872) INFO: 127.0.0.1:56960 - "GET /v1/models HTTP/1.1" 200 OK
|
| 217 |
+
(APIServer pid=872) INFO: 127.0.0.1:56960 - "GET /v1/models HTTP/1.1" 200 OK
|