patdev commited on
Commit
ba0032f
·
verified ·
1 Parent(s): 915c68d

journal de demarrage

Browse files
Files changed (1) hide show
  1. etat/6tqc2otcxvsry3.log +54 -1
etat/6tqc2otcxvsry3.log CHANGED
@@ -1,4 +1,4 @@
1
- === bootstrap v82-humming-extras-cu13 | 15:56:19 UTC ===
2
  [VL] 15:45:50 bootstrap v82-humming-extras-cu13
3
  [VL] 15:45:51 pilote 550.54.15, CUDA runtime 13.0
4
  [VL] 15:45:51 compat CUDA 13 active (LD_LIBRARY_PATH)
@@ -16,6 +16,10 @@
16
  [VL] 15:46:48 cache de compilation, cle e32bc8688250c69d (NVIDIA_A100_80GB_PCIe, vllm 0.27.1)
17
  [VL] 15:46:49 pas de cache pour cette cle : premiere compilation, il sera publie ensuite
18
  [VL] 15:46:49 journal distant : https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/etat/6tqc2otcxvsry3.log
 
 
 
 
19
 
20
  === nvidia-smi ===
21
  75441 MiB, 81920 MiB
@@ -162,3 +166,52 @@
162
  (EngineCore pid=1604) INFO 08-23 15:56:10 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
163
  (EngineCore pid=1604) INFO 08-23 15:56:11 [core.py:348] init engine (profile, create kv cache, warmup model) took 273.96 s (compilation: 23.71 s)
164
  (EngineCore pid=1604) [fastokens] patch_transformers: successfully patched transformers v5.15.1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ === bootstrap v82-humming-extras-cu13 | 15:56:48 UTC ===
2
  [VL] 15:45:50 bootstrap v82-humming-extras-cu13
3
  [VL] 15:45:51 pilote 550.54.15, CUDA runtime 13.0
4
  [VL] 15:45:51 compat CUDA 13 active (LD_LIBRARY_PATH)
 
16
  [VL] 15:46:48 cache de compilation, cle e32bc8688250c69d (NVIDIA_A100_80GB_PCIe, vllm 0.27.1)
17
  [VL] 15:46:49 pas de cache pour cette cle : premiere compilation, il sera publie ensuite
18
  [VL] 15:46:49 journal distant : https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/etat/6tqc2otcxvsry3.log
19
+ [VL] 15:56:40 READY nemotron
20
+ [VL] 15:56:44 pont Anthropic en ligne sur 8080 (/v1/messages)
21
+ [VL] 15:56:44 serveur en ligne sur 18081 (API OpenAI, modele 'ornith')
22
+ [VL] 15:56:44 surveillance de https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/vllm_bootstrap.sh et du pont (rechargement en place)
23
 
24
  === nvidia-smi ===
25
  75441 MiB, 81920 MiB
 
166
  (EngineCore pid=1604) INFO 08-23 15:56:10 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
167
  (EngineCore pid=1604) INFO 08-23 15:56:11 [core.py:348] init engine (profile, create kv cache, warmup model) took 273.96 s (compilation: 23.71 s)
168
  (EngineCore pid=1604) [fastokens] patch_transformers: successfully patched transformers v5.15.1
169
+ (EngineCore pid=1604) INFO 08-23 15:56:20 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
170
+ (APIServer pid=872) INFO 08-23 15:56:20 [api_server.py:678] Supported tasks: ['generate']
171
+ (APIServer pid=872) INFO 08-23 15:56:21 [parser_manager.py:37] "auto" tool choice has been enabled.
172
+ (APIServer pid=872) INFO 08-23 15:56:31 [hf.py:540] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
173
+ (APIServer pid=872) WARNING 08-23 15:56:32 [model.py:1637] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 1.0, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
174
+ (APIServer pid=872) INFO 08-23 15:56:35 [api_server.py:682] Starting vLLM server on http://0.0.0.0:18081
175
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:37] Available routes are:
176
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
177
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /docs, Methods: GET, HEAD
178
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
179
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
180
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /load, Methods: GET
181
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /version, Methods: GET
182
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /health, Methods: GET
183
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /metrics, Methods: GET
184
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /tokenize, Methods: POST
185
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /detokenize, Methods: POST
186
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/models, Methods: GET
187
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /ping, Methods: GET
188
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /ping, Methods: POST
189
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /invocations, Methods: POST
190
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
191
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
192
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/responses, Methods: POST
193
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
194
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
195
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/completions, Methods: POST
196
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/messages, Methods: POST
197
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
198
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /generative_scoring, Methods: POST
199
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
200
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
201
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
202
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/completions/render, Methods: POST
203
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
204
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
205
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
206
+ (APIServer pid=872) INFO 08-23 15:56:35 [launcher.py:99] API server: waiting for HTTP server to start
207
+ (APIServer pid=872) INFO: Started server process [872]
208
+ (APIServer pid=872) INFO: Waiting for application startup.
209
+ (APIServer pid=872) INFO: Application startup complete.
210
+ (APIServer pid=872) INFO 08-23 15:56:36 [launcher.py:105] API server: HTTP server started
211
+ (APIServer pid=872) INFO: 127.0.0.1:35050 - "GET /health HTTP/1.1" 200 OK
212
+ (APIServer pid=872) INFO: 127.0.0.1:35060 - "GET /health HTTP/1.1" 200 OK
213
+ (APIServer pid=872) INFO: 127.0.0.1:56960 - "GET /health HTTP/1.1" 200 OK
214
+ (APIServer pid=872) INFO: 127.0.0.1:56960 - "GET /health HTTP/1.1" 200 OK
215
+ (APIServer pid=872) INFO: 127.0.0.1:56960 - "GET /v1/models HTTP/1.1" 200 OK
216
+ (APIServer pid=872) INFO: 127.0.0.1:56960 - "GET /v1/models HTTP/1.1" 200 OK
217
+ (APIServer pid=872) INFO: 127.0.0.1:56960 - "GET /v1/models HTTP/1.1" 200 OK