patdev commited on
Commit
77daea1
·
verified ·
1 Parent(s): 89af12e

Upload etat/carte-d0wveneybza9f4.log with huggingface_hub

Browse files
Files changed (1) hide show
  1. etat/carte-d0wveneybza9f4.log +9 -0
etat/carte-d0wveneybza9f4.log CHANGED
@@ -6,3 +6,12 @@ vllm 0.27.1 cap (12, 0) SM 170
6
  [ 60 s] Using FLASHINFER attention backend out of potential backends: ['FLASHINFER', 'TRITON_ATTN'].
7
  [240 s] Using uncalibrated q_scale 1.0 and/or prob_scale 1.0 with fp8 attention. This may cause accuracy issues. Pleas
8
  [300 s] Using flashinfer Mamba SSU backend.
 
 
 
 
 
 
 
 
 
 
6
  [ 60 s] Using FLASHINFER attention backend out of potential backends: ['FLASHINFER', 'TRITON_ATTN'].
7
  [240 s] Using uncalibrated q_scale 1.0 and/or prob_scale 1.0 with fp8 attention. This may cause accuracy issues. Pleas
8
  [300 s] Using flashinfer Mamba SSU backend.
9
+ Using FLASHINFER attention backend out of potential backends: ['FLASHINFER', 'TRITON_ATTN'].
10
+ Using 'MARLIN' NvFp4 MoE backend out of potential backends: ['FLASHINFER_TRTLLM', 'FLASHINFER_CUTEDSL', 'FLASHINFER_CUTEDSL_BATCHED', 'FLASH
11
+ Using MarlinNvFp4LinearKernel for NVFP4 GEMM
12
+ GPU KV cache size: 319,488 tokens
13
+ Maximum concurrency for 65,536 tokens per request: 4.88x
14
+ NE DEMARRE PAS
15
+ > ValueError: max_num_seqs (256) exceeds available Mamba cache blocks (117). Each decode sequence requires one Mamba cache block, so CUDA graph capture cannot proc
16
+ > RuntimeError: Engine core initialization failed. See root cause above. Failed core proc(s): {}
17
+ [CARTE] 19:01:44 TERMINE