patdev/linux-desktop / state /agent /legacy-vllm.log
patdev's picture
download
raw
215 kB
WARNING 09-19 23:12:52 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-19 23:12:52 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-19 23:12:52 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
(APIServer pid=2057) INFO 09-19 23:12:53 [api_utils.py:347]
(APIServer pid=2057) INFO 09-19 23:12:53 [api_utils.py:347] █ █ █▄ ▄█
(APIServer pid=2057) INFO 09-19 23:12:53 [api_utils.py:347] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.29.0
(APIServer pid=2057) INFO 09-19 23:12:53 [api_utils.py:347] █▄█▀ █ █ █ █ model Qwen/Qwen3-1.7B
(APIServer pid=2057) INFO 09-19 23:12:53 [api_utils.py:347] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
(APIServer pid=2057) INFO 09-19 23:12:53 [api_utils.py:347]
(APIServer pid=2057) INFO 09-19 23:12:53 [api_utils.py:286] non-default args: {'model_tag': 'Qwen/Qwen3-1.7B', 'enable_auto_tool_choice': True, 'tool_call_parser': 'hermes', 'host': '127.0.0.1', 'port': 43123, 'uvicorn_log_level': 'warning', 'model': 'Qwen/Qwen3-1.7B', 'dtype': 'bfloat16', 'max_model_len': 32768, 'enforce_eager': True, 'served_model_name': ['local-qwen3-1.7b'], 'generation_config': 'vllm', 'reasoning_parser': 'qwen3', 'enable_prefix_caching': True, 'max_num_batched_tokens': 2048, 'max_num_seqs': 1}
(APIServer pid=2057) INFO 09-19 23:13:00 [model.py:684] Resolved architecture: Qwen3ForCausalLM
(APIServer pid=2057) INFO 09-19 23:13:00 [model.py:2021] Using max model len 32768
(APIServer pid=2057) Parse safetensors files: 0%| | 0/2 [00:00<?, ?it/s] Parse safetensors files: 50%|█████ | 1/2 [00:00<00:00, 5.72it/s] Parse safetensors files: 100%|██████████| 2/2 [00:00<00:00, 9.78it/s]
(APIServer pid=2057) WARNING 09-19 23:13:00 [vllm.py:1371] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
(APIServer pid=2057) WARNING 09-19 23:13:00 [vllm.py:1406] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
(APIServer pid=2057) INFO 09-19 23:13:00 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
(APIServer pid=2057) WARNING 09-19 23:13:00 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(APIServer pid=2057) INFO 09-19 23:13:02 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
WARNING 09-19 23:13:09 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-19 23:13:09 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-19 23:13:09 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
(EngineCore pid=2383) INFO 09-19 23:13:09 [core.py:123] Initializing a V1 LLM engine (v0.29.0) with config: model='Qwen/Qwen3-1.7B', speculative_config=None, tokenizer='Qwen/Qwen3-1.7B', skip_tokenizer_init=False, tokenizer_mode=auto, revision=main, tokenizer_revision=main, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=local-qwen3-1.7b, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto')
(EngineCore pid=2383) INFO 09-19 23:13:09 [multiproc_executor.py:153] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip=10.111.38.55 (local), world_size=1, local_world_size=1
(EngineCore pid=2383) INFO 09-19 23:13:09 [ompmultiprocessing.py:185] OpenMP thread binding info:
(EngineCore pid=2383) INFO 09-19 23:13:09 [ompmultiprocessing.py:185] VLLM_CPU_OMP_THREADS_BIND='0-1', auto_setup=False, skip_setup=False
(EngineCore pid=2383) INFO 09-19 23:13:09 [ompmultiprocessing.py:185] local_world_size=1, reserve_cpu_num=1
(EngineCore pid=2383) INFO 09-19 23:13:09 [ompmultiprocessing.py:185] local_rank=0, core ids=[0, 1]
(EngineCore pid=2383) INFO 09-19 23:13:09 [ompmultiprocessing.py:185] reserved_cpus=[]
WARNING 09-19 23:13:16 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-19 23:13:16 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-19 23:13:16 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
get_mempolicy: Operation not permitted
[W919 23:13:17.340243711 utils.cpp:41] Warning: numa_migrate_pages failed. errno: 1 (function init_cpu_memory_env)
set_mempolicy: Operation not permitted
[W919 23:13:17.340420637 utils.cpp:65] Warning: numa_set_membind failed. errno: 1 (function init_cpu_memory_env)
WARNING 09-19 23:13:17 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(Worker pid=2493) INFO 09-19 23:13:17 [parallel_state.py:1775] world_size=1 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_5b4bbca93469468f8d66263b39e3c7fc backend=gloo
(Worker pid=2493) INFO 09-19 23:13:17 [parallel_state.py:2119] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A, EPLB rank N/A
(Worker pid=2493) INFO 09-19 23:13:17 [cpu_model_runner.py:131] Starting to load model Qwen/Qwen3-1.7B...
(Worker pid=2493) INFO 09-19 23:13:36 [weight_utils.py:526] Time spent downloading weights for Qwen/Qwen3-1.7B: 18.003470 seconds
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] WorkerProc failed to start.
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] Traceback (most recent call last):
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 911, in worker_main
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] worker = WorkerProc(*args, **kwargs)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] return func(*args, **kwargs)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 680, in __init__
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] self.worker.load_model()
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_worker.py", line 495, in load_model
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] self.model_runner.load_model(load_dummy_weights=load_dummy_weights)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] return func(*args, **kwargs)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/cpu_model_runner.py", line 132, in load_model
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] self.model = get_model(vllm_config=self.vllm_config)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/model_loader/__init__.py", line 137, in get_model
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] return loader.load_model(
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] return func(*args, **kwargs)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/model_loader/base_loader.py", line 64, in load_model
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] self.load_weights(model, model_config)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] return func(*args, **kwargs)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/model_loader/default_loader.py", line 427, in load_weights
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] loaded_weights = model.load_weights(self.get_all_weights(model_config, model))
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/models/qwen3.py", line 339, in load_weights
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] return loader.load_weights(weights)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/model_loader/reload/torchao_decorator.py", line 50, in patched_model_load_weights
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] return original_load_weights(self, weights, *args, **kwargs)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/models/utils.py", line 444, in load_weights
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] autoloaded_weights = set(self._load_module("", self.module, weights))
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/models/utils.py", line 378, in _load_module
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] for child_prefix, child_weights in self._groupby_prefix(weights):
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/models/utils.py", line 256, in _groupby_prefix
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] for prefix, group in itertools.groupby(weights_by_parts, key=lambda x: x[0][0]):
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/models/utils.py", line 251, in <genexpr>
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] weights_by_parts = (
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/models/utils.py", line 451, in _filter_skipped
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] for name, weight in weights:
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/models/utils.py", line 140, in apply
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] for name, data in weights:
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/model_loader/default_loader.py", line 333, in get_all_weights
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] yield from self._get_weights_iterator(primary_weights)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/model_loader/default_loader.py", line 319, in <genexpr>
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] return ((source.prefix + name, tensor) for (name, tensor) in weights_iterator)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/model_loader/weight_utils.py", line 857, in safetensors_weights_iterator
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] total_bytes = _get_checkpoints_size_bytes(sorted_files)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/model_loader/weight_utils.py", line 689, in _get_checkpoints_size_bytes
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] return sum(os.path.getsize(f) for f in files)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "/opt/vllm/lib/python3.11/site-packages/vllm/model_executor/model_loader/weight_utils.py", line 689, in <genexpr>
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] return sum(os.path.getsize(f) for f in files)
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] ^^^^^^^^^^^^^^^^^^
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] File "<frozen genericpath>", line 50, in getsize
(Worker pid=2493) ERROR 09-19 23:13:36 [multiproc_executor.py:944] FileNotFoundError: [Errno 2] No such file or directory: '/data/hf/hub/models--Qwen--Qwen3-1.7B/snapshots/70d244cc86ccca08cf5af4e1e306ecf908b1ad5e/model-00002-of-00002.safetensors'
(EngineCore pid=2383) INFO 09-19 23:13:36 [multiproc_executor.py:472] [shutdown] Executor: waiting for worker exit count=1
(EngineCore pid=2383) INFO 09-19 23:13:37 [multiproc_executor.py:479] [shutdown] Executor: all workers exited gracefully
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] EngineCore failed to start.
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] Traceback (most recent call last):
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1336, in run_engine_core
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] engine_core = EngineCoreProc(*args, engine_index=dp_rank, **kwargs)
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] File "/opt/vllm/lib/python3.11/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] return func(*args, **kwargs)
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] ^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1093, in __init__
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] super().__init__(
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 134, in __init__
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] self.model_executor = executor_class(vllm_config)
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] ^^^^^^^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 116, in __init__
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] super().__init__(vllm_config)
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] File "/opt/vllm/lib/python3.11/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] return func(*args, **kwargs)
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] ^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/abstract.py", line 110, in __init__
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] self._init_executor()
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 214, in _init_executor
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] self.workers = WorkerProc.wait_for_ready(unready_workers)
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 808, in wait_for_ready
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] raise e from None
(EngineCore pid=2383) ERROR 09-19 23:13:37 [core.py:1374] Exception: WorkerProc initialization failed due to an exception in a background process. See stack trace for root cause.
(EngineCore pid=2383) Process EngineCore:
(EngineCore pid=2383) Traceback (most recent call last):
(EngineCore pid=2383) File "/usr/lib/python3.11/multiprocessing/process.py", line 314, in _bootstrap
(EngineCore pid=2383) self.run()
(EngineCore pid=2383) File "/usr/lib/python3.11/multiprocessing/process.py", line 108, in run
(EngineCore pid=2383) self._target(*self._args, **self._kwargs)
(EngineCore pid=2383) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1378, in run_engine_core
(EngineCore pid=2383) raise e
(EngineCore pid=2383) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1336, in run_engine_core
(EngineCore pid=2383) engine_core = EngineCoreProc(*args, engine_index=dp_rank, **kwargs)
(EngineCore pid=2383) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=2383) File "/opt/vllm/lib/python3.11/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(EngineCore pid=2383) return func(*args, **kwargs)
(EngineCore pid=2383) ^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=2383) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1093, in __init__
(EngineCore pid=2383) super().__init__(
(EngineCore pid=2383) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 134, in __init__
(EngineCore pid=2383) self.model_executor = executor_class(vllm_config)
(EngineCore pid=2383) ^^^^^^^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=2383) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 116, in __init__
(EngineCore pid=2383) super().__init__(vllm_config)
(EngineCore pid=2383) File "/opt/vllm/lib/python3.11/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(EngineCore pid=2383) return func(*args, **kwargs)
(EngineCore pid=2383) ^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=2383) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/abstract.py", line 110, in __init__
(EngineCore pid=2383) self._init_executor()
(EngineCore pid=2383) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 214, in _init_executor
(EngineCore pid=2383) self.workers = WorkerProc.wait_for_ready(unready_workers)
(EngineCore pid=2383) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=2383) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 808, in wait_for_ready
(EngineCore pid=2383) raise e from None
(EngineCore pid=2383) Exception: WorkerProc initialization failed due to an exception in a background process. See stack trace for root cause.
(APIServer pid=2057) Traceback (most recent call last):
(APIServer pid=2057) File "/opt/vllm/bin/vllm", line 8, in <module>
(APIServer pid=2057) sys.exit(main())
(APIServer pid=2057) ^^^^^^
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/cli/main.py", line 97, in main
(APIServer pid=2057) args.dispatch_function(args)
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/cli/serve.py", line 153, in cmd
(APIServer pid=2057) uvloop.run(run_server(args))
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/uvloop/__init__.py", line 92, in run
(APIServer pid=2057) return runner.run(wrapper())
(APIServer pid=2057) ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2057) File "/usr/lib/python3.11/asyncio/runners.py", line 118, in run
(APIServer pid=2057) return self._loop.run_until_complete(task)
(APIServer pid=2057) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2057) File "uvloop/loop.pyx", line 1518, in uvloop.loop.Loop.run_until_complete
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/uvloop/__init__.py", line 48, in wrapper
(APIServer pid=2057) return await main
(APIServer pid=2057) ^^^^^^^^^^
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/launchers/api_server/entry.py", line 176, in run_server
(APIServer pid=2057) await run_server_worker(listen_address, sock, args, **uvicorn_kwargs)
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/launchers/api_server/entry.py", line 190, in run_server_worker
(APIServer pid=2057) async with build_async_engine_client(
(APIServer pid=2057) File "/usr/lib/python3.11/contextlib.py", line 204, in __aenter__
(APIServer pid=2057) return await anext(self.gen)
(APIServer pid=2057) ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/launchers/api_server/entry.py", line 58, in build_async_engine_client
(APIServer pid=2057) async with build_async_engine_client_from_engine_args(
(APIServer pid=2057) File "/usr/lib/python3.11/contextlib.py", line 204, in __aenter__
(APIServer pid=2057) return await anext(self.gen)
(APIServer pid=2057) ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/launchers/api_server/entry.py", line 94, in build_async_engine_client_from_engine_args
(APIServer pid=2057) async_llm = AsyncLLM.from_vllm_config(
(APIServer pid=2057) ^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/async_llm.py", line 227, in from_vllm_config
(APIServer pid=2057) return cls(
(APIServer pid=2057) ^^^^
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/async_llm.py", line 156, in __init__
(APIServer pid=2057) self.engine_core = EngineCoreClient.make_async_mp_client(
(APIServer pid=2057) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(APIServer pid=2057) return func(*args, **kwargs)
(APIServer pid=2057) ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core_client.py", line 139, in make_async_mp_client
(APIServer pid=2057) return AsyncMPClient(*client_args)
(APIServer pid=2057) ^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(APIServer pid=2057) return func(*args, **kwargs)
(APIServer pid=2057) ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core_client.py", line 991, in __init__
(APIServer pid=2057) super().__init__(
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core_client.py", line 609, in __init__
(APIServer pid=2057) with launch_core_engines(
(APIServer pid=2057) File "/usr/lib/python3.11/contextlib.py", line 144, in __exit__
(APIServer pid=2057) next(self.gen)
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/utils.py", line 1240, in launch_core_engines
(APIServer pid=2057) wait_for_engine_startup(
(APIServer pid=2057) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/utils.py", line 1320, in wait_for_engine_startup
(APIServer pid=2057) raise RuntimeError(
(APIServer pid=2057) RuntimeError: Engine core initialization failed. See root cause above. Failed core proc(s): {}
WARNING 09-19 23:36:44 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-19 23:36:44 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-19 23:36:44 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
(APIServer pid=874) INFO 09-19 23:36:45 [api_utils.py:347]
(APIServer pid=874) INFO 09-19 23:36:45 [api_utils.py:347] █ █ █▄ ▄█
(APIServer pid=874) INFO 09-19 23:36:45 [api_utils.py:347] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.29.0
(APIServer pid=874) INFO 09-19 23:36:45 [api_utils.py:347] █▄█▀ █ █ █ █ model /data/local-agent/models/Qwen3-0.6B
(APIServer pid=874) INFO 09-19 23:36:45 [api_utils.py:347] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
(APIServer pid=874) INFO 09-19 23:36:45 [api_utils.py:347]
(APIServer pid=874) INFO 09-19 23:36:45 [api_utils.py:286] non-default args: {'model_tag': '/data/local-agent/models/Qwen3-0.6B', 'enable_auto_tool_choice': True, 'tool_call_parser': 'hermes', 'host': '127.0.0.1', 'port': 43123, 'uvicorn_log_level': 'warning', 'model': '/data/local-agent/models/Qwen3-0.6B', 'dtype': 'bfloat16', 'max_model_len': 32768, 'enforce_eager': True, 'served_model_name': ['local-qwen3-0.6b'], 'generation_config': 'vllm', 'reasoning_parser': 'qwen3', 'enable_prefix_caching': True, 'max_num_batched_tokens': 2048, 'max_num_seqs': 1}
(APIServer pid=874) INFO 09-19 23:36:54 [model.py:684] Resolved architecture: Qwen3ForCausalLM
(APIServer pid=874) INFO 09-19 23:36:54 [model.py:2021] Using max model len 32768
(APIServer pid=874) WARNING 09-19 23:36:55 [vllm.py:1371] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
(APIServer pid=874) WARNING 09-19 23:36:55 [vllm.py:1406] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
(APIServer pid=874) INFO 09-19 23:36:55 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
(APIServer pid=874) WARNING 09-19 23:36:55 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(APIServer pid=874) INFO 09-19 23:36:56 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
WARNING 09-19 23:37:04 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-19 23:37:04 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-19 23:37:04 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
(EngineCore pid=1183) INFO 09-19 23:37:05 [core.py:123] Initializing a V1 LLM engine (v0.29.0) with config: model='/data/local-agent/models/Qwen3-0.6B', speculative_config=None, tokenizer='/data/local-agent/models/Qwen3-0.6B', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=local-qwen3-0.6b, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto')
(EngineCore pid=1183) INFO 09-19 23:37:05 [multiproc_executor.py:153] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip=10.111.73.11 (local), world_size=1, local_world_size=1
(EngineCore pid=1183) INFO 09-19 23:37:05 [ompmultiprocessing.py:185] OpenMP thread binding info:
(EngineCore pid=1183) INFO 09-19 23:37:05 [ompmultiprocessing.py:185] VLLM_CPU_OMP_THREADS_BIND='0-1', auto_setup=False, skip_setup=False
(EngineCore pid=1183) INFO 09-19 23:37:05 [ompmultiprocessing.py:185] local_world_size=1, reserve_cpu_num=1
(EngineCore pid=1183) INFO 09-19 23:37:05 [ompmultiprocessing.py:185] local_rank=0, core ids=[0, 1]
(EngineCore pid=1183) INFO 09-19 23:37:05 [ompmultiprocessing.py:185] reserved_cpus=[]
WARNING 09-19 23:37:10 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-19 23:37:10 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-19 23:37:10 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
get_mempolicy: Operation not permitted
[W919 23:37:11.354201496 utils.cpp:41] Warning: numa_migrate_pages failed. errno: 1 (function init_cpu_memory_env)
set_mempolicy: Operation not permitted
[W919 23:37:11.354296562 utils.cpp:65] Warning: numa_set_membind failed. errno: 1 (function init_cpu_memory_env)
WARNING 09-19 23:37:11 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(Worker pid=1294) INFO 09-19 23:37:11 [parallel_state.py:1775] world_size=1 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_ce881f712a554a45a13298efc6eea455 backend=gloo
(Worker pid=1294) INFO 09-19 23:37:11 [parallel_state.py:2119] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A, EPLB rank N/A
(Worker pid=1294) INFO 09-19 23:37:11 [cpu_model_runner.py:131] Starting to load model /data/local-agent/models/Qwen3-0.6B...
(Worker pid=1294) INFO 09-19 23:37:12 [weight_utils.py:863] Filesystem type for checkpoints: FUSE. Checkpoint size: 1.40 GiB. Available RAM: 12.92 GiB.
(Worker pid=1294) INFO 09-19 23:37:12 [weight_utils.py:886] Auto-prefetch is disabled because the filesystem (FUSE) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
(Worker pid=1294) Loading safetensors checkpoint shards: 0% Completed | 0/1 [00:00<?, ?it/s]
(Worker pid=1294) Loading safetensors checkpoint shards: 100% Completed | 1/1 [02:02<00:00, 122.74s/it]
(Worker pid=1294) Loading safetensors checkpoint shards: 100% Completed | 1/1 [02:02<00:00, 122.74s/it]
(Worker pid=1294)
(Worker pid=1294) INFO 09-19 23:39:15 [default_loader.py:430] Loading weights took 122.81 seconds
(EngineCore pid=1183) WARNING 09-19 23:39:16 [torch_utils.py:265] OMP_NUM_THREADS=2 is set; leaving Torch threads at 2 for serving. Multi-threaded torch CPU ops during serving can degrade performance through spin-wait contention and cgroup CPU-quota throttling.
(EngineCore pid=1183) INFO 09-19 23:39:16 [utils.py:306] Using LBHNC KV cache layout.
(Worker pid=1294) INFO 09-19 23:39:16 [cpu_worker.py:255] Explicitly set (4.0/14.9) GiB for KV cache on node 0.
(EngineCore pid=1183) INFO 09-19 23:39:16 [kv_cache_utils.py:2032] GPU KV cache size: 37,376 tokens, Maximum concurrency for 32,768 tokens per request: 1.14x
(EngineCore pid=1183) INFO 09-19 23:39:17 [core.py:368] init engine (profile, create kv cache, warmup model) took 1.53 s
(EngineCore pid=1183) WARNING 09-19 23:39:20 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(EngineCore pid=1183) WARNING 09-19 23:39:20 [vllm.py:1371] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
(EngineCore pid=1183) WARNING 09-19 23:39:20 [vllm.py:1406] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
(EngineCore pid=1183) INFO 09-19 23:39:20 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
(EngineCore pid=1183) INFO 09-19 23:39:20 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
(APIServer pid=874) INFO 09-19 23:39:20 [entry.py:135] Supported tasks: ['generate']
(APIServer pid=874) INFO 09-19 23:39:20 [parser_manager.py:36] "auto" tool choice has been enabled.
(APIServer pid=874) INFO 09-19 23:39:21 [hf.py:547] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
(APIServer pid=874) INFO 09-19 23:39:21 [entry.py:139] Starting vLLM server on http://127.0.0.1:43123
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:61] Available routes are:
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /openapi.json, Methods: HEAD, GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /docs, Methods: HEAD, GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /docs/oauth2-redirect, Methods: HEAD, GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /redoc, Methods: HEAD, GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /load, Methods: GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /version, Methods: GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /health, Methods: GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /metrics, Methods: GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /tokenize, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /detokenize, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/models, Methods: GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /ping, Methods: GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /ping, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /invocations, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/chat/completions, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/chat/completions/batch, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/responses, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/responses/{response_id}, Methods: GET
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/responses/{response_id}/cancel, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/completions, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/messages, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/messages/count_tokens, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /generative_scoring, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /scale_elastic_ep, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /is_scaling_elastic_ep, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/chat/completions/render, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/messages/render, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/completions/render, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/chat/completions/derender, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /v1/completions/derender, Methods: POST
(APIServer pid=874) INFO 09-19 23:39:21 [launcher.py:70] Route: /inference/v1/generate, Methods: POST
(Worker pid=1294) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=1294) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=1294) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=1294) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=1294) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=1294) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=1294) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=1294) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=1294) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=1294) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=1294) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=1294) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] WorkerProc hit an exception.
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] Traceback (most recent call last):
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 1047, in _execute_worker_rpc
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] output = func(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/utils/_contextlib.py", line 124, in decorate_context
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return func(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_worker.py", line 1105, in sample_tokens
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return self.model_runner.sample_tokens(grammar_output)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/utils/_contextlib.py", line 124, in decorate_context
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return func(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_model_runner.py", line 4664, in sample_tokens
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] sampler_output = self._sample(logits, spec_decode_metadata)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_model_runner.py", line 3731, in _sample
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return self.sampler(
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return self._call_impl(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1789, in _call_impl
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return forward_call(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/sampler.py", line 103, in forward
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] sampled, processed_logprobs = self.sample(logits, sampling_metadata)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/sampler.py", line 287, in sample
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] random_sampled, processed_logprobs = self.topk_topp_sampler(
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return self._call_impl(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1789, in _call_impl
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return forward_call(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/ops/topk_topp_sampler.py", line 204, in forward_cpu
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return compiled_random_sample(logits), logits_to_return
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_dynamo/eval_frame.py", line 1183, in compile_wrapper
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] raise e.remove_dynamo_frames() from None # see TORCHDYNAMO_VERBOSE=1
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1079, in _compile_fx_inner
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] raise InductorError(e, currentframe()).with_traceback(
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1059, in _compile_fx_inner
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] mb_compiled_graph = fx_codegen_and_compile(
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1847, in fx_codegen_and_compile
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return scheme.codegen_and_compile(gm, example_inputs, inputs_to_check, graph_kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1608, in codegen_and_compile
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] compiled_module = graph.compile_to_module()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2669, in compile_to_module
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return self._compile_to_module()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2679, in _compile_to_module
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] mod = self._compile_to_module_lines(wrapper_code)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2754, in _compile_to_module_lines
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] mod = PyCodeCache.load_by_key_path(
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4622, in load_by_key_path
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] mod = _reload_python_module(key, path, set_sys_modules=set_sys_modules)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/runtime/compile_tasks.py", line 35, in _reload_python_module
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] exec(code, mod.__dict__, mod.__dict__)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/tmp/torchinductor_agent/hu/chuaszb6rgufguvuy3xrzwcgcm6eug2vmkmnvpi3hie6u35dza3e.py", line 33, in <module>
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] cpp_fused__softmax_argmax_div_exponential_0 = async_compile.cpp_pybinding(['const float*', 'const int64_t*', 'float*', 'float*', 'float*', 'int64_t*', 'const int64_t'], r'''
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/async_compile.py", line 570, in cpp_pybinding
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return CppPythonBindingsCodeCache.load_pybinding(argtypes, source_code)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4084, in load_pybinding
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return cls.load_pybinding_async(*args, **kwargs)()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4076, in future
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] result = get_result()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 3853, in load_fn
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] result = worker_fn()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 3882, in _worker_compile_cpp
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] builder.build()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 2571, in build
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] run_compile_cmd(build_cmd, cwd=_build_tmp_dir)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 718, in run_compile_cmd
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] _run_compile_cmd(cmd_line, cwd)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 713, in _run_compile_cmd
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] raise exc.CppCompileError(cmd, output) from e
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] torch._inductor.exc.InductorError: CppCompileError: C++ compile error
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] Command:
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] g++ /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp -D TORCH_INDUCTOR_CPP_WRAPPER -D STANDALONE_TORCH_HEADER -D TORCH_INDUCTOR_PRECOMPILE_HEADERS -D C10_USING_CUSTOM_GENERATED_MACROS -D CPU_CAPABILITY_AVX512 -O3 -DNDEBUG -fno-omit-frame-pointer -g1 -fno-trapping-math -funsafe-math-optimizations -ffinite-math-only -fno-signed-zeros -fno-finite-math-only -fno-unsafe-math-optimizations -fmath-errno -ffp-contract=off -fexcess-precision=fast -fno-tree-loop-vectorize -march=native -shared -fPIC -Wall -std=c++20 -Wno-unused-variable -Wno-unknown-pragmas -pedantic -fopenmp -include /tmp/torchinductor_agent/precompiled_headers/c3d7kkupnir2q25xqv5t7pw3ruepjadpjxcxa2m555qoy5o5bca4.h -I/usr/include/python3.11 -I/opt/vllm/lib/python3.11/site-packages/torch/include -I/opt/vllm/lib/python3.11/site-packages/torch/include/torch/csrc/api/include -mavx512f -mavx512dq -mavx512vl -mavx512bw -mfma -mavx512vnni -mavx512vl -o /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.so -ltorch -ltorch_cpu -ltorch_python -lgomp -L/usr/lib/x86_64-linux-gnu -L/opt/vllm/lib/python3.11/site-packages/torch/lib
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] Output:
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp:266:10: fatal error: Python.h: No such file or directory
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] 266 | #include <Python.h>
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] | ^~~~~~~~~~
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] compilation terminated.
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] Set TORCHDYNAMO_VERBOSE=1 for the internal stack trace (please do this especially if you're reporting a bug to PyTorch). For even more developer context, set TORCH_LOGS="+dynamo"
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] Traceback (most recent call last):
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 1047, in _execute_worker_rpc
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] output = func(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/utils/_contextlib.py", line 124, in decorate_context
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return func(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_worker.py", line 1105, in sample_tokens
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return self.model_runner.sample_tokens(grammar_output)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/utils/_contextlib.py", line 124, in decorate_context
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return func(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_model_runner.py", line 4664, in sample_tokens
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] sampler_output = self._sample(logits, spec_decode_metadata)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_model_runner.py", line 3731, in _sample
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return self.sampler(
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return self._call_impl(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1789, in _call_impl
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return forward_call(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/sampler.py", line 103, in forward
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] sampled, processed_logprobs = self.sample(logits, sampling_metadata)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/sampler.py", line 287, in sample
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] random_sampled, processed_logprobs = self.topk_topp_sampler(
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return self._call_impl(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1789, in _call_impl
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return forward_call(*args, **kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/ops/topk_topp_sampler.py", line 204, in forward_cpu
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return compiled_random_sample(logits), logits_to_return
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_dynamo/eval_frame.py", line 1183, in compile_wrapper
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] raise e.remove_dynamo_frames() from None # see TORCHDYNAMO_VERBOSE=1
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1079, in _compile_fx_inner
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] raise InductorError(e, currentframe()).with_traceback(
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1059, in _compile_fx_inner
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] mb_compiled_graph = fx_codegen_and_compile(
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1847, in fx_codegen_and_compile
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return scheme.codegen_and_compile(gm, example_inputs, inputs_to_check, graph_kwargs)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1608, in codegen_and_compile
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] compiled_module = graph.compile_to_module()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2669, in compile_to_module
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return self._compile_to_module()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2679, in _compile_to_module
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] mod = self._compile_to_module_lines(wrapper_code)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2754, in _compile_to_module_lines
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] mod = PyCodeCache.load_by_key_path(
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4622, in load_by_key_path
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] mod = _reload_python_module(key, path, set_sys_modules=set_sys_modules)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/runtime/compile_tasks.py", line 35, in _reload_python_module
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] exec(code, mod.__dict__, mod.__dict__)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/tmp/torchinductor_agent/hu/chuaszb6rgufguvuy3xrzwcgcm6eug2vmkmnvpi3hie6u35dza3e.py", line 33, in <module>
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] cpp_fused__softmax_argmax_div_exponential_0 = async_compile.cpp_pybinding(['const float*', 'const int64_t*', 'float*', 'float*', 'float*', 'int64_t*', 'const int64_t'], r'''
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/async_compile.py", line 570, in cpp_pybinding
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return CppPythonBindingsCodeCache.load_pybinding(argtypes, source_code)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4084, in load_pybinding
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] return cls.load_pybinding_async(*args, **kwargs)()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4076, in future
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] result = get_result()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 3853, in load_fn
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] result = worker_fn()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] ^^^^^^^^^^^
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 3882, in _worker_compile_cpp
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] builder.build()
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 2571, in build
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] run_compile_cmd(build_cmd, cwd=_build_tmp_dir)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 718, in run_compile_cmd
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] _run_compile_cmd(cmd_line, cwd)
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 713, in _run_compile_cmd
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] raise exc.CppCompileError(cmd, output) from e
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] torch._inductor.exc.InductorError: CppCompileError: C++ compile error
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] Command:
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] g++ /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp -D TORCH_INDUCTOR_CPP_WRAPPER -D STANDALONE_TORCH_HEADER -D TORCH_INDUCTOR_PRECOMPILE_HEADERS -D C10_USING_CUSTOM_GENERATED_MACROS -D CPU_CAPABILITY_AVX512 -O3 -DNDEBUG -fno-omit-frame-pointer -g1 -fno-trapping-math -funsafe-math-optimizations -ffinite-math-only -fno-signed-zeros -fno-finite-math-only -fno-unsafe-math-optimizations -fmath-errno -ffp-contract=off -fexcess-precision=fast -fno-tree-loop-vectorize -march=native -shared -fPIC -Wall -std=c++20 -Wno-unused-variable -Wno-unknown-pragmas -pedantic -fopenmp -include /tmp/torchinductor_agent/precompiled_headers/c3d7kkupnir2q25xqv5t7pw3ruepjadpjxcxa2m555qoy5o5bca4.h -I/usr/include/python3.11 -I/opt/vllm/lib/python3.11/site-packages/torch/include -I/opt/vllm/lib/python3.11/site-packages/torch/include/torch/csrc/api/include -mavx512f -mavx512dq -mavx512vl -mavx512bw -mfma -mavx512vnni -mavx512vl -o /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.so -ltorch -ltorch_cpu -ltorch_python -lgomp -L/usr/lib/x86_64-linux-gnu -L/opt/vllm/lib/python3.11/site-packages/torch/lib
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] Output:
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp:266:10: fatal error: Python.h: No such file or directory
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] 266 | #include <Python.h>
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] | ^~~~~~~~~~
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] compilation terminated.
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055] Set TORCHDYNAMO_VERBOSE=1 for the internal stack trace (please do this especially if you're reporting a bug to PyTorch). For even more developer context, set TORCH_LOGS="+dynamo"
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(Worker pid=1294) ERROR 09-19 23:41:07 [multiproc_executor.py:1055]
(EngineCore pid=1183) ERROR 09-19 23:41:07 [dump_input.py:72] Dumping input data for V1 LLM engine (v0.29.0) with config: model='/data/local-agent/models/Qwen3-0.6B', speculative_config=None, tokenizer='/data/local-agent/models/Qwen3-0.6B', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=local-qwen3-0.6b, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto'),
(EngineCore pid=1183) ERROR 09-19 23:41:07 [dump_input.py:79] Dumping scheduler output for model execution: SchedulerOutput(scheduled_new_reqs=[NewRequestData(req_id=chatcmpl-af952ee6038dc9a6-8a70259d,prompt_token_ids_len=20,prefill_token_ids_len=None,mm_features=[],sampling_params=SamplingParams(n=1, presence_penalty=0.0, frequency_penalty=0.0, repetition_penalty=1.0, temperature=0.7, top_p=1.0, top_k=0, min_p=0.0, seed=None, stop=[], stop_token_ids=[151643], bad_words=[], thinking_token_budget=None, include_stop_str_in_output=False, ignore_eos=False, max_tokens=256, min_tokens=0, logprobs=None, prompt_logprobs=None, skip_special_tokens=False, spaces_between_special_tokens=True, structured_outputs=None, extra_args=None),block_ids=([1],),num_computed_tokens=0,lora_request=None,prompt_embeds_shape=None)], scheduled_cached_reqs=CachedRequestData(req_ids=[],resumed_req_ids=set(),new_token_ids_lens=[],all_token_ids_lens={},new_block_ids=[],num_computed_tokens=[],num_output_tokens=[]), num_scheduled_tokens={chatcmpl-af952ee6038dc9a6-8a70259d: 20}, total_num_scheduled_tokens=20, scheduled_spec_decode_tokens={}, scheduled_encoder_inputs={}, num_common_prefix_blocks=[1], finished_req_ids=[], free_encoder_mm_hashes=[], scheduled_encoder_input_stats=null, preempted_req_ids=[], has_structured_output_requests=false, pending_structured_output_tokens=false, num_invalid_spec_tokens=null, kv_connector_metadata=null, has_sync_kv_loads=false, ec_connector_metadata=null, ec_manager_metadata=null, new_block_ids_to_zero=null, kv_cache_block_copies=null, kv_connector_block_state=null, num_spec_tokens_to_schedule=0)
(EngineCore pid=1183) ERROR 09-19 23:41:07 [dump_input.py:81] Dumping scheduler stats: SchedulerStats(num_running_reqs=1, num_waiting_reqs=0, num_skipped_waiting_reqs=0, step_counter=0, current_wave=0, kv_cache_usage=0.0034364261168384758, iteration_details=None, prefix_cache_stats=PrefixCacheStats(reset=False, requests=1, queries=20, hits=0, preempted_requests=0, preempted_queries=0, preempted_hits=0), connector_prefix_cache_stats=None, kv_cache_eviction_events=[], spec_decoding_stats=None, kv_connector_stats=None, waiting_lora_adapters={}, running_lora_adapters={}, cudagraph_stats=None, perf_stats=None)
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] EngineCore encountered a fatal error.
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] Traceback (most recent call last):
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1360, in run_engine_core
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] engine_core.run_busy_loop()
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/fault_tolerance/engine_core_sentinel.py", line 179, in run_with_fault_tolerance
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] busy_loop_func(self)
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1419, in run_busy_loop
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] self._process_engine_step()
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1472, in _process_engine_step
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] outputs, model_executed = self.step_fn()
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] ^^^^^^^^^^^^^^
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 617, in step
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] model_output = self.model_executor.sample_tokens(grammar_output)
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 356, in sample_tokens
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] return self.collective_rpc(
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] ^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 448, in collective_rpc
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] return future if non_block else future.result()
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] ^^^^^^^^^^^^^^^
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 99, in result
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] return super().result()
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] ^^^^^^^^^^^^^^^^
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/usr/lib/python3.11/concurrent/futures/_base.py", line 449, in result
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] return self.__get_result()
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] ^^^^^^^^^^^^^^^^^^^
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/usr/lib/python3.11/concurrent/futures/_base.py", line 401, in __get_result
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] raise self._exception
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 103, in _wait_for_response
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] response = self.aggregate(self.get_response())
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] ^^^^^^^^^^^^^^^^^^^
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 437, in get_response
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] raise RuntimeError(
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] RuntimeError: Worker failed with error 'CppCompileError: C++ compile error
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376]
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] Command:
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] g++ /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp -D TORCH_INDUCTOR_CPP_WRAPPER -D STANDALONE_TORCH_HEADER -D TORCH_INDUCTOR_PRECOMPILE_HEADERS -D C10_USING_CUSTOM_GENERATED_MACROS -D CPU_CAPABILITY_AVX512 -O3 -DNDEBUG -fno-omit-frame-pointer -g1 -fno-trapping-math -funsafe-math-optimizations -ffinite-math-only -fno-signed-zeros -fno-finite-math-only -fno-unsafe-math-optimizations -fmath-errno -ffp-contract=off -fexcess-precision=fast -fno-tree-loop-vectorize -march=native -shared -fPIC -Wall -std=c++20 -Wno-unused-variable -Wno-unknown-pragmas -pedantic -fopenmp -include /tmp/torchinductor_agent/precompiled_headers/c3d7kkupnir2q25xqv5t7pw3ruepjadpjxcxa2m555qoy5o5bca4.h -I/usr/include/python3.11 -I/opt/vllm/lib/python3.11/site-packages/torch/include -I/opt/vllm/lib/python3.11/site-packages/torch/include/torch/csrc/api/include -mavx512f -mavx512dq -mavx512vl -mavx512bw -mfma -mavx512vnni -mavx512vl -o /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.so -ltorch -ltorch_cpu -ltorch_python -lgomp -L/usr/lib/x86_64-linux-gnu -L/opt/vllm/lib/python3.11/site-packages/torch/lib
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376]
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] Output:
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp:266:10: fatal error: Python.h: No such file or directory
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] 266 | #include <Python.h>
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] | ^~~~~~~~~~
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] compilation terminated.
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376]
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376]
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] Set TORCHDYNAMO_VERBOSE=1 for the internal stack trace (please do this especially if you're reporting a bug to PyTorch). For even more developer context, set TORCH_LOGS="+dynamo"
(EngineCore pid=1183) ERROR 09-19 23:41:07 [core.py:1376] ', please check the stack trace above for the root cause
(Worker pid=1294) INFO 09-19 23:41:07 [multiproc_executor.py:836] Parent process exited, terminating worker queues
(EngineCore pid=1183) INFO 09-19 23:41:07 [multiproc_executor.py:472] [shutdown] Executor: waiting for worker exit count=1
(APIServer pid=874) ERROR 09-19 23:41:07 [async_llm.py:819] AsyncLLM output_handler failed.
(APIServer pid=874) ERROR 09-19 23:41:07 [async_llm.py:819] Traceback (most recent call last):
(APIServer pid=874) ERROR 09-19 23:41:07 [async_llm.py:819] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/async_llm.py", line 765, in output_handler
(APIServer pid=874) ERROR 09-19 23:41:07 [async_llm.py:819] outputs = await engine_core.get_output_async()
(APIServer pid=874) ERROR 09-19 23:41:07 [async_llm.py:819] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=874) ERROR 09-19 23:41:07 [async_llm.py:819] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core_client.py", line 1105, in get_output_async
(APIServer pid=874) ERROR 09-19 23:41:07 [async_llm.py:819] raise self._format_exception(outputs) from None
(APIServer pid=874) ERROR 09-19 23:41:07 [async_llm.py:819] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] Error in chat completion stream generator.
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] Traceback (most recent call last):
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 516, in chat_completion_stream_generator
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] async for res in result_generator:
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/async_llm.py", line 682, in generate
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] out = q.get_nowait() or await q.get()
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] ^^^^^^^^^^^^^
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/output_processor.py", line 88, in get
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] raise output
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/async_llm.py", line 765, in output_handler
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] outputs = await engine_core.get_output_async()
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core_client.py", line 1105, in get_output_async
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] raise self._format_exception(outputs) from None
(APIServer pid=874) ERROR 09-19 23:41:07 [serving.py:906] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.
(EngineCore pid=1183) INFO 09-19 23:41:10 [multiproc_executor.py:479] [shutdown] Executor: all workers exited gracefully
(EngineCore pid=1183) Process EngineCore:
(EngineCore pid=1183) Traceback (most recent call last):
(EngineCore pid=1183) File "/usr/lib/python3.11/multiprocessing/process.py", line 314, in _bootstrap
(EngineCore pid=1183) self.run()
(EngineCore pid=1183) File "/usr/lib/python3.11/multiprocessing/process.py", line 108, in run
(EngineCore pid=1183) self._target(*self._args, **self._kwargs)
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1378, in run_engine_core
(EngineCore pid=1183) raise e
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1360, in run_engine_core
(EngineCore pid=1183) engine_core.run_busy_loop()
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/fault_tolerance/engine_core_sentinel.py", line 179, in run_with_fault_tolerance
(EngineCore pid=1183) busy_loop_func(self)
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1419, in run_busy_loop
(EngineCore pid=1183) self._process_engine_step()
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1472, in _process_engine_step
(EngineCore pid=1183) outputs, model_executed = self.step_fn()
(EngineCore pid=1183) ^^^^^^^^^^^^^^
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 617, in step
(EngineCore pid=1183) model_output = self.model_executor.sample_tokens(grammar_output)
(EngineCore pid=1183) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 356, in sample_tokens
(EngineCore pid=1183) return self.collective_rpc(
(EngineCore pid=1183) ^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 448, in collective_rpc
(EngineCore pid=1183) return future if non_block else future.result()
(EngineCore pid=1183) ^^^^^^^^^^^^^^^
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 99, in result
(EngineCore pid=1183) return super().result()
(EngineCore pid=1183) ^^^^^^^^^^^^^^^^
(EngineCore pid=1183) File "/usr/lib/python3.11/concurrent/futures/_base.py", line 449, in result
(EngineCore pid=1183) return self.__get_result()
(EngineCore pid=1183) ^^^^^^^^^^^^^^^^^^^
(EngineCore pid=1183) File "/usr/lib/python3.11/concurrent/futures/_base.py", line 401, in __get_result
(EngineCore pid=1183) raise self._exception
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 103, in _wait_for_response
(EngineCore pid=1183) response = self.aggregate(self.get_response())
(EngineCore pid=1183) ^^^^^^^^^^^^^^^^^^^
(EngineCore pid=1183) File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 437, in get_response
(EngineCore pid=1183) raise RuntimeError(
(EngineCore pid=1183) RuntimeError: Worker failed with error 'CppCompileError: C++ compile error
(EngineCore pid=1183)
(EngineCore pid=1183) Command:
(EngineCore pid=1183) g++ /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp -D TORCH_INDUCTOR_CPP_WRAPPER -D STANDALONE_TORCH_HEADER -D TORCH_INDUCTOR_PRECOMPILE_HEADERS -D C10_USING_CUSTOM_GENERATED_MACROS -D CPU_CAPABILITY_AVX512 -O3 -DNDEBUG -fno-omit-frame-pointer -g1 -fno-trapping-math -funsafe-math-optimizations -ffinite-math-only -fno-signed-zeros -fno-finite-math-only -fno-unsafe-math-optimizations -fmath-errno -ffp-contract=off -fexcess-precision=fast -fno-tree-loop-vectorize -march=native -shared -fPIC -Wall -std=c++20 -Wno-unused-variable -Wno-unknown-pragmas -pedantic -fopenmp -include /tmp/torchinductor_agent/precompiled_headers/c3d7kkupnir2q25xqv5t7pw3ruepjadpjxcxa2m555qoy5o5bca4.h -I/usr/include/python3.11 -I/opt/vllm/lib/python3.11/site-packages/torch/include -I/opt/vllm/lib/python3.11/site-packages/torch/include/torch/csrc/api/include -mavx512f -mavx512dq -mavx512vl -mavx512bw -mfma -mavx512vnni -mavx512vl -o /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.so -ltorch -ltorch_cpu -ltorch_python -lgomp -L/usr/lib/x86_64-linux-gnu -L/opt/vllm/lib/python3.11/site-packages/torch/lib
(EngineCore pid=1183)
(EngineCore pid=1183) Output:
(EngineCore pid=1183) /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp:266:10: fatal error: Python.h: No such file or directory
(EngineCore pid=1183) 266 | #include <Python.h>
(EngineCore pid=1183) | ^~~~~~~~~~
(EngineCore pid=1183) compilation terminated.
(EngineCore pid=1183)
(EngineCore pid=1183)
(EngineCore pid=1183) Set TORCHDYNAMO_VERBOSE=1 for the internal stack trace (please do this especially if you're reporting a bug to PyTorch). For even more developer context, set TORCH_LOGS="+dynamo"
(EngineCore pid=1183) ', please check the stack trace above for the root cause
(APIServer pid=874) INFO 09-19 23:41:11 [core_client.py:689] [shutdown] MPClient: start timeout=0s
(APIServer pid=874) INFO 09-19 23:41:11 [core_client.py:691] [shutdown] MPClient: stopping engine manager
(APIServer pid=874) INFO 09-19 23:41:11 [utils.py:620] [shutdown] Process manager: send sigterm to process EngineCore
(APIServer pid=874) WARNING 09-19 23:41:11 [utils.py:640] [shutdown] Process manager: force killing remaining processes count=1
(APIServer pid=874) WARNING 09-19 23:41:11 [utils.py:645] [shutdown] Process manager: force killing remaining process EngineCore pid 1183
(APIServer pid=874) INFO 09-19 23:41:11 [core_client.py:693] [shutdown] MPClient: engine manager stopped
(APIServer pid=874) INFO 09-19 23:41:11 [core_client.py:694] [shutdown] MPClient: cleaning up background resources
(APIServer pid=874) INFO 09-19 23:41:11 [core_client.py:696] [shutdown] MPClient: complete
WARNING 09-19 23:58:14 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-19 23:58:14 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-19 23:58:14 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
(APIServer pid=1390) INFO 09-19 23:58:16 [api_utils.py:347]
(APIServer pid=1390) INFO 09-19 23:58:16 [api_utils.py:347] █ █ █▄ ▄█
(APIServer pid=1390) INFO 09-19 23:58:16 [api_utils.py:347] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.29.0
(APIServer pid=1390) INFO 09-19 23:58:16 [api_utils.py:347] █▄█▀ █ █ █ █ model /data/local-agent/models/Qwen3-0.6B
(APIServer pid=1390) INFO 09-19 23:58:16 [api_utils.py:347] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
(APIServer pid=1390) INFO 09-19 23:58:16 [api_utils.py:347]
(APIServer pid=1390) INFO 09-19 23:58:16 [api_utils.py:286] non-default args: {'model_tag': '/data/local-agent/models/Qwen3-0.6B', 'enable_auto_tool_choice': True, 'tool_call_parser': 'hermes', 'host': '127.0.0.1', 'port': 43123, 'uvicorn_log_level': 'warning', 'model': '/data/local-agent/models/Qwen3-0.6B', 'dtype': 'bfloat16', 'max_model_len': 32768, 'enforce_eager': True, 'served_model_name': ['local-qwen3-0.6b'], 'generation_config': 'vllm', 'reasoning_parser': 'qwen3', 'enable_prefix_caching': True, 'max_num_batched_tokens': 2048, 'max_num_seqs': 1}
(APIServer pid=1390) INFO 09-19 23:58:23 [model.py:684] Resolved architecture: Qwen3ForCausalLM
(APIServer pid=1390) INFO 09-19 23:58:23 [model.py:2021] Using max model len 32768
(APIServer pid=1390) WARNING 09-19 23:58:24 [vllm.py:1371] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
(APIServer pid=1390) WARNING 09-19 23:58:24 [vllm.py:1406] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
(APIServer pid=1390) INFO 09-19 23:58:24 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
(APIServer pid=1390) WARNING 09-19 23:58:24 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(APIServer pid=1390) INFO 09-19 23:58:25 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
WARNING 09-19 23:58:33 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-19 23:58:33 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-19 23:58:33 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
(EngineCore pid=1613) INFO 09-19 23:58:34 [core.py:123] Initializing a V1 LLM engine (v0.29.0) with config: model='/data/local-agent/models/Qwen3-0.6B', speculative_config=None, tokenizer='/data/local-agent/models/Qwen3-0.6B', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=local-qwen3-0.6b, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto')
(EngineCore pid=1613) INFO 09-19 23:58:34 [multiproc_executor.py:153] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip=10.111.85.252 (local), world_size=1, local_world_size=1
(EngineCore pid=1613) INFO 09-19 23:58:34 [ompmultiprocessing.py:185] OpenMP thread binding info:
(EngineCore pid=1613) INFO 09-19 23:58:34 [ompmultiprocessing.py:185] VLLM_CPU_OMP_THREADS_BIND='0-1', auto_setup=False, skip_setup=False
(EngineCore pid=1613) INFO 09-19 23:58:34 [ompmultiprocessing.py:185] local_world_size=1, reserve_cpu_num=1
(EngineCore pid=1613) INFO 09-19 23:58:34 [ompmultiprocessing.py:185] local_rank=0, core ids=[0, 1]
(EngineCore pid=1613) INFO 09-19 23:58:34 [ompmultiprocessing.py:185] reserved_cpus=[]
WARNING 09-19 23:58:39 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-19 23:58:39 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-19 23:58:39 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
get_mempolicy: Operation not permitted
[W919 23:58:40.723893483 utils.cpp:41] Warning: numa_migrate_pages failed. errno: 1 (function init_cpu_memory_env)
set_mempolicy: Operation not permitted
[W919 23:58:40.723962591 utils.cpp:65] Warning: numa_set_membind failed. errno: 1 (function init_cpu_memory_env)
WARNING 09-19 23:58:40 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(Worker pid=1678) INFO 09-19 23:58:40 [parallel_state.py:1775] world_size=1 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_6fe9091ac78d40bbb35a2df08bbaf817 backend=gloo
(Worker pid=1678) INFO 09-19 23:58:40 [parallel_state.py:2119] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A, EPLB rank N/A
(Worker pid=1678) INFO 09-19 23:58:41 [cpu_model_runner.py:131] Starting to load model /data/local-agent/models/Qwen3-0.6B...
(Worker pid=1678) INFO 09-19 23:58:41 [weight_utils.py:863] Filesystem type for checkpoints: FUSE. Checkpoint size: 1.40 GiB. Available RAM: 12.61 GiB.
(Worker pid=1678) INFO 09-19 23:58:41 [weight_utils.py:886] Auto-prefetch is disabled because the filesystem (FUSE) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
(Worker pid=1678) Loading safetensors checkpoint shards: 0% Completed | 0/1 [00:00<?, ?it/s]
(Worker pid=1678) Loading safetensors checkpoint shards: 100% Completed | 1/1 [02:03<00:00, 123.21s/it]
(Worker pid=1678) Loading safetensors checkpoint shards: 100% Completed | 1/1 [02:03<00:00, 123.21s/it]
(Worker pid=1678)
(Worker pid=1678) INFO 09-20 00:00:44 [default_loader.py:430] Loading weights took 123.27 seconds
(EngineCore pid=1613) WARNING 09-20 00:00:45 [torch_utils.py:265] OMP_NUM_THREADS=2 is set; leaving Torch threads at 2 for serving. Multi-threaded torch CPU ops during serving can degrade performance through spin-wait contention and cgroup CPU-quota throttling.
(EngineCore pid=1613) INFO 09-20 00:00:45 [utils.py:306] Using LBHNC KV cache layout.
(Worker pid=1678) INFO 09-20 00:00:45 [cpu_worker.py:255] Explicitly set (4.0/14.9) GiB for KV cache on node 0.
(EngineCore pid=1613) INFO 09-20 00:00:45 [kv_cache_utils.py:2032] GPU KV cache size: 37,376 tokens, Maximum concurrency for 32,768 tokens per request: 1.14x
(EngineCore pid=1613) INFO 09-20 00:00:46 [core.py:368] init engine (profile, create kv cache, warmup model) took 1.32 s
(EngineCore pid=1613) WARNING 09-20 00:00:49 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(EngineCore pid=1613) WARNING 09-20 00:00:49 [vllm.py:1371] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
(EngineCore pid=1613) WARNING 09-20 00:00:49 [vllm.py:1406] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
(EngineCore pid=1613) INFO 09-20 00:00:49 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
(EngineCore pid=1613) INFO 09-20 00:00:49 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
(APIServer pid=1390) INFO 09-20 00:00:49 [entry.py:135] Supported tasks: ['generate']
(APIServer pid=1390) INFO 09-20 00:00:49 [parser_manager.py:36] "auto" tool choice has been enabled.
(APIServer pid=1390) INFO 09-20 00:00:50 [hf.py:547] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
(APIServer pid=1390) INFO 09-20 00:00:50 [entry.py:139] Starting vLLM server on http://127.0.0.1:43123
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:61] Available routes are:
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /openapi.json, Methods: HEAD, GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /docs, Methods: HEAD, GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /docs/oauth2-redirect, Methods: HEAD, GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /redoc, Methods: HEAD, GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /load, Methods: GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /version, Methods: GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /health, Methods: GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /metrics, Methods: GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /tokenize, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /detokenize, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/models, Methods: GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /ping, Methods: GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /ping, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /invocations, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/chat/completions, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/chat/completions/batch, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/responses, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/responses/{response_id}, Methods: GET
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/responses/{response_id}/cancel, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/completions, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/messages, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/messages/count_tokens, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /generative_scoring, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /scale_elastic_ep, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /is_scaling_elastic_ep, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/chat/completions/render, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/messages/render, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/completions/render, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/chat/completions/derender, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /v1/completions/derender, Methods: POST
(APIServer pid=1390) INFO 09-20 00:00:50 [launcher.py:70] Route: /inference/v1/generate, Methods: POST
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] Error in create_messages: This model's maximum context length is 32768 tokens. However, you requested 32000 output tokens and your prompt contains at least 769 input tokens, for a total of at least 32769 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] Traceback (most recent call last):
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/api_router.py", line 73, in create_messages
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] generator = await handler.create_messages(request, raw_request)
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/serving.py", line 610, in create_messages
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] generator = await self.create_chat_completion(chat_req, raw_request)
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 254, in create_chat_completion
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] return await self._with_kv_transfer_rejection_cleanup(
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/generate/base/serving.py", line 300, in _with_kv_transfer_rejection_cleanup
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] return await awaitable
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 275, in _create_chat_completion
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] result = await self.render_chat_request(request)
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 240, in render_chat_request
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] return await self.online_renderer.render_chat(request)
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/online_renderer.py", line 192, in render_chat
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] conversation, engine_inputs = await self.preprocess_chat(
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/online_renderer.py", line 426, in preprocess_chat
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] (conversation,), (engine_input,) = await renderer.render_chat_async(
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/base.py", line 1090, in render_chat_async
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params)
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/base.py", line 645, in tokenize_prompts_async
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] return await asyncio.gather(
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/base.py", line 638, in tokenize_prompt_async
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] return await self._tokenize_singleton_prompt_async(prompt, params)
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/base.py", line 571, in _tokenize_singleton_prompt_async
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type]
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/params.py", line 505, in apply_post_tokenization
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key]
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/params.py", line 489, in _validate_tokens
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] tokens = validator(tokenizer, tokens)
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/params.py", line 464, in _token_len_check
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] raise VLLMValidationError(
(APIServer pid=1390) ERROR 09-20 00:02:07 [api_router.py:75] vllm.exceptions.VLLMValidationError: This model's maximum context length is 32768 tokens. However, you requested 32000 output tokens and your prompt contains at least 769 input tokens, for a total of at least 32769 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] Error in create_messages: This model's maximum context length is 32768 tokens. However, you requested 32000 output tokens and your prompt contains at least 769 input tokens, for a total of at least 32769 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] Traceback (most recent call last):
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/api_router.py", line 73, in create_messages
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] generator = await handler.create_messages(request, raw_request)
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/serving.py", line 610, in create_messages
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] generator = await self.create_chat_completion(chat_req, raw_request)
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 254, in create_chat_completion
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] return await self._with_kv_transfer_rejection_cleanup(
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/generate/base/serving.py", line 300, in _with_kv_transfer_rejection_cleanup
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] return await awaitable
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 275, in _create_chat_completion
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] result = await self.render_chat_request(request)
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 240, in render_chat_request
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] return await self.online_renderer.render_chat(request)
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/online_renderer.py", line 192, in render_chat
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] conversation, engine_inputs = await self.preprocess_chat(
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/online_renderer.py", line 426, in preprocess_chat
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] (conversation,), (engine_input,) = await renderer.render_chat_async(
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/base.py", line 1090, in render_chat_async
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params)
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/base.py", line 645, in tokenize_prompts_async
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] return await asyncio.gather(
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/base.py", line 638, in tokenize_prompt_async
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] return await self._tokenize_singleton_prompt_async(prompt, params)
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/base.py", line 571, in _tokenize_singleton_prompt_async
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type]
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/params.py", line 505, in apply_post_tokenization
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key]
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/params.py", line 489, in _validate_tokens
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] tokens = validator(tokenizer, tokens)
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/renderers/params.py", line 464, in _token_len_check
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] raise VLLMValidationError(
(APIServer pid=1390) ERROR 09-20 00:05:30 [api_router.py:75] vllm.exceptions.VLLMValidationError: This model's maximum context length is 32768 tokens. However, you requested 32000 output tokens and your prompt contains at least 769 input tokens, for a total of at least 32769 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)
Fetching 10 files: 0%| | 0/10 [00:00<?, ?it/s] Fetching 10 files: 40%|████ | 4/10 [00:00<00:00, 14.83it/s] Fetching 10 files: 80%|████████ | 8/10 [00:00<00:00, 19.32it/s] Fetching 10 files: 100%|██████████| 10/10 [00:04<00:00, 2.32it/s]
✓ Downloaded
path: /data/local-agent/models/Qwen__Qwen3-0.6B
WARNING 09-20 01:15:23 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-20 01:15:23 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-20 01:15:23 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
(APIServer pid=4269) INFO 09-20 01:15:25 [api_utils.py:347]
(APIServer pid=4269) INFO 09-20 01:15:25 [api_utils.py:347] █ █ █▄ ▄█
(APIServer pid=4269) INFO 09-20 01:15:25 [api_utils.py:347] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.29.0
(APIServer pid=4269) INFO 09-20 01:15:25 [api_utils.py:347] █▄█▀ █ █ █ █ model /data/local-agent/models/Qwen__Qwen3-0.6B
(APIServer pid=4269) INFO 09-20 01:15:25 [api_utils.py:347] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
(APIServer pid=4269) INFO 09-20 01:15:25 [api_utils.py:347]
(APIServer pid=4269) INFO 09-20 01:15:25 [api_utils.py:286] non-default args: {'model_tag': '/data/local-agent/models/Qwen__Qwen3-0.6B', 'enable_auto_tool_choice': True, 'tool_call_parser': 'hermes', 'host': '127.0.0.1', 'port': 43123, 'uvicorn_log_level': 'warning', 'model': '/data/local-agent/models/Qwen__Qwen3-0.6B', 'dtype': 'bfloat16', 'max_model_len': 65536, 'enforce_eager': True, 'served_model_name': ['local-qwen3-0.6b'], 'hf_overrides': {'rope_scaling': {'rope_type': 'yarn', 'factor': 2.0, 'original_max_position_embeddings': 32768}}, 'generation_config': 'vllm', 'reasoning_parser': 'qwen3', 'kv_cache_memory_bytes': 4026531840, 'kv_cache_dtype': 'fp8', 'enable_prefix_caching': True, 'max_num_batched_tokens': 2048, 'max_num_seqs': 1}
(APIServer pid=4269) INFO 09-20 01:15:32 [model.py:684] Resolved architecture: Qwen3ForCausalLM
(APIServer pid=4269) INFO 09-20 01:15:32 [model.py:2021] Using max model len 65536
(APIServer pid=4269) INFO 09-20 01:15:32 [cache.py:335] Using fp8 data type to store kv cache. It reduces the GPU memory footprint and boosts the performance. Meanwhile, it may cause accuracy drop without a proper scaling factor
(APIServer pid=4269) WARNING 09-20 01:15:32 [vllm.py:1371] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
(APIServer pid=4269) WARNING 09-20 01:15:32 [vllm.py:1406] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
(APIServer pid=4269) INFO 09-20 01:15:32 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
(APIServer pid=4269) WARNING 09-20 01:15:32 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(APIServer pid=4269) INFO 09-20 01:15:33 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
WARNING 09-20 01:15:40 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-20 01:15:40 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-20 01:15:40 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
(EngineCore pid=4461) INFO 09-20 01:15:41 [core.py:123] Initializing a V1 LLM engine (v0.29.0) with config: model='/data/local-agent/models/Qwen__Qwen3-0.6B', speculative_config=None, tokenizer='/data/local-agent/models/Qwen__Qwen3-0.6B', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=65536, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=fp8, device_config=cpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=local-qwen3-0.6b, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto')
(EngineCore pid=4461) INFO 09-20 01:15:41 [multiproc_executor.py:153] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip=10.111.72.214 (local), world_size=1, local_world_size=1
(EngineCore pid=4461) INFO 09-20 01:15:41 [ompmultiprocessing.py:185] OpenMP thread binding info:
(EngineCore pid=4461) INFO 09-20 01:15:41 [ompmultiprocessing.py:185] VLLM_CPU_OMP_THREADS_BIND='0-1', auto_setup=False, skip_setup=False
(EngineCore pid=4461) INFO 09-20 01:15:41 [ompmultiprocessing.py:185] local_world_size=1, reserve_cpu_num=1
(EngineCore pid=4461) INFO 09-20 01:15:41 [ompmultiprocessing.py:185] local_rank=0, core ids=[0, 1]
(EngineCore pid=4461) INFO 09-20 01:15:41 [ompmultiprocessing.py:185] reserved_cpus=[]
WARNING 09-20 01:15:46 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-20 01:15:46 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-20 01:15:46 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
get_mempolicy: Operation not permitted
[W920 01:15:47.211343620 utils.cpp:41] Warning: numa_migrate_pages failed. errno: 1 (function init_cpu_memory_env)
set_mempolicy: Operation not permitted
[W920 01:15:47.211453140 utils.cpp:65] Warning: numa_set_membind failed. errno: 1 (function init_cpu_memory_env)
WARNING 09-20 01:15:47 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(Worker pid=4526) INFO 09-20 01:15:47 [parallel_state.py:1775] world_size=1 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_d0ea46b75d4d492d9490bf63f7fc9ec8 backend=gloo
(Worker pid=4526) INFO 09-20 01:15:47 [parallel_state.py:2119] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A, EPLB rank N/A
(Worker pid=4526) INFO 09-20 01:15:47 [cpu_model_runner.py:131] Starting to load model /data/local-agent/models/Qwen__Qwen3-0.6B...
(Worker pid=4526) INFO 09-20 01:15:47 [weight_utils.py:863] Filesystem type for checkpoints: FUSE. Checkpoint size: 1.40 GiB. Available RAM: 0.71 GiB.
(Worker pid=4526) INFO 09-20 01:15:47 [weight_utils.py:893] Auto-prefetch is disabled because the filesystem (FUSE) is not a recognized network FS (NFS/Lustre) and the checkpoint size (1.40 GiB) exceeds 90% of available RAM (0.71 GiB).
(Worker pid=4526) Loading safetensors checkpoint shards: 0% Completed | 0/1 [00:00<?, ?it/s]
(Worker pid=4526) Loading safetensors checkpoint shards: 100% Completed | 1/1 [02:08<00:00, 128.18s/it]
(Worker pid=4526) Loading safetensors checkpoint shards: 100% Completed | 1/1 [02:08<00:00, 128.18s/it]
(Worker pid=4526)
(Worker pid=4526) INFO 09-20 01:17:56 [default_loader.py:430] Loading weights took 128.23 seconds
(EngineCore pid=4461) WARNING 09-20 01:17:56 [torch_utils.py:265] OMP_NUM_THREADS=2 is set; leaving Torch threads at 2 for serving. Multi-threaded torch CPU ops during serving can degrade performance through spin-wait contention and cgroup CPU-quota throttling.
(EngineCore pid=4461) INFO 09-20 01:17:56 [utils.py:306] Using LBHNC KV cache layout.
(Worker pid=4526) INFO 09-20 01:17:56 [cpu_worker.py:255] Explicitly set (3.75/14.9) GiB for KV cache on node 0.
(EngineCore pid=4461) INFO 09-20 01:17:56 [kv_cache_utils.py:2032] GPU KV cache size: 70,144 tokens, Maximum concurrency for 65,536 tokens per request: 1.07x
(EngineCore pid=4461) INFO 09-20 01:17:58 [core.py:368] init engine (profile, create kv cache, warmup model) took 1.25 s
(EngineCore pid=4461) WARNING 09-20 01:17:59 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(EngineCore pid=4461) WARNING 09-20 01:17:59 [vllm.py:1371] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
(EngineCore pid=4461) WARNING 09-20 01:17:59 [vllm.py:1406] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
(EngineCore pid=4461) INFO 09-20 01:17:59 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
(EngineCore pid=4461) INFO 09-20 01:17:59 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
(APIServer pid=4269) INFO 09-20 01:17:59 [entry.py:135] Supported tasks: ['generate']
(APIServer pid=4269) INFO 09-20 01:18:00 [parser_manager.py:36] "auto" tool choice has been enabled.
(APIServer pid=4269) INFO 09-20 01:18:00 [hf.py:547] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
(APIServer pid=4269) INFO 09-20 01:18:01 [entry.py:139] Starting vLLM server on http://127.0.0.1:43123
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:61] Available routes are:
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /openapi.json, Methods: HEAD, GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /docs, Methods: HEAD, GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /docs/oauth2-redirect, Methods: HEAD, GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /redoc, Methods: HEAD, GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /load, Methods: GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /version, Methods: GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /health, Methods: GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /metrics, Methods: GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /tokenize, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /detokenize, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/models, Methods: GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /ping, Methods: GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /ping, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /invocations, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/chat/completions, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/chat/completions/batch, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/responses, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/responses/{response_id}, Methods: GET
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/responses/{response_id}/cancel, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/completions, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/messages, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/messages/count_tokens, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /generative_scoring, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /scale_elastic_ep, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /is_scaling_elastic_ep, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/chat/completions/render, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/messages/render, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/completions/render, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/chat/completions/derender, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /v1/completions/derender, Methods: POST
(APIServer pid=4269) INFO 09-20 01:18:01 [launcher.py:70] Route: /inference/v1/generate, Methods: POST
WARNING 09-20 02:36:16 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-20 02:36:16 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-20 02:36:16 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
(APIServer pid=2949) INFO 09-20 02:36:18 [api_utils.py:347]
(APIServer pid=2949) INFO 09-20 02:36:18 [api_utils.py:347] █ █ █▄ ▄█
(APIServer pid=2949) INFO 09-20 02:36:18 [api_utils.py:347] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.29.0
(APIServer pid=2949) INFO 09-20 02:36:18 [api_utils.py:347] █▄█▀ █ █ █ █ model /data/local-agent/models/Qwen__Qwen3-0.6B
(APIServer pid=2949) INFO 09-20 02:36:18 [api_utils.py:347] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
(APIServer pid=2949) INFO 09-20 02:36:18 [api_utils.py:347]
(APIServer pid=2949) INFO 09-20 02:36:18 [api_utils.py:286] non-default args: {'model_tag': '/data/local-agent/models/Qwen__Qwen3-0.6B', 'enable_auto_tool_choice': True, 'tool_call_parser': 'hermes', 'host': '127.0.0.1', 'port': 43123, 'uvicorn_log_level': 'warning', 'model': '/data/local-agent/models/Qwen__Qwen3-0.6B', 'dtype': 'bfloat16', 'max_model_len': 65536, 'enforce_eager': True, 'served_model_name': ['local-qwen3-0.6b'], 'hf_overrides': {'rope_scaling': {'rope_type': 'yarn', 'factor': 2.0, 'original_max_position_embeddings': 32768}}, 'generation_config': 'vllm', 'reasoning_parser': 'qwen3', 'kv_cache_memory_bytes': 4026531840, 'kv_cache_dtype': 'fp8', 'enable_prefix_caching': True, 'max_num_batched_tokens': 2048, 'max_num_seqs': 1}
(APIServer pid=2949) INFO 09-20 02:36:27 [model.py:684] Resolved architecture: Qwen3ForCausalLM
(APIServer pid=2949) INFO 09-20 02:36:27 [model.py:2021] Using max model len 65536
(APIServer pid=2949) INFO 09-20 02:36:27 [cache.py:335] Using fp8 data type to store kv cache. It reduces the GPU memory footprint and boosts the performance. Meanwhile, it may cause accuracy drop without a proper scaling factor
(APIServer pid=2949) WARNING 09-20 02:36:28 [vllm.py:1371] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
(APIServer pid=2949) WARNING 09-20 02:36:28 [vllm.py:1406] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
(APIServer pid=2949) INFO 09-20 02:36:28 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
(APIServer pid=2949) WARNING 09-20 02:36:28 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(APIServer pid=2949) INFO 09-20 02:36:29 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
WARNING 09-20 02:36:40 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-20 02:36:40 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-20 02:36:40 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
(EngineCore pid=3190) INFO 09-20 02:36:41 [core.py:123] Initializing a V1 LLM engine (v0.29.0) with config: model='/data/local-agent/models/Qwen__Qwen3-0.6B', speculative_config=None, tokenizer='/data/local-agent/models/Qwen__Qwen3-0.6B', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=65536, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=fp8, device_config=cpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=local-qwen3-0.6b, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto')
(EngineCore pid=3190) INFO 09-20 02:36:41 [multiproc_executor.py:153] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip=10.111.114.63 (local), world_size=1, local_world_size=1
(EngineCore pid=3190) INFO 09-20 02:36:41 [ompmultiprocessing.py:185] OpenMP thread binding info:
(EngineCore pid=3190) INFO 09-20 02:36:41 [ompmultiprocessing.py:185] VLLM_CPU_OMP_THREADS_BIND='0-1', auto_setup=False, skip_setup=False
(EngineCore pid=3190) INFO 09-20 02:36:41 [ompmultiprocessing.py:185] local_world_size=1, reserve_cpu_num=1
(EngineCore pid=3190) INFO 09-20 02:36:41 [ompmultiprocessing.py:185] local_rank=0, core ids=[0, 1]
(EngineCore pid=3190) INFO 09-20 02:36:41 [ompmultiprocessing.py:185] reserved_cpus=[]
WARNING 09-20 02:36:48 [importing.py:45] Triton is installed, but doesn't include CPU backend. Disabling Triton.
INFO 09-20 02:36:48 [importing.py:74] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 09-20 02:36:48 [importing.py:98] Triton not installed or not compatible; certain GPU-related functions will not be available.
get_mempolicy: Operation not permitted
[W920 02:36:49.738867355 utils.cpp:41] Warning: numa_migrate_pages failed. errno: 1 (function init_cpu_memory_env)
set_mempolicy: Operation not permitted
[W920 02:36:49.738974755 utils.cpp:65] Warning: numa_set_membind failed. errno: 1 (function init_cpu_memory_env)
WARNING 09-20 02:36:49 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(Worker pid=3277) INFO 09-20 02:36:49 [parallel_state.py:1775] world_size=1 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_31d4d5c5b6b744a9a97a0b47a41d70e2 backend=gloo
(Worker pid=3277) INFO 09-20 02:36:49 [parallel_state.py:2119] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A, EPLB rank N/A
(Worker pid=3277) INFO 09-20 02:36:49 [cpu_model_runner.py:131] Starting to load model /data/local-agent/models/Qwen__Qwen3-0.6B...
(Worker pid=3277) INFO 09-20 02:36:50 [weight_utils.py:863] Filesystem type for checkpoints: FUSE. Checkpoint size: 1.40 GiB. Available RAM: 6.78 GiB.
(Worker pid=3277) INFO 09-20 02:36:50 [weight_utils.py:886] Auto-prefetch is disabled because the filesystem (FUSE) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
(Worker pid=3277) Loading safetensors checkpoint shards: 0% Completed | 0/1 [00:00<?, ?it/s]
(Worker pid=3277) Loading safetensors checkpoint shards: 100% Completed | 1/1 [02:07<00:00, 127.61s/it]
(Worker pid=3277) Loading safetensors checkpoint shards: 100% Completed | 1/1 [02:07<00:00, 127.62s/it]
(Worker pid=3277)
(Worker pid=3277) INFO 09-20 02:38:58 [default_loader.py:430] Loading weights took 127.68 seconds
(EngineCore pid=3190) WARNING 09-20 02:38:58 [torch_utils.py:265] OMP_NUM_THREADS=2 is set; leaving Torch threads at 2 for serving. Multi-threaded torch CPU ops during serving can degrade performance through spin-wait contention and cgroup CPU-quota throttling.
(EngineCore pid=3190) INFO 09-20 02:38:58 [utils.py:306] Using LBHNC KV cache layout.
(Worker pid=3277) INFO 09-20 02:38:58 [cpu_worker.py:255] Explicitly set (3.75/14.9) GiB for KV cache on node 0.
(EngineCore pid=3190) INFO 09-20 02:38:58 [kv_cache_utils.py:2032] GPU KV cache size: 70,144 tokens, Maximum concurrency for 65,536 tokens per request: 1.07x
(EngineCore pid=3190) INFO 09-20 02:39:00 [core.py:368] init engine (profile, create kv cache, warmup model) took 1.23 s
(EngineCore pid=3190) WARNING 09-20 02:39:02 [vllm.py:663] Model Runner V2 requires Triton; using the V1 model runner instead.
(EngineCore pid=3190) WARNING 09-20 02:39:02 [vllm.py:1371] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
(EngineCore pid=3190) WARNING 09-20 02:39:02 [vllm.py:1406] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
(EngineCore pid=3190) INFO 09-20 02:39:02 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
(EngineCore pid=3190) INFO 09-20 02:39:02 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
(APIServer pid=2949) INFO 09-20 02:39:02 [entry.py:135] Supported tasks: ['generate']
(APIServer pid=2949) INFO 09-20 02:39:02 [parser_manager.py:36] "auto" tool choice has been enabled.
(APIServer pid=2949) INFO 09-20 02:39:03 [hf.py:547] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
(APIServer pid=2949) INFO 09-20 02:39:03 [entry.py:139] Starting vLLM server on http://127.0.0.1:43123
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:61] Available routes are:
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /openapi.json, Methods: HEAD, GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /docs, Methods: HEAD, GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /docs/oauth2-redirect, Methods: HEAD, GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /redoc, Methods: HEAD, GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /load, Methods: GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /version, Methods: GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /health, Methods: GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /metrics, Methods: GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /tokenize, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /detokenize, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/models, Methods: GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /ping, Methods: GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /ping, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /invocations, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/chat/completions, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/chat/completions/batch, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/responses, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/responses/{response_id}, Methods: GET
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/responses/{response_id}/cancel, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/completions, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/messages, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/messages/count_tokens, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /generative_scoring, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /scale_elastic_ep, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /is_scaling_elastic_ep, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/chat/completions/render, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/messages/render, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/completions/render, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/chat/completions/derender, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /v1/completions/derender, Methods: POST
(APIServer pid=2949) INFO 09-20 02:39:03 [launcher.py:70] Route: /inference/v1/generate, Methods: POST
(Worker pid=3277) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=3277) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=3277) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=3277) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=3277) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=3277) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=3277) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=3277) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=3277) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=3277) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=3277) /opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py:1361: UserWarning: Can't find Python.h in /usr/include/python3.11
(Worker pid=3277) warnings.warn(f"Can't find Python.h in {str(include_dir)}")
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] WorkerProc hit an exception.
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] Traceback (most recent call last):
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 1047, in _execute_worker_rpc
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] output = func(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/utils/_contextlib.py", line 124, in decorate_context
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return func(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_worker.py", line 1105, in sample_tokens
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return self.model_runner.sample_tokens(grammar_output)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/utils/_contextlib.py", line 124, in decorate_context
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return func(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_model_runner.py", line 4664, in sample_tokens
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] sampler_output = self._sample(logits, spec_decode_metadata)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_model_runner.py", line 3731, in _sample
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return self.sampler(
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return self._call_impl(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1789, in _call_impl
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return forward_call(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/sampler.py", line 103, in forward
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] sampled, processed_logprobs = self.sample(logits, sampling_metadata)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/sampler.py", line 287, in sample
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] random_sampled, processed_logprobs = self.topk_topp_sampler(
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return self._call_impl(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1789, in _call_impl
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return forward_call(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/ops/topk_topp_sampler.py", line 204, in forward_cpu
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return compiled_random_sample(logits), logits_to_return
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_dynamo/eval_frame.py", line 1183, in compile_wrapper
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] raise e.remove_dynamo_frames() from None # see TORCHDYNAMO_VERBOSE=1
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1079, in _compile_fx_inner
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] raise InductorError(e, currentframe()).with_traceback(
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1059, in _compile_fx_inner
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] mb_compiled_graph = fx_codegen_and_compile(
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1847, in fx_codegen_and_compile
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return scheme.codegen_and_compile(gm, example_inputs, inputs_to_check, graph_kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1608, in codegen_and_compile
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] compiled_module = graph.compile_to_module()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2669, in compile_to_module
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return self._compile_to_module()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2679, in _compile_to_module
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] mod = self._compile_to_module_lines(wrapper_code)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2754, in _compile_to_module_lines
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] mod = PyCodeCache.load_by_key_path(
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4622, in load_by_key_path
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] mod = _reload_python_module(key, path, set_sys_modules=set_sys_modules)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/runtime/compile_tasks.py", line 35, in _reload_python_module
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] exec(code, mod.__dict__, mod.__dict__)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/tmp/torchinductor_agent/hu/chuaszb6rgufguvuy3xrzwcgcm6eug2vmkmnvpi3hie6u35dza3e.py", line 33, in <module>
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] cpp_fused__softmax_argmax_div_exponential_0 = async_compile.cpp_pybinding(['const float*', 'const int64_t*', 'float*', 'float*', 'float*', 'int64_t*', 'const int64_t'], r'''
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/async_compile.py", line 570, in cpp_pybinding
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return CppPythonBindingsCodeCache.load_pybinding(argtypes, source_code)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4084, in load_pybinding
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return cls.load_pybinding_async(*args, **kwargs)()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4076, in future
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] result = get_result()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 3853, in load_fn
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] result = worker_fn()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 3882, in _worker_compile_cpp
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] builder.build()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 2571, in build
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] run_compile_cmd(build_cmd, cwd=_build_tmp_dir)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 718, in run_compile_cmd
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] _run_compile_cmd(cmd_line, cwd)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 713, in _run_compile_cmd
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] raise exc.CppCompileError(cmd, output) from e
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] torch._inductor.exc.InductorError: CppCompileError: C++ compile error
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] Command:
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] g++ /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp -D TORCH_INDUCTOR_CPP_WRAPPER -D STANDALONE_TORCH_HEADER -D TORCH_INDUCTOR_PRECOMPILE_HEADERS -D C10_USING_CUSTOM_GENERATED_MACROS -D CPU_CAPABILITY_AVX512 -O3 -DNDEBUG -fno-omit-frame-pointer -g1 -fno-trapping-math -funsafe-math-optimizations -ffinite-math-only -fno-signed-zeros -fno-finite-math-only -fno-unsafe-math-optimizations -fmath-errno -ffp-contract=off -fexcess-precision=fast -fno-tree-loop-vectorize -march=native -shared -fPIC -Wall -std=c++20 -Wno-unused-variable -Wno-unknown-pragmas -pedantic -fopenmp -include /tmp/torchinductor_agent/precompiled_headers/c3d7kkupnir2q25xqv5t7pw3ruepjadpjxcxa2m555qoy5o5bca4.h -I/usr/include/python3.11 -I/opt/vllm/lib/python3.11/site-packages/torch/include -I/opt/vllm/lib/python3.11/site-packages/torch/include/torch/csrc/api/include -mavx512f -mavx512dq -mavx512vl -mavx512bw -mfma -mavx512vnni -mavx512vl -o /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.so -ltorch -ltorch_cpu -ltorch_python -lgomp -L/usr/lib/x86_64-linux-gnu -L/opt/vllm/lib/python3.11/site-packages/torch/lib
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] Output:
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp:266:10: fatal error: Python.h: No such file or directory
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] 266 | #include <Python.h>
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] | ^~~~~~~~~~
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] compilation terminated.
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] Set TORCHDYNAMO_VERBOSE=1 for the internal stack trace (please do this especially if you're reporting a bug to PyTorch). For even more developer context, set TORCH_LOGS="+dynamo"
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] Traceback (most recent call last):
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 1047, in _execute_worker_rpc
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] output = func(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/utils/_contextlib.py", line 124, in decorate_context
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return func(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_worker.py", line 1105, in sample_tokens
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return self.model_runner.sample_tokens(grammar_output)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/utils/_contextlib.py", line 124, in decorate_context
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return func(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_model_runner.py", line 4664, in sample_tokens
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] sampler_output = self._sample(logits, spec_decode_metadata)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/worker/gpu_model_runner.py", line 3731, in _sample
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return self.sampler(
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return self._call_impl(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1789, in _call_impl
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return forward_call(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/sampler.py", line 103, in forward
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] sampled, processed_logprobs = self.sample(logits, sampling_metadata)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/sampler.py", line 287, in sample
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] random_sampled, processed_logprobs = self.topk_topp_sampler(
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return self._call_impl(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1789, in _call_impl
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return forward_call(*args, **kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/sample/ops/topk_topp_sampler.py", line 204, in forward_cpu
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return compiled_random_sample(logits), logits_to_return
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_dynamo/eval_frame.py", line 1183, in compile_wrapper
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] raise e.remove_dynamo_frames() from None # see TORCHDYNAMO_VERBOSE=1
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1079, in _compile_fx_inner
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] raise InductorError(e, currentframe()).with_traceback(
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1059, in _compile_fx_inner
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] mb_compiled_graph = fx_codegen_and_compile(
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1847, in fx_codegen_and_compile
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return scheme.codegen_and_compile(gm, example_inputs, inputs_to_check, graph_kwargs)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/compile_fx.py", line 1608, in codegen_and_compile
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] compiled_module = graph.compile_to_module()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2669, in compile_to_module
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return self._compile_to_module()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2679, in _compile_to_module
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] mod = self._compile_to_module_lines(wrapper_code)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/graph.py", line 2754, in _compile_to_module_lines
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] mod = PyCodeCache.load_by_key_path(
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4622, in load_by_key_path
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] mod = _reload_python_module(key, path, set_sys_modules=set_sys_modules)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/runtime/compile_tasks.py", line 35, in _reload_python_module
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] exec(code, mod.__dict__, mod.__dict__)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/tmp/torchinductor_agent/hu/chuaszb6rgufguvuy3xrzwcgcm6eug2vmkmnvpi3hie6u35dza3e.py", line 33, in <module>
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] cpp_fused__softmax_argmax_div_exponential_0 = async_compile.cpp_pybinding(['const float*', 'const int64_t*', 'float*', 'float*', 'float*', 'int64_t*', 'const int64_t'], r'''
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/async_compile.py", line 570, in cpp_pybinding
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return CppPythonBindingsCodeCache.load_pybinding(argtypes, source_code)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4084, in load_pybinding
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] return cls.load_pybinding_async(*args, **kwargs)()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 4076, in future
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] result = get_result()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 3853, in load_fn
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] result = worker_fn()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] ^^^^^^^^^^^
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/codecache.py", line 3882, in _worker_compile_cpp
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] builder.build()
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 2571, in build
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] run_compile_cmd(build_cmd, cwd=_build_tmp_dir)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 718, in run_compile_cmd
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] _run_compile_cmd(cmd_line, cwd)
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] File "/opt/vllm/lib/python3.11/site-packages/torch/_inductor/cpp_builder.py", line 713, in _run_compile_cmd
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] raise exc.CppCompileError(cmd, output) from e
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] torch._inductor.exc.InductorError: CppCompileError: C++ compile error
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] Command:
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] g++ /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp -D TORCH_INDUCTOR_CPP_WRAPPER -D STANDALONE_TORCH_HEADER -D TORCH_INDUCTOR_PRECOMPILE_HEADERS -D C10_USING_CUSTOM_GENERATED_MACROS -D CPU_CAPABILITY_AVX512 -O3 -DNDEBUG -fno-omit-frame-pointer -g1 -fno-trapping-math -funsafe-math-optimizations -ffinite-math-only -fno-signed-zeros -fno-finite-math-only -fno-unsafe-math-optimizations -fmath-errno -ffp-contract=off -fexcess-precision=fast -fno-tree-loop-vectorize -march=native -shared -fPIC -Wall -std=c++20 -Wno-unused-variable -Wno-unknown-pragmas -pedantic -fopenmp -include /tmp/torchinductor_agent/precompiled_headers/c3d7kkupnir2q25xqv5t7pw3ruepjadpjxcxa2m555qoy5o5bca4.h -I/usr/include/python3.11 -I/opt/vllm/lib/python3.11/site-packages/torch/include -I/opt/vllm/lib/python3.11/site-packages/torch/include/torch/csrc/api/include -mavx512f -mavx512dq -mavx512vl -mavx512bw -mfma -mavx512vnni -mavx512vl -o /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.so -ltorch -ltorch_cpu -ltorch_python -lgomp -L/usr/lib/x86_64-linux-gnu -L/opt/vllm/lib/python3.11/site-packages/torch/lib
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] Output:
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp:266:10: fatal error: Python.h: No such file or directory
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] 266 | #include <Python.h>
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] | ^~~~~~~~~~
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] compilation terminated.
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055] Set TORCHDYNAMO_VERBOSE=1 for the internal stack trace (please do this especially if you're reporting a bug to PyTorch). For even more developer context, set TORCH_LOGS="+dynamo"
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(Worker pid=3277) ERROR 09-20 02:43:06 [multiproc_executor.py:1055]
(EngineCore pid=3190) ERROR 09-20 02:43:06 [dump_input.py:72] Dumping input data for V1 LLM engine (v0.29.0) with config: model='/data/local-agent/models/Qwen__Qwen3-0.6B', speculative_config=None, tokenizer='/data/local-agent/models/Qwen__Qwen3-0.6B', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=65536, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=fp8, device_config=cpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=local-qwen3-0.6b, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto'),
(EngineCore pid=3190) ERROR 09-20 02:43:06 [dump_input.py:79] Dumping scheduler output for model execution: SchedulerOutput(scheduled_new_reqs=[NewRequestData(req_id=chatcmpl-8bf01c222d9df789-95df03f2,prompt_token_ids_len=19407,prefill_token_ids_len=None,mm_features=[],sampling_params=SamplingParams(n=1, presence_penalty=0.0, frequency_penalty=0.0, repetition_penalty=1.0, temperature=1.0, top_p=1.0, top_k=0, min_p=0.0, seed=None, stop=[], stop_token_ids=[151643], bad_words=[], thinking_token_budget=None, include_stop_str_in_output=False, ignore_eos=False, max_tokens=32000, min_tokens=0, logprobs=None, prompt_logprobs=None, skip_special_tokens=False, spaces_between_special_tokens=True, structured_outputs=None, extra_args=None),block_ids=([1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16],),num_computed_tokens=0,lora_request=None,prompt_embeds_shape=None)], scheduled_cached_reqs=CachedRequestData(req_ids=[],resumed_req_ids=set(),new_token_ids_lens=[],all_token_ids_lens={},new_block_ids=[],num_computed_tokens=[],num_output_tokens=[]), num_scheduled_tokens={chatcmpl-8bf01c222d9df789-95df03f2: 2048}, total_num_scheduled_tokens=2048, scheduled_spec_decode_tokens={}, scheduled_encoder_inputs={}, num_common_prefix_blocks=[16], finished_req_ids=[], free_encoder_mm_hashes=[], scheduled_encoder_input_stats=null, preempted_req_ids=[], has_structured_output_requests=false, pending_structured_output_tokens=false, num_invalid_spec_tokens=null, kv_connector_metadata=null, has_sync_kv_loads=false, ec_connector_metadata=null, ec_manager_metadata=null, new_block_ids_to_zero=null, kv_cache_block_copies=null, kv_connector_block_state=null, num_spec_tokens_to_schedule=0)
(EngineCore pid=3190) ERROR 09-20 02:43:06 [dump_input.py:81] Dumping scheduler stats: SchedulerStats(num_running_reqs=1, num_waiting_reqs=0, num_skipped_waiting_reqs=0, step_counter=0, current_wave=0, kv_cache_usage=0.029250457038391242, iteration_details=None, prefix_cache_stats=PrefixCacheStats(reset=False, requests=1, queries=19407, hits=0, preempted_requests=0, preempted_queries=0, preempted_hits=0), connector_prefix_cache_stats=None, kv_cache_eviction_events=[], spec_decoding_stats=None, kv_connector_stats=None, waiting_lora_adapters={}, running_lora_adapters={}, cudagraph_stats=None, perf_stats=None)
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] EngineCore encountered a fatal error.
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] Traceback (most recent call last):
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1360, in run_engine_core
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] engine_core.run_busy_loop()
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/fault_tolerance/engine_core_sentinel.py", line 179, in run_with_fault_tolerance
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] busy_loop_func(self)
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1419, in run_busy_loop
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] self._process_engine_step()
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 1472, in _process_engine_step
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] outputs, model_executed = self.step_fn()
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] ^^^^^^^^^^^^^^
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 617, in step
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] model_output = self.model_executor.sample_tokens(grammar_output)
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 356, in sample_tokens
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] return self.collective_rpc(
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] ^^^^^^^^^^^^^^^^^^^^
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 448, in collective_rpc
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] return future if non_block else future.result()
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] ^^^^^^^^^^^^^^^
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 99, in result
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] return super().result()
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] ^^^^^^^^^^^^^^^^
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/usr/lib/python3.11/concurrent/futures/_base.py", line 449, in result
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] return self.__get_result()
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] ^^^^^^^^^^^^^^^^^^^
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/usr/lib/python3.11/concurrent/futures/_base.py", line 401, in __get_result
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] raise self._exception
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 103, in _wait_for_response
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] response = self.aggregate(self.get_response())
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] ^^^^^^^^^^^^^^^^^^^
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/executor/multiproc_executor.py", line 437, in get_response
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] raise RuntimeError(
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] RuntimeError: Worker failed with error 'CppCompileError: C++ compile error
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376]
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] Command:
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] g++ /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp -D TORCH_INDUCTOR_CPP_WRAPPER -D STANDALONE_TORCH_HEADER -D TORCH_INDUCTOR_PRECOMPILE_HEADERS -D C10_USING_CUSTOM_GENERATED_MACROS -D CPU_CAPABILITY_AVX512 -O3 -DNDEBUG -fno-omit-frame-pointer -g1 -fno-trapping-math -funsafe-math-optimizations -ffinite-math-only -fno-signed-zeros -fno-finite-math-only -fno-unsafe-math-optimizations -fmath-errno -ffp-contract=off -fexcess-precision=fast -fno-tree-loop-vectorize -march=native -shared -fPIC -Wall -std=c++20 -Wno-unused-variable -Wno-unknown-pragmas -pedantic -fopenmp -include /tmp/torchinductor_agent/precompiled_headers/c3d7kkupnir2q25xqv5t7pw3ruepjadpjxcxa2m555qoy5o5bca4.h -I/usr/include/python3.11 -I/opt/vllm/lib/python3.11/site-packages/torch/include -I/opt/vllm/lib/python3.11/site-packages/torch/include/torch/csrc/api/include -mavx512f -mavx512dq -mavx512vl -mavx512bw -mfma -mavx512vnni -mavx512vl -o /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.so -ltorch -ltorch_cpu -ltorch_python -lgomp -L/usr/lib/x86_64-linux-gnu -L/opt/vllm/lib/python3.11/site-packages/torch/lib
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376]
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] Output:
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] /tmp/torchinductor_agent/af/cafpbp4rm3egecxdyphk5pgiy6wpo263bwpa23ctanu2wvejsmlu.main.cpp:266:10: fatal error: Python.h: No such file or directory
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] 266 | #include <Python.h>
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] | ^~~~~~~~~~
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] compilation terminated.
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376]
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376]
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] Set TORCHDYNAMO_VERBOSE=1 for the internal stack trace (please do this especially if you're reporting a bug to PyTorch). For even more developer context, set TORCH_LOGS="+dynamo"
(EngineCore pid=3190) ERROR 09-20 02:43:06 [core.py:1376] ', please check the stack trace above for the root cause
(Worker pid=3277) INFO 09-20 02:43:06 [multiproc_executor.py:836] Parent process exited, terminating worker queues
(EngineCore pid=3190) INFO 09-20 02:43:06 [multiproc_executor.py:472] [shutdown] Executor: waiting for worker exit count=1
(APIServer pid=2949) ERROR 09-20 02:43:06 [async_llm.py:819] AsyncLLM output_handler failed.
(APIServer pid=2949) ERROR 09-20 02:43:06 [async_llm.py:819] Traceback (most recent call last):
(APIServer pid=2949) ERROR 09-20 02:43:06 [async_llm.py:819] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/async_llm.py", line 765, in output_handler
(APIServer pid=2949) ERROR 09-20 02:43:06 [async_llm.py:819] outputs = await engine_core.get_output_async()
(APIServer pid=2949) ERROR 09-20 02:43:06 [async_llm.py:819] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:06 [async_llm.py:819] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core_client.py", line 1105, in get_output_async
(APIServer pid=2949) ERROR 09-20 02:43:06 [async_llm.py:819] raise self._format_exception(outputs) from None
(APIServer pid=2949) ERROR 09-20 02:43:06 [async_llm.py:819] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] Error in chat completion stream generator.
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] Traceback (most recent call last):
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 516, in chat_completion_stream_generator
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] async for res in result_generator:
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/async_llm.py", line 682, in generate
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] out = q.get_nowait() or await q.get()
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] ^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/output_processor.py", line 88, in get
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] raise output
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/async_llm.py", line 765, in output_handler
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] outputs = await engine_core.get_output_async()
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] File "/opt/vllm/lib/python3.11/site-packages/vllm/v1/engine/core_client.py", line 1105, in get_output_async
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] raise self._format_exception(outputs) from None
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:906] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] Error in message stream converter.
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] Traceback (most recent call last):
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/serving.py", line 804, in message_stream_converter
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] origin_chunk = ChatCompletionStreamResponse.model_validate_json(
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] File "/opt/vllm/lib/python3.11/site-packages/pydantic/main.py", line 782, in model_validate_json
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] return cls.__pydantic_validator__.validate_json(
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] pydantic_core._pydantic_core.ValidationError: 2 validation errors for ChatCompletionStreamResponse
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] model
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] Field required [type=missing, input_value={'error': {'message': 'En...am': None, 'code': 500}}, input_type=dict]
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] For further information visit https://errors.pydantic.dev/2.13/v/missing
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] choices
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] Field required [type=missing, input_value={'error': {'message': 'En...am': None, 'code': 500}}, input_type=dict]
(APIServer pid=2949) ERROR 09-20 02:43:06 [serving.py:1009] For further information visit https://errors.pydantic.dev/2.13/v/missing
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] Error in create_messages: EngineCore encountered an issue. See stack trace (above) for the root cause.
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] Traceback (most recent call last):
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/api_router.py", line 73, in create_messages
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] generator = await handler.create_messages(request, raw_request)
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/serving.py", line 610, in create_messages
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] generator = await self.create_chat_completion(chat_req, raw_request)
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 254, in create_chat_completion
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] return await self._with_kv_transfer_rejection_cleanup(
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/generate/base/serving.py", line 300, in _with_kv_transfer_rejection_cleanup
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] return await awaitable
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] ^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 275, in _create_chat_completion
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] result = await self.render_chat_request(request)
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 238, in render_chat_request
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] self._preflight(request.n or 1)
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/generate/base/serving.py", line 184, in _preflight
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] raise self.engine_client.dead_error
(APIServer pid=2949) ERROR 09-20 02:43:06 [api_router.py:75] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] Error in create_messages: EngineCore encountered an issue. See stack trace (above) for the root cause.
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] Traceback (most recent call last):
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/api_router.py", line 73, in create_messages
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] generator = await handler.create_messages(request, raw_request)
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/serving.py", line 610, in create_messages
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] generator = await self.create_chat_completion(chat_req, raw_request)
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 254, in create_chat_completion
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] return await self._with_kv_transfer_rejection_cleanup(
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/generate/base/serving.py", line 300, in _with_kv_transfer_rejection_cleanup
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] return await awaitable
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] ^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 275, in _create_chat_completion
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] result = await self.render_chat_request(request)
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 238, in render_chat_request
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] self._preflight(request.n or 1)
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/generate/base/serving.py", line 184, in _preflight
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] raise self.engine_client.dead_error
(APIServer pid=2949) ERROR 09-20 02:43:07 [api_router.py:75] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] Error in create_messages: EngineCore encountered an issue. See stack trace (above) for the root cause.
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] Traceback (most recent call last):
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/api_router.py", line 73, in create_messages
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] generator = await handler.create_messages(request, raw_request)
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/anthropic/serving.py", line 610, in create_messages
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] generator = await self.create_chat_completion(chat_req, raw_request)
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 254, in create_chat_completion
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] return await self._with_kv_transfer_rejection_cleanup(
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/generate/base/serving.py", line 300, in _with_kv_transfer_rejection_cleanup
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] return await awaitable
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] ^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 275, in _create_chat_completion
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] result = await self.render_chat_request(request)
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 238, in render_chat_request
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] self._preflight(request.n or 1)
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] File "/opt/vllm/lib/python3.11/site-packages/vllm/entrypoints/generate/base/serving.py", line 184, in _preflight
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] raise self.engine_client.dead_error
(APIServer pid=2949) ERROR 09-20 02:43:08 [api_router.py:75] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.
(EngineCore pid=3190) INFO 09-20 02:43:09 [multiproc_executor.py:479] [shutdown] Executor: all workers exited gracefully
(APIServer pid=2949) INFO 09-20 02:43:09 [core_client.py:689] [shutdown] MPClient: start timeout=0s
(APIServer pid=2949) INFO 09-20 02:43:09 [core_client.py:691] [shutdown] MPClient: stopping engine manager
(APIServer pid=2949) INFO 09-20 02:43:09 [utils.py:620] [shutdown] Process manager: send sigterm to process EngineCore
(APIServer pid=2949) WARNING 09-20 02:43:09 [utils.py:640] [shutdown] Process manager: force killing remaining processes count=1
(APIServer pid=2949) WARNING 09-20 02:43:09 [utils.py:645] [shutdown] Process manager: force killing remaining process EngineCore pid 3190
(APIServer pid=2949) INFO 09-20 02:43:09 [core_client.py:693] [shutdown] MPClient: engine manager stopped
(APIServer pid=2949) INFO 09-20 02:43:09 [core_client.py:694] [shutdown] MPClient: cleaning up background resources
(APIServer pid=2949) INFO 09-20 02:43:09 [core_client.py:696] [shutdown] MPClient: complete

Xet Storage Details

Size:
215 kB
·
Xet hash:
411c8f19d6637fd2478a7a8233c1531bd1f07750a24150ee0942197219feb5ce

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.