vLLM Pre-built Wheel for Python 3.13 & CUDA 12.8 (A100 / SM 8.0)
Pre-built wheel of latest vLLM compiled from source for systems running Python 3.13 with CUDA 12.8 and PyTorch 2.10.0+cu128.
Build Information
- vLLM Commit:
a7ff44355c1cadd1d85b82f4adc9694ce7864b56(main branch) - Python Version:
3.13(cp313) - CUDA Toolkit / Driver: CUDA
12.8(NVCC12.8.61) - Host Compiler: GCC
12.4.0 - PyTorch Version:
2.10.0+cu128 - Target GPU Architecture:
sm_80(NVIDIA A100-SXM4 / PCIe) - Included Extensions:
_C_stable_libtorch.abi3.so(Ops registered via STABLE_TORCH_LIBRARY)_moe_C_stable_libtorch.abi3.so(Marlin MoE kernels)_vllm_fa2_C.abi3.so(FlashAttention-2)_vllm_fa3_C.abi3.so(FlashAttention-3)cumem_allocator.abi3.so,fs_io_C.abi3.so,spinloop.abi3.so
Installation
pip install https://huggingface.co/ruwwww/vllm-wheels/resolve/main/vllm-0.1.dev1+ga7ff44355.cu128-cp313-cp313-linux_x86_64.whl
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support