vLLM Pre-built Wheel for Python 3.13 & CUDA 12.8 (A100 / SM 8.0)

Pre-built wheel of latest vLLM compiled from source for systems running Python 3.13 with CUDA 12.8 and PyTorch 2.10.0+cu128.

Build Information

  • vLLM Commit: a7ff44355c1cadd1d85b82f4adc9694ce7864b56 (main branch)
  • Python Version: 3.13 (cp313)
  • CUDA Toolkit / Driver: CUDA 12.8 (NVCC 12.8.61)
  • Host Compiler: GCC 12.4.0
  • PyTorch Version: 2.10.0+cu128
  • Target GPU Architecture: sm_80 (NVIDIA A100-SXM4 / PCIe)
  • Included Extensions:
    • _C_stable_libtorch.abi3.so (Ops registered via STABLE_TORCH_LIBRARY)
    • _moe_C_stable_libtorch.abi3.so (Marlin MoE kernels)
    • _vllm_fa2_C.abi3.so (FlashAttention-2)
    • _vllm_fa3_C.abi3.so (FlashAttention-3)
    • cumem_allocator.abi3.so, fs_io_C.abi3.so, spinloop.abi3.so

Installation

pip install https://huggingface.co/ruwwww/vllm-wheels/resolve/main/vllm-0.1.dev1+ga7ff44355.cu128-cp313-cp313-linux_x86_64.whl
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support