Heads-up for 0.5.19 users: the new prefill CUDA graph holds 1.81 GB and starves quantized-KV long-context prefill
#4 opened 18 days ago
by
nakakennnn
Request: wheel built against PyTorch 2.13 (current build caps users at SGLang <= 0.5.17)
4
#3 opened 25 days ago
by
nakakennnn
Expose reasoning effort capability in the OpenAI-compatible API
#2 opened about 1 month ago
by
lhfe