File size: 5,911 Bytes
730c5bb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 | # Serving on vLLM and SGLang
`mindxtrain serve <config.yaml> --to vllm|sglang` launches an OpenAI-compatible server on the
trained checkpoint, waits until it answers, and can hand it to mindX as the fallback model. It sits
beside `--to ollama` (merge + `ollama create`) and `--to bankml` (merge + `bankml create`, see
[bankml.md](bankml.md)). Code: `mindxtrain/deploy/openai_server_push.py`.
vLLM and SGLang are reached **only** as a subprocess and over HTTP — neither is imported by
mindXtrain (clean-room policy). SGLang's presence is checked with `importlib.util.find_spec`,
which does not import it.
```
mindxtrain serve run.yaml --to vllm [--tag NAME] [--checkpoint DIR] [--merge] [--host 127.0.0.1]
[--port N] [--dtype auto] [--server-bin PATH]
[--server-arg ARG ...] [--ready-timeout 600]
[--cpu-kvcache-gib 4] [--register-as-fallback]
mindxtrain serve run.yaml --to sglang ...same options...
mindxtrain serve run.yaml --to vllm --dry-run # print the argv (and env); run nothing
mindxtrain serve run.yaml --to vllm --stop # SIGTERM the recorded server, verify it is gone
```
## What runs
1. **Checkpoint.** `--checkpoint`, else `out/runs/<run>/quantized/` when the config quantizes and
that directory exists, else `out/runs/<run>/checkpoint/`.
2. **LoRA, native or merged.** A directory with `adapter_config.json` is a LoRA. By default it is
served **natively over the base** (no merge):
- vLLM: `--enable-lora --lora-modules <tag>=<adapter> --max-lora-rank R` (R = the adapter's `r`
rounded up to a vLLM choice: 1, 8, 16, 32, 64, 128, 256, 320, 512). The base is served as
`<tag>-base`; the adapter is listed on `/v1/models` as `<tag>`, which is the client's `model`.
- SGLang: `--enable-lora --lora-paths <tag>=<adapter> --max-lora-rank r`. The base is served as
`<tag>-base`; a client selects the adapter with `model: "<tag>-base:<tag>"`.
`--merge` folds the adapter in first (`merge_lora_adapter`, needs `uv sync --extra ml`) and
serves the merged directory as `<tag>`.
3. **argv.** vLLM: `vllm serve <model> --served-model-name … --host --port --dtype
--max-model-len <serve.max_model_len> --tensor-parallel-size <serve.tensor_parallel>
[--quantization fp8|mxfp4|gptq]`. SGLang: `python -m sglang.launch_server --model-path …
--served-model-name … --host --port --dtype --context-length <serve.max_model_len>
--tp <serve.tensor_parallel>` plus `--device cpu` on CPU or `--mem-fraction-static 0.85` on GPU.
Anything else goes through `--server-arg`, verbatim (e.g. `--server-arg=--enforce-eager`).
4. **Detached.** The server runs in its own session; stdout+stderr go to
`out/runs/<run>/serve/<to>/server.log`, its pid to `server.pid`, and the argv/env/time to
`launch.json`. It outlives the CLI.
5. **Ready.** `/v1/models` is polled (and, for vLLM, `/health` must be 200) until the expected
name is listed, the process exits, or `--ready-timeout` passes. On a timeout the server is
**left running** (big models load slowly) — watch the log or `--stop` it.
6. **mindX.** With `--register-as-fallback`, mindX's fallback model is swapped to
`{provider: "vllm"|"sglang", model: <client model>}` (best-effort, as for Ollama).
## CPU and GPU
With no `/dev/kfd` (ROCm) or `/dev/nvidia0` the CPU backends are used: `--dtype auto` becomes
`bfloat16` (the vLLM CPU guide's recommendation; float16 is unstable or unsupported on CPU), vLLM
gets `VLLM_CPU_KVCACHE_SPACE=<--cpu-kvcache-gib>` unless the environment already sets it (other
CPU knobs such as `VLLM_CPU_OMP_THREADS_BIND` pass through from your env), and SGLang gets
`--device cpu`. A config that needs a GPU is **refused** on such a host, with the reason: an FP8 /
MXFP4 / GPTQ checkpoint, or `serve.tensor_parallel > 1`. Neither upstream documents LoRA on CPU;
if your build rejects it, use `--merge`.
## Installing the servers
- vLLM: `uv sync --extra serve` (GPU / ROCm). CPU-only per the
[vLLM CPU guide](https://docs.vllm.ai/en/latest/getting_started/installation/cpu.html), e.g.
`uv pip install vllm --torch-backend cpu`. Or point `--server-bin` at a `vllm` elsewhere.
- SGLang is not a mindXtrain extra: install it into the interpreter that will run it
(`uv pip install sglang`, [docs.sglang.io](https://docs.sglang.io)) and pass that interpreter as
`--server-bin` if it is not this one.
## Exit codes
`0` ready / dry run / stopped / not running · `1` checkpoint missing · `2` refused: server missing,
GPU needed, already running · `3` merge failed, server exited early, stop failed, error ·
`4` not ready before `--ready-timeout` (still running).
## Where each flag comes from
vLLM: the [`vllm serve` CLI reference](https://docs.vllm.ai/en/latest/cli/serve.html)
(`--served-model-name`, `--host`, `--port`, `--dtype`, `--max-model-len`, `--tensor-parallel-size`,
`--quantization`, `--enable-lora`, `--lora-modules`, `--max-lora-rank` and its choices), the
[LoRA page](https://docs.vllm.ai/en/latest/features/lora.html) (`name=path`, adapters on
`/v1/models`), the [CPU installation page](https://docs.vllm.ai/en/latest/getting_started/installation/cpu.html)
(`VLLM_CPU_KVCACHE_SPACE`, bfloat16), and [online serving](https://docs.vllm.ai/en/latest/serving/online_serving/)
(`/health`, `/v1/models`). SGLang: [Server Arguments](https://docs.sglang.io/advanced_features/server_arguments.html)
(`--model-path`, `--served-model-name`, `--host`, `--port`, `--dtype`, `--context-length`,
`--device`, `--tp`, `--mem-fraction-static`, `--enable-lora`, `--lora-paths`, `--max-lora-rank`)
and the [LoRA page](https://docs.sglang.io/advanced_features/lora.html) (`base:adapter` model
syntax). Checked against the docs current on 2026-10-02 (`uv.lock` resolves vLLM 0.20.1); no
server was installed or launched to write this.
|