# Serving on vLLM and SGLang `mindxtrain serve --to vllm|sglang` launches an OpenAI-compatible server on the trained checkpoint, waits until it answers, and can hand it to mindX as the fallback model. It sits beside `--to ollama` (merge + `ollama create`) and `--to bankml` (merge + `bankml create`, see [bankml.md](bankml.md)). Code: `mindxtrain/deploy/openai_server_push.py`. vLLM and SGLang are reached **only** as a subprocess and over HTTP — neither is imported by mindXtrain (clean-room policy). SGLang's presence is checked with `importlib.util.find_spec`, which does not import it. ``` mindxtrain serve run.yaml --to vllm [--tag NAME] [--checkpoint DIR] [--merge] [--host 127.0.0.1] [--port N] [--dtype auto] [--server-bin PATH] [--server-arg ARG ...] [--ready-timeout 600] [--cpu-kvcache-gib 4] [--register-as-fallback] mindxtrain serve run.yaml --to sglang ...same options... mindxtrain serve run.yaml --to vllm --dry-run # print the argv (and env); run nothing mindxtrain serve run.yaml --to vllm --stop # SIGTERM the recorded server, verify it is gone ``` ## What runs 1. **Checkpoint.** `--checkpoint`, else `out/runs//quantized/` when the config quantizes and that directory exists, else `out/runs//checkpoint/`. 2. **LoRA, native or merged.** A directory with `adapter_config.json` is a LoRA. By default it is served **natively over the base** (no merge): - vLLM: `--enable-lora --lora-modules = --max-lora-rank R` (R = the adapter's `r` rounded up to a vLLM choice: 1, 8, 16, 32, 64, 128, 256, 320, 512). The base is served as `-base`; the adapter is listed on `/v1/models` as ``, which is the client's `model`. - SGLang: `--enable-lora --lora-paths = --max-lora-rank r`. The base is served as `-base`; a client selects the adapter with `model: "-base:"`. `--merge` folds the adapter in first (`merge_lora_adapter`, needs `uv sync --extra ml`) and serves the merged directory as ``. 3. **argv.** vLLM: `vllm serve --served-model-name … --host --port --dtype --max-model-len --tensor-parallel-size [--quantization fp8|mxfp4|gptq]`. SGLang: `python -m sglang.launch_server --model-path … --served-model-name … --host --port --dtype --context-length --tp ` plus `--device cpu` on CPU or `--mem-fraction-static 0.85` on GPU. Anything else goes through `--server-arg`, verbatim (e.g. `--server-arg=--enforce-eager`). 4. **Detached.** The server runs in its own session; stdout+stderr go to `out/runs//serve//server.log`, its pid to `server.pid`, and the argv/env/time to `launch.json`. It outlives the CLI. 5. **Ready.** `/v1/models` is polled (and, for vLLM, `/health` must be 200) until the expected name is listed, the process exits, or `--ready-timeout` passes. On a timeout the server is **left running** (big models load slowly) — watch the log or `--stop` it. 6. **mindX.** With `--register-as-fallback`, mindX's fallback model is swapped to `{provider: "vllm"|"sglang", model: }` (best-effort, as for Ollama). ## CPU and GPU With no `/dev/kfd` (ROCm) or `/dev/nvidia0` the CPU backends are used: `--dtype auto` becomes `bfloat16` (the vLLM CPU guide's recommendation; float16 is unstable or unsupported on CPU), vLLM gets `VLLM_CPU_KVCACHE_SPACE=<--cpu-kvcache-gib>` unless the environment already sets it (other CPU knobs such as `VLLM_CPU_OMP_THREADS_BIND` pass through from your env), and SGLang gets `--device cpu`. A config that needs a GPU is **refused** on such a host, with the reason: an FP8 / MXFP4 / GPTQ checkpoint, or `serve.tensor_parallel > 1`. Neither upstream documents LoRA on CPU; if your build rejects it, use `--merge`. ## Installing the servers - vLLM: `uv sync --extra serve` (GPU / ROCm). CPU-only per the [vLLM CPU guide](https://docs.vllm.ai/en/latest/getting_started/installation/cpu.html), e.g. `uv pip install vllm --torch-backend cpu`. Or point `--server-bin` at a `vllm` elsewhere. - SGLang is not a mindXtrain extra: install it into the interpreter that will run it (`uv pip install sglang`, [docs.sglang.io](https://docs.sglang.io)) and pass that interpreter as `--server-bin` if it is not this one. ## Exit codes `0` ready / dry run / stopped / not running · `1` checkpoint missing · `2` refused: server missing, GPU needed, already running · `3` merge failed, server exited early, stop failed, error · `4` not ready before `--ready-timeout` (still running). ## Where each flag comes from vLLM: the [`vllm serve` CLI reference](https://docs.vllm.ai/en/latest/cli/serve.html) (`--served-model-name`, `--host`, `--port`, `--dtype`, `--max-model-len`, `--tensor-parallel-size`, `--quantization`, `--enable-lora`, `--lora-modules`, `--max-lora-rank` and its choices), the [LoRA page](https://docs.vllm.ai/en/latest/features/lora.html) (`name=path`, adapters on `/v1/models`), the [CPU installation page](https://docs.vllm.ai/en/latest/getting_started/installation/cpu.html) (`VLLM_CPU_KVCACHE_SPACE`, bfloat16), and [online serving](https://docs.vllm.ai/en/latest/serving/online_serving/) (`/health`, `/v1/models`). SGLang: [Server Arguments](https://docs.sglang.io/advanced_features/server_arguments.html) (`--model-path`, `--served-model-name`, `--host`, `--port`, `--dtype`, `--context-length`, `--device`, `--tp`, `--mem-fraction-static`, `--enable-lora`, `--lora-paths`, `--max-lora-rank`) and the [LoRA page](https://docs.sglang.io/advanced_features/lora.html) (`base:adapter` model syntax). Checked against the docs current on 2026-10-02 (`uv.lock` resolves vLLM 0.20.1); no server was installed or launched to write this.