File size: 5,911 Bytes
730c5bb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
# Serving on vLLM and SGLang

`mindxtrain serve <config.yaml> --to vllm|sglang` launches an OpenAI-compatible server on the
trained checkpoint, waits until it answers, and can hand it to mindX as the fallback model. It sits
beside `--to ollama` (merge + `ollama create`) and `--to bankml` (merge + `bankml create`, see
[bankml.md](bankml.md)). Code: `mindxtrain/deploy/openai_server_push.py`.

vLLM and SGLang are reached **only** as a subprocess and over HTTP — neither is imported by
mindXtrain (clean-room policy). SGLang's presence is checked with `importlib.util.find_spec`,
which does not import it.

```
mindxtrain serve run.yaml --to vllm   [--tag NAME] [--checkpoint DIR] [--merge] [--host 127.0.0.1]
                                      [--port N] [--dtype auto] [--server-bin PATH]
                                      [--server-arg ARG ...] [--ready-timeout 600]
                                      [--cpu-kvcache-gib 4] [--register-as-fallback]
mindxtrain serve run.yaml --to sglang ...same options...
mindxtrain serve run.yaml --to vllm --dry-run      # print the argv (and env); run nothing
mindxtrain serve run.yaml --to vllm --stop         # SIGTERM the recorded server, verify it is gone
```

## What runs

1. **Checkpoint.** `--checkpoint`, else `out/runs/<run>/quantized/` when the config quantizes and
   that directory exists, else `out/runs/<run>/checkpoint/`.
2. **LoRA, native or merged.** A directory with `adapter_config.json` is a LoRA. By default it is
   served **natively over the base** (no merge):
   - vLLM: `--enable-lora --lora-modules <tag>=<adapter> --max-lora-rank R` (R = the adapter's `r`
     rounded up to a vLLM choice: 1, 8, 16, 32, 64, 128, 256, 320, 512). The base is served as
     `<tag>-base`; the adapter is listed on `/v1/models` as `<tag>`, which is the client's `model`.
   - SGLang: `--enable-lora --lora-paths <tag>=<adapter> --max-lora-rank r`. The base is served as
     `<tag>-base`; a client selects the adapter with `model: "<tag>-base:<tag>"`.

   `--merge` folds the adapter in first (`merge_lora_adapter`, needs `uv sync --extra ml`) and
   serves the merged directory as `<tag>`.
3. **argv.** vLLM: `vllm serve <model> --served-model-name … --host --port --dtype
   --max-model-len <serve.max_model_len> --tensor-parallel-size <serve.tensor_parallel>
   [--quantization fp8|mxfp4|gptq]`. SGLang: `python -m sglang.launch_server --model-path …
   --served-model-name … --host --port --dtype --context-length <serve.max_model_len>
   --tp <serve.tensor_parallel>` plus `--device cpu` on CPU or `--mem-fraction-static 0.85` on GPU.
   Anything else goes through `--server-arg`, verbatim (e.g. `--server-arg=--enforce-eager`).
4. **Detached.** The server runs in its own session; stdout+stderr go to
   `out/runs/<run>/serve/<to>/server.log`, its pid to `server.pid`, and the argv/env/time to
   `launch.json`. It outlives the CLI.
5. **Ready.** `/v1/models` is polled (and, for vLLM, `/health` must be 200) until the expected
   name is listed, the process exits, or `--ready-timeout` passes. On a timeout the server is
   **left running** (big models load slowly) — watch the log or `--stop` it.
6. **mindX.** With `--register-as-fallback`, mindX's fallback model is swapped to
   `{provider: "vllm"|"sglang", model: <client model>}` (best-effort, as for Ollama).

## CPU and GPU

With no `/dev/kfd` (ROCm) or `/dev/nvidia0` the CPU backends are used: `--dtype auto` becomes
`bfloat16` (the vLLM CPU guide's recommendation; float16 is unstable or unsupported on CPU), vLLM
gets `VLLM_CPU_KVCACHE_SPACE=<--cpu-kvcache-gib>` unless the environment already sets it (other
CPU knobs such as `VLLM_CPU_OMP_THREADS_BIND` pass through from your env), and SGLang gets
`--device cpu`. A config that needs a GPU is **refused** on such a host, with the reason: an FP8 /
MXFP4 / GPTQ checkpoint, or `serve.tensor_parallel > 1`. Neither upstream documents LoRA on CPU;
if your build rejects it, use `--merge`.

## Installing the servers

- vLLM: `uv sync --extra serve` (GPU / ROCm). CPU-only per the
  [vLLM CPU guide](https://docs.vllm.ai/en/latest/getting_started/installation/cpu.html), e.g.
  `uv pip install vllm --torch-backend cpu`. Or point `--server-bin` at a `vllm` elsewhere.
- SGLang is not a mindXtrain extra: install it into the interpreter that will run it
  (`uv pip install sglang`, [docs.sglang.io](https://docs.sglang.io)) and pass that interpreter as
  `--server-bin` if it is not this one.

## Exit codes

`0` ready / dry run / stopped / not running · `1` checkpoint missing · `2` refused: server missing,
GPU needed, already running · `3` merge failed, server exited early, stop failed, error ·
`4` not ready before `--ready-timeout` (still running).

## Where each flag comes from

vLLM: the [`vllm serve` CLI reference](https://docs.vllm.ai/en/latest/cli/serve.html)
(`--served-model-name`, `--host`, `--port`, `--dtype`, `--max-model-len`, `--tensor-parallel-size`,
`--quantization`, `--enable-lora`, `--lora-modules`, `--max-lora-rank` and its choices), the
[LoRA page](https://docs.vllm.ai/en/latest/features/lora.html) (`name=path`, adapters on
`/v1/models`), the [CPU installation page](https://docs.vllm.ai/en/latest/getting_started/installation/cpu.html)
(`VLLM_CPU_KVCACHE_SPACE`, bfloat16), and [online serving](https://docs.vllm.ai/en/latest/serving/online_serving/)
(`/health`, `/v1/models`). SGLang: [Server Arguments](https://docs.sglang.io/advanced_features/server_arguments.html)
(`--model-path`, `--served-model-name`, `--host`, `--port`, `--dtype`, `--context-length`,
`--device`, `--tp`, `--mem-fraction-static`, `--enable-lora`, `--lora-paths`, `--max-lora-rank`)
and the [LoRA page](https://docs.sglang.io/advanced_features/lora.html) (`base:adapter` model
syntax). Checked against the docs current on 2026-10-02 (`uv.lock` resolves vLLM 0.20.1); no
server was installed or launched to write this.