mindXtrain / docs /serve.md
Gregory-L's picture
bankml: the verified CPU engine as backend, serve target and imprint probe (#1)
730c5bb
|
Raw History Blame Contribute Delete
5.91 kB

Serving on vLLM and SGLang

mindxtrain serve <config.yaml> --to vllm|sglang launches an OpenAI-compatible server on the trained checkpoint, waits until it answers, and can hand it to mindX as the fallback model. It sits beside --to ollama (merge + ollama create) and --to bankml (merge + bankml create, see bankml.md). Code: mindxtrain/deploy/openai_server_push.py.

vLLM and SGLang are reached only as a subprocess and over HTTP — neither is imported by mindXtrain (clean-room policy). SGLang's presence is checked with importlib.util.find_spec, which does not import it.

mindxtrain serve run.yaml --to vllm   [--tag NAME] [--checkpoint DIR] [--merge] [--host 127.0.0.1]
                                      [--port N] [--dtype auto] [--server-bin PATH]
                                      [--server-arg ARG ...] [--ready-timeout 600]
                                      [--cpu-kvcache-gib 4] [--register-as-fallback]
mindxtrain serve run.yaml --to sglang ...same options...
mindxtrain serve run.yaml --to vllm --dry-run      # print the argv (and env); run nothing
mindxtrain serve run.yaml --to vllm --stop         # SIGTERM the recorded server, verify it is gone

What runs

  1. Checkpoint. --checkpoint, else out/runs/<run>/quantized/ when the config quantizes and that directory exists, else out/runs/<run>/checkpoint/.

  2. LoRA, native or merged. A directory with adapter_config.json is a LoRA. By default it is served natively over the base (no merge):

    • vLLM: --enable-lora --lora-modules <tag>=<adapter> --max-lora-rank R (R = the adapter's r rounded up to a vLLM choice: 1, 8, 16, 32, 64, 128, 256, 320, 512). The base is served as <tag>-base; the adapter is listed on /v1/models as <tag>, which is the client's model.
    • SGLang: --enable-lora --lora-paths <tag>=<adapter> --max-lora-rank r. The base is served as <tag>-base; a client selects the adapter with model: "<tag>-base:<tag>".

    --merge folds the adapter in first (merge_lora_adapter, needs uv sync --extra ml) and serves the merged directory as <tag>.

  3. argv. vLLM: vllm serve <model> --served-model-name … --host --port --dtype --max-model-len <serve.max_model_len> --tensor-parallel-size <serve.tensor_parallel> [--quantization fp8|mxfp4|gptq]. SGLang: python -m sglang.launch_server --model-path … --served-model-name … --host --port --dtype --context-length <serve.max_model_len> --tp <serve.tensor_parallel> plus --device cpu on CPU or --mem-fraction-static 0.85 on GPU. Anything else goes through --server-arg, verbatim (e.g. --server-arg=--enforce-eager).

  4. Detached. The server runs in its own session; stdout+stderr go to out/runs/<run>/serve/<to>/server.log, its pid to server.pid, and the argv/env/time to launch.json. It outlives the CLI.

  5. Ready. /v1/models is polled (and, for vLLM, /health must be 200) until the expected name is listed, the process exits, or --ready-timeout passes. On a timeout the server is left running (big models load slowly) — watch the log or --stop it.

  6. mindX. With --register-as-fallback, mindX's fallback model is swapped to {provider: "vllm"|"sglang", model: <client model>} (best-effort, as for Ollama).

CPU and GPU

With no /dev/kfd (ROCm) or /dev/nvidia0 the CPU backends are used: --dtype auto becomes bfloat16 (the vLLM CPU guide's recommendation; float16 is unstable or unsupported on CPU), vLLM gets VLLM_CPU_KVCACHE_SPACE=<--cpu-kvcache-gib> unless the environment already sets it (other CPU knobs such as VLLM_CPU_OMP_THREADS_BIND pass through from your env), and SGLang gets --device cpu. A config that needs a GPU is refused on such a host, with the reason: an FP8 / MXFP4 / GPTQ checkpoint, or serve.tensor_parallel > 1. Neither upstream documents LoRA on CPU; if your build rejects it, use --merge.

Installing the servers

  • vLLM: uv sync --extra serve (GPU / ROCm). CPU-only per the vLLM CPU guide, e.g. uv pip install vllm --torch-backend cpu. Or point --server-bin at a vllm elsewhere.
  • SGLang is not a mindXtrain extra: install it into the interpreter that will run it (uv pip install sglang, docs.sglang.io) and pass that interpreter as --server-bin if it is not this one.

Exit codes

0 ready / dry run / stopped / not running · 1 checkpoint missing · 2 refused: server missing, GPU needed, already running · 3 merge failed, server exited early, stop failed, error · 4 not ready before --ready-timeout (still running).

Where each flag comes from

vLLM: the vllm serve CLI reference (--served-model-name, --host, --port, --dtype, --max-model-len, --tensor-parallel-size, --quantization, --enable-lora, --lora-modules, --max-lora-rank and its choices), the LoRA page (name=path, adapters on /v1/models), the CPU installation page (VLLM_CPU_KVCACHE_SPACE, bfloat16), and online serving (/health, /v1/models). SGLang: Server Arguments (--model-path, --served-model-name, --host, --port, --dtype, --context-length, --device, --tp, --mem-fraction-static, --enable-lora, --lora-paths, --max-lora-rank) and the LoRA page (base:adapter model syntax). Checked against the docs current on 2026-10-02 (uv.lock resolves vLLM 0.20.1); no server was installed or launched to write this.