Download docs/serve.md from PYTHAI/mindXtrain: direct link, hf CLI and curl.
- Browser
- Download file 5.91 kB
-
https://huggingface.co/PYTHAI/mindXtrain/resolve/main/docs/serve.md
- Command line
-
hf download hf://PYTHAI/mindXtrain/docs/serve.md
-
curl -L -o serve.md https://huggingface.co/PYTHAI/mindXtrain/resolve/main/docs/serve.md
Serving on vLLM and SGLang
mindxtrain serve <config.yaml> --to vllm|sglang launches an OpenAI-compatible server on the
trained checkpoint, waits until it answers, and can hand it to mindX as the fallback model. It sits
beside --to ollama (merge + ollama create) and --to bankml (merge + bankml create, see
bankml.md). Code: mindxtrain/deploy/openai_server_push.py.
vLLM and SGLang are reached only as a subprocess and over HTTP — neither is imported by
mindXtrain (clean-room policy). SGLang's presence is checked with importlib.util.find_spec,
which does not import it.
mindxtrain serve run.yaml --to vllm [--tag NAME] [--checkpoint DIR] [--merge] [--host 127.0.0.1]
[--port N] [--dtype auto] [--server-bin PATH]
[--server-arg ARG ...] [--ready-timeout 600]
[--cpu-kvcache-gib 4] [--register-as-fallback]
mindxtrain serve run.yaml --to sglang ...same options...
mindxtrain serve run.yaml --to vllm --dry-run # print the argv (and env); run nothing
mindxtrain serve run.yaml --to vllm --stop # SIGTERM the recorded server, verify it is gone
What runs
Checkpoint.
--checkpoint, elseout/runs/<run>/quantized/when the config quantizes and that directory exists, elseout/runs/<run>/checkpoint/.LoRA, native or merged. A directory with
adapter_config.jsonis a LoRA. By default it is served natively over the base (no merge):- vLLM:
--enable-lora --lora-modules <tag>=<adapter> --max-lora-rank R(R = the adapter'srrounded up to a vLLM choice: 1, 8, 16, 32, 64, 128, 256, 320, 512). The base is served as<tag>-base; the adapter is listed on/v1/modelsas<tag>, which is the client'smodel. - SGLang:
--enable-lora --lora-paths <tag>=<adapter> --max-lora-rank r. The base is served as<tag>-base; a client selects the adapter withmodel: "<tag>-base:<tag>".
--mergefolds the adapter in first (merge_lora_adapter, needsuv sync --extra ml) and serves the merged directory as<tag>.- vLLM:
argv. vLLM:
vllm serve <model> --served-model-name … --host --port --dtype --max-model-len <serve.max_model_len> --tensor-parallel-size <serve.tensor_parallel> [--quantization fp8|mxfp4|gptq]. SGLang:python -m sglang.launch_server --model-path … --served-model-name … --host --port --dtype --context-length <serve.max_model_len> --tp <serve.tensor_parallel>plus--device cpuon CPU or--mem-fraction-static 0.85on GPU. Anything else goes through--server-arg, verbatim (e.g.--server-arg=--enforce-eager).Detached. The server runs in its own session; stdout+stderr go to
out/runs/<run>/serve/<to>/server.log, its pid toserver.pid, and the argv/env/time tolaunch.json. It outlives the CLI.Ready.
/v1/modelsis polled (and, for vLLM,/healthmust be 200) until the expected name is listed, the process exits, or--ready-timeoutpasses. On a timeout the server is left running (big models load slowly) — watch the log or--stopit.mindX. With
--register-as-fallback, mindX's fallback model is swapped to{provider: "vllm"|"sglang", model: <client model>}(best-effort, as for Ollama).
CPU and GPU
With no /dev/kfd (ROCm) or /dev/nvidia0 the CPU backends are used: --dtype auto becomes
bfloat16 (the vLLM CPU guide's recommendation; float16 is unstable or unsupported on CPU), vLLM
gets VLLM_CPU_KVCACHE_SPACE=<--cpu-kvcache-gib> unless the environment already sets it (other
CPU knobs such as VLLM_CPU_OMP_THREADS_BIND pass through from your env), and SGLang gets
--device cpu. A config that needs a GPU is refused on such a host, with the reason: an FP8 /
MXFP4 / GPTQ checkpoint, or serve.tensor_parallel > 1. Neither upstream documents LoRA on CPU;
if your build rejects it, use --merge.
Installing the servers
- vLLM:
uv sync --extra serve(GPU / ROCm). CPU-only per the vLLM CPU guide, e.g.uv pip install vllm --torch-backend cpu. Or point--server-binat avllmelsewhere. - SGLang is not a mindXtrain extra: install it into the interpreter that will run it
(
uv pip install sglang, docs.sglang.io) and pass that interpreter as--server-binif it is not this one.
Exit codes
0 ready / dry run / stopped / not running · 1 checkpoint missing · 2 refused: server missing,
GPU needed, already running · 3 merge failed, server exited early, stop failed, error ·
4 not ready before --ready-timeout (still running).
Where each flag comes from
vLLM: the vllm serve CLI reference
(--served-model-name, --host, --port, --dtype, --max-model-len, --tensor-parallel-size,
--quantization, --enable-lora, --lora-modules, --max-lora-rank and its choices), the
LoRA page (name=path, adapters on
/v1/models), the CPU installation page
(VLLM_CPU_KVCACHE_SPACE, bfloat16), and online serving
(/health, /v1/models). SGLang: Server Arguments
(--model-path, --served-model-name, --host, --port, --dtype, --context-length,
--device, --tp, --mem-fraction-static, --enable-lora, --lora-paths, --max-lora-rank)
and the LoRA page (base:adapter model
syntax). Checked against the docs current on 2026-10-02 (uv.lock resolves vLLM 0.20.1); no
server was installed or launched to write this.