|
Download docs/serve.md from PYTHAI/mindXtrain: direct link, hf CLI and curl.
- Browser
- Download file 5.91 kB
-
https://huggingface.co/PYTHAI/mindXtrain/resolve/main/docs/serve.md
- Command line
-
hf download hf://PYTHAI/mindXtrain/docs/serve.md
-
curl -L -o serve.md https://huggingface.co/PYTHAI/mindXtrain/resolve/main/docs/serve.md
5.91 kB
| # Serving on vLLM and SGLang | |
| `mindxtrain serve <config.yaml> --to vllm|sglang` launches an OpenAI-compatible server on the | |
| trained checkpoint, waits until it answers, and can hand it to mindX as the fallback model. It sits | |
| beside `--to ollama` (merge + `ollama create`) and `--to bankml` (merge + `bankml create`, see | |
| [bankml.md](bankml.md)). Code: `mindxtrain/deploy/openai_server_push.py`. | |
| vLLM and SGLang are reached **only** as a subprocess and over HTTP — neither is imported by | |
| mindXtrain (clean-room policy). SGLang's presence is checked with `importlib.util.find_spec`, | |
| which does not import it. | |
| ``` | |
| mindxtrain serve run.yaml --to vllm [--tag NAME] [--checkpoint DIR] [--merge] [--host 127.0.0.1] | |
| [--port N] [--dtype auto] [--server-bin PATH] | |
| [--server-arg ARG ...] [--ready-timeout 600] | |
| [--cpu-kvcache-gib 4] [--register-as-fallback] | |
| mindxtrain serve run.yaml --to sglang ...same options... | |
| mindxtrain serve run.yaml --to vllm --dry-run # print the argv (and env); run nothing | |
| mindxtrain serve run.yaml --to vllm --stop # SIGTERM the recorded server, verify it is gone | |
| ``` | |
| ## What runs | |
| 1. **Checkpoint.** `--checkpoint`, else `out/runs/<run>/quantized/` when the config quantizes and | |
| that directory exists, else `out/runs/<run>/checkpoint/`. | |
| 2. **LoRA, native or merged.** A directory with `adapter_config.json` is a LoRA. By default it is | |
| served **natively over the base** (no merge): | |
| - vLLM: `--enable-lora --lora-modules <tag>=<adapter> --max-lora-rank R` (R = the adapter's `r` | |
| rounded up to a vLLM choice: 1, 8, 16, 32, 64, 128, 256, 320, 512). The base is served as | |
| `<tag>-base`; the adapter is listed on `/v1/models` as `<tag>`, which is the client's `model`. | |
| - SGLang: `--enable-lora --lora-paths <tag>=<adapter> --max-lora-rank r`. The base is served as | |
| `<tag>-base`; a client selects the adapter with `model: "<tag>-base:<tag>"`. | |
| `--merge` folds the adapter in first (`merge_lora_adapter`, needs `uv sync --extra ml`) and | |
| serves the merged directory as `<tag>`. | |
| 3. **argv.** vLLM: `vllm serve <model> --served-model-name … --host --port --dtype | |
| --max-model-len <serve.max_model_len> --tensor-parallel-size <serve.tensor_parallel> | |
| [--quantization fp8|mxfp4|gptq]`. SGLang: `python -m sglang.launch_server --model-path … | |
| --served-model-name … --host --port --dtype --context-length <serve.max_model_len> | |
| --tp <serve.tensor_parallel>` plus `--device cpu` on CPU or `--mem-fraction-static 0.85` on GPU. | |
| Anything else goes through `--server-arg`, verbatim (e.g. `--server-arg=--enforce-eager`). | |
| 4. **Detached.** The server runs in its own session; stdout+stderr go to | |
| `out/runs/<run>/serve/<to>/server.log`, its pid to `server.pid`, and the argv/env/time to | |
| `launch.json`. It outlives the CLI. | |
| 5. **Ready.** `/v1/models` is polled (and, for vLLM, `/health` must be 200) until the expected | |
| name is listed, the process exits, or `--ready-timeout` passes. On a timeout the server is | |
| **left running** (big models load slowly) — watch the log or `--stop` it. | |
| 6. **mindX.** With `--register-as-fallback`, mindX's fallback model is swapped to | |
| `{provider: "vllm"|"sglang", model: <client model>}` (best-effort, as for Ollama). | |
| ## CPU and GPU | |
| With no `/dev/kfd` (ROCm) or `/dev/nvidia0` the CPU backends are used: `--dtype auto` becomes | |
| `bfloat16` (the vLLM CPU guide's recommendation; float16 is unstable or unsupported on CPU), vLLM | |
| gets `VLLM_CPU_KVCACHE_SPACE=<--cpu-kvcache-gib>` unless the environment already sets it (other | |
| CPU knobs such as `VLLM_CPU_OMP_THREADS_BIND` pass through from your env), and SGLang gets | |
| `--device cpu`. A config that needs a GPU is **refused** on such a host, with the reason: an FP8 / | |
| MXFP4 / GPTQ checkpoint, or `serve.tensor_parallel > 1`. Neither upstream documents LoRA on CPU; | |
| if your build rejects it, use `--merge`. | |
| ## Installing the servers | |
| - vLLM: `uv sync --extra serve` (GPU / ROCm). CPU-only per the | |
| [vLLM CPU guide](https://docs.vllm.ai/en/latest/getting_started/installation/cpu.html), e.g. | |
| `uv pip install vllm --torch-backend cpu`. Or point `--server-bin` at a `vllm` elsewhere. | |
| - SGLang is not a mindXtrain extra: install it into the interpreter that will run it | |
| (`uv pip install sglang`, [docs.sglang.io](https://docs.sglang.io)) and pass that interpreter as | |
| `--server-bin` if it is not this one. | |
| ## Exit codes | |
| `0` ready / dry run / stopped / not running · `1` checkpoint missing · `2` refused: server missing, | |
| GPU needed, already running · `3` merge failed, server exited early, stop failed, error · | |
| `4` not ready before `--ready-timeout` (still running). | |
| ## Where each flag comes from | |
| vLLM: the [`vllm serve` CLI reference](https://docs.vllm.ai/en/latest/cli/serve.html) | |
| (`--served-model-name`, `--host`, `--port`, `--dtype`, `--max-model-len`, `--tensor-parallel-size`, | |
| `--quantization`, `--enable-lora`, `--lora-modules`, `--max-lora-rank` and its choices), the | |
| [LoRA page](https://docs.vllm.ai/en/latest/features/lora.html) (`name=path`, adapters on | |
| `/v1/models`), the [CPU installation page](https://docs.vllm.ai/en/latest/getting_started/installation/cpu.html) | |
| (`VLLM_CPU_KVCACHE_SPACE`, bfloat16), and [online serving](https://docs.vllm.ai/en/latest/serving/online_serving/) | |
| (`/health`, `/v1/models`). SGLang: [Server Arguments](https://docs.sglang.io/advanced_features/server_arguments.html) | |
| (`--model-path`, `--served-model-name`, `--host`, `--port`, `--dtype`, `--context-length`, | |
| `--device`, `--tp`, `--mem-fraction-static`, `--enable-lora`, `--lora-paths`, `--max-lora-rank`) | |
| and the [LoRA page](https://docs.sglang.io/advanced_features/lora.html) (`base:adapter` model | |
| syntax). Checked against the docs current on 2026-10-02 (`uv.lock` resolves vLLM 0.20.1); no | |
| server was installed or launched to write this. | |