Dual-5090 voice stack (STT + LLM + TTS)
Bootstrap scripts that bring up a known-good Arabic voice stack on 2× RTX 5090
(Vast.ai / similar). Tuned for Blackwell (sm_120): no FP8 KV / ngram on the LLM,
Moss TTS forced to SDPA, CUDA graphs off, HF cache on /dev/shm when disk is tight.
| Service | Port | GPU | Model | Runtime |
|---|---|---|---|---|
| LLM | 8000 |
0 (shared w/ STT) | ibm-granite/granite-4.1-8b |
SGLang |
| STT | 8002 |
0 (shared w/ LLM) | CohereLabs/cohere-transcribe-arabic-07-2026 |
vLLM 0.19 |
| TTS | 8003 |
1 (dedicated) | Rabe3/Moss-Saudi-3 |
sglang-omni |
One-time setup on a new machine
# 1) HF token (write/read for this private repo; also needed for gated STT)
export HF_TOKEN=hf_...
# 2) Clone this private repo (replace USER with your HF username)
git clone https://USER:$HF_TOKEN@huggingface.co/USER/vast-voice-stack ~/agent-deploy
cd ~/agent-deploy
chmod +x *.sh
# 3) Disk: need ~40G+ free for two CUDA envs, OR leave HF cache on /dev/shm
# (default in common.sh). Prefer a host volume / larger disk when possible.
# 4) Launch (sequential install — do not parallelize first boot)
./deploy_all.sh
Upload from a machine that already has the scripts:
export HF_TOKEN=hf_... # token for the target account
./push_to_hf.sh # creates private USER/vast-voice-stack and uploads
Health checks: curl -sf localhost:8000/health / :8002 / :8003.
Restart clean:
./deploy_all.sh restart-clean
Load test (Locust)
./run_locust_ui.sh --local-only # or without flag to expose via Vast portal port 10100
Custom metrics: llm_ttfb, tts_ttfa, stt_ttfb, pipe_e2e. On this layout TTS is
the concurrency ceiling (~5–6 RPS full-WAV), not the LLM.
What these scripts already fix
- Disk: sequential env install,
PIP_NO_CACHE_DIR, shared~/sglang-envfor LLM+TTS, HF cache →/dev/shm/hf_homewhen overlay is ~32G. - sglang-omni: not on PyPI — clone +
uv pip install -ewith protobuf override. - STT: transformers from git for
cohere_asr; prefer local HF snapshot if present. - GPU1 budget: TTS dedicated at
MEM_FRACTION=0.72(colocated Moss OOM’d at 0.85). - GPU0 budget: STT
MEM_UTIL=0.18+ LLMMEM_FRACTION=0.68. - LLM on Blackwell: bf16 KV, no ngram speculative decoding (FP8/ngram caused device asserts).
- Moss on Blackwell: patch FA→SDPA in vocoder + audio tokenizer;
cuda_graph: false.
Load test:
./run_loadtest.sh # TTS + pipeline concurrency ramp (stream+pcm)
./run_loadtest.sh pipe # pipeline only
./run_locust_ui.sh # interactive UI
Env knobs
| Variable | Default | Meaning |
|---|---|---|
HF_TOKEN |
— | Required for gated STT download |
MODEL |
per-script | Override model id or local path |
PORT |
8000/8002/8003 | Bind port |
MEM_FRACTION / MEM_UTIL |
see scripts | GPU memory share |
LOG_DIR |
~/agent-logs |
Server logs |
Layout
common.sh # toolchain, HF home, venvs, omni + Blackwell patches
start_stt.sh / start_llm.sh / start_tts.sh
deploy_all.sh
moss_tts_local_blackwell.yaml # Moss pipeline config for 5090
locustfile.py / run_locust_ui.sh
Notes
- Models are not vendored here (~30G). Scripts pull them via Hugging Face Hub.
/dev/shmHF cache is lost on reboot — re-download or pointHF_HOMEat a persistent volume.- TTS returns full WAV (not streaming); latency grows under load even when LLM TTFT stays low.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support