Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
DEPLOY: pont muet sous charge (fuite de flux), correctifs
Browse files
DEPLOY.md
CHANGED
|
@@ -500,3 +500,21 @@ patch natif). `VL_BRIDGE=off` = natif + patch effort, pour un usage léger.
|
|
| 500 |
Leviers réels pour ce régime : moins de réflexion pour les sous-agents (`effort: low`), moins
|
| 501 |
d'agents en parallèle, outils qui renvoient moins de texte ; au-delà, du calcul GPU
|
| 502 |
(`reports/2026-08-22-materiel-prix-runpod-hf.md`).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 500 |
Leviers réels pour ce régime : moins de réflexion pour les sous-agents (`effort: low`), moins
|
| 501 |
d'agents en parallèle, outils qui renvoient moins de texte ; au-delà, du calcul GPU
|
| 502 |
(`reports/2026-08-22-materiel-prix-runpod-hf.md`).
|
| 503 |
+
|
| 504 |
+
### Pont muet sous charge (22/08 nuit) : fuite de flux amont, pool httpx saturé
|
| 505 |
+
|
| 506 |
+
Symptôme : agents Claude Code à 0 jeton / « failed », `/metrics` via le proxy en timeout ; sur
|
| 507 |
+
le pod : vLLM sain (health 2 ms, 0 requête, GPU 0 %), **pont vivant (3 % CPU, 117 fd) mais
|
| 508 |
+
`/health` en timeout**. Journal du pont : 189 `POST /v1/messages` pour 100 traces de fin ;
|
| 509 |
+
`ss` : 45 connexions amont vers vLLM pour 2 clients. Les flux abandonnés côté client (retry de
|
| 510 |
+
Claude Code, coupure proxy) gardaient leur connexion amont — vLLM générait pour personne — et à
|
| 511 |
+
100 connexions (limite httpx par défaut) le pool bloquait tout, `/health` compris (il passe par
|
| 512 |
+
l'amont).
|
| 513 |
+
|
| 514 |
+
Correctifs : (1) `stream_anthropic` reçoit la `Request` et **ferme l'amont quand le client est
|
| 515 |
+
parti** (`request.is_disconnected()` à chaque ping et tous les 32 jetons) → vLLM annule la
|
| 516 |
+
génération ; (2) `httpx.Limits(max_connections=512, max_keepalive_connections=64)` et
|
| 517 |
+
`httpx.Timeout(TIMEOUT, connect=10, pool=10)` → une saturation devient un 503 rapide, plus un
|
| 518 |
+
blocage ; (3) chien de garde dans le watcher du bootstrap (v69) : deux `/health` KO d'affilée
|
| 519 |
+
→ relance du pont seul, journal conservé dans `/tmp/proxy.hung.*.log` (et, posé à la main ce
|
| 520 |
+
soir, `/opt/pont_watchdog.sh`). Le pont se recharge à chaud depuis le Hub sans toucher à vLLM.
|