Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
DEPLOY : audit 13/13, pont vs natif
Browse files
DEPLOY.md
CHANGED
|
@@ -315,3 +315,28 @@ reconstitue un JSON valide, `stop_reason=tool_use`), `tool_result.is_error`,
|
|
| 315 |
`redacted_thinking` dans l'historique, `temperature/top_p/top_k/metadata`,
|
| 316 |
préremplissage par un dernier message assistant, modèle inconnu (repli sur le
|
| 317 |
modèle servi, 200).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 315 |
`redacted_thinking` dans l'historique, `temperature/top_p/top_k/metadata`,
|
| 316 |
préremplissage par un dernier message assistant, modèle inconnu (repli sur le
|
| 317 |
modèle servi, 200).
|
| 318 |
+
|
| 319 |
+
Vérifié après v51 : `cache_read_input_tokens` = 23 360 sur 24 047 au second
|
| 320 |
+
appel d'un prompt de 24 k (bloc de cache = 784 jetons en mode mamba align : un
|
| 321 |
+
prompt d'un seul bloc ne rapporte rien, c'est normal). Audit : **13/13**.
|
| 322 |
+
|
| 323 |
+
### Le pont est-il encore nécessaire ? (vLLM 0.27.1 sert `/v1/messages` nativement)
|
| 324 |
+
|
| 325 |
+
Oui, vLLM expose déjà `/v1/messages` et `/v1/messages/count_tokens` sur le même
|
| 326 |
+
port que l'API OpenAI (`vllm/entrypoints/anthropic/`) : blocs `thinking` avec
|
| 327 |
+
signature, `tool_use`/`input_json_delta`, images, `stop_sequence`, `ping`,
|
| 328 |
+
`cache_read_input_tokens`, `redacted_thinking`, raisonnement de l'historique
|
| 329 |
+
→ `reasoning`, fusion des messages système inline. Testé sur le pod :
|
| 330 |
+
`curl 127.0.0.1:18081/v1/messages` répond correctement.
|
| 331 |
+
|
| 332 |
+
Ce que le pont ajoute encore (mesuré ou vérifié dans les sources) :
|
| 333 |
+
- le **niveau de raisonnement** : `thinking.budget_tokens` → `thinking_token_budget`
|
| 334 |
+
(absent du serveur natif, qui ne connaît que `output_config.effort`) ;
|
| 335 |
+
- les **alias** `claude-*` / `[1m]` avec `context_window` dans `/v1/models`
|
| 336 |
+
(Claude Code lit `id`/`display_name` pour la découverte) ;
|
| 337 |
+
- le **keepalive pendant le préremplissage** (ping toutes les 15 s : le proxy
|
| 338 |
+
RunPod coupe à ~125 s de silence — un prefill de 634 k dure 4 min) ;
|
| 339 |
+
- le **garde agentique** (refus des commandes destructrices, 4/4 au test), la
|
| 340 |
+
console, la trace par requête, la transmission intacte des erreurs amont.
|
| 341 |
+
Pour un usage direct sans RunPod ni garde, le natif suffit : pointer
|
| 342 |
+
`ANTHROPIC_BASE_URL` sur le port vLLM et `ANTHROPIC_MODEL=ornith`.
|