Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
DEPLOY : patch effort natif
Browse files
DEPLOY.md
CHANGED
|
@@ -380,3 +380,17 @@ warmup (65 s à froid) · 5 s graphes CUDA. v55 (préparée, non publiée) :
|
|
| 380 |
`--limit-mm-per-prompt '{"video":0}'` et `HF_HUB_OFFLINE=1` quand le modèle est
|
| 381 |
en cache → visé ~1 min 45. Le vrai levier reste le nombre de redémarrages : en
|
| 382 |
natif tout changement de flag/alias en coûte un ; le pont rechargeait en 5 s.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 380 |
`--limit-mm-per-prompt '{"video":0}'` et `HF_HUB_OFFLINE=1` quand le modèle est
|
| 381 |
en cache → visé ~1 min 45. Le vrai levier reste le nombre de redémarrages : en
|
| 382 |
natif tout changement de flag/alias en coûte un ; le pont rechargeait en 5 s.
|
| 383 |
+
|
| 384 |
+
### Niveau de raisonnement en natif : patch vLLM (v55)
|
| 385 |
+
|
| 386 |
+
Constat (sources 0.27.1) : le serveur Anthropic natif ne lit pas `thinking`
|
| 387 |
+
(champ absent de `AnthropicMessagesRequest`, jeté par pydantic) et ne traduit
|
| 388 |
+
`output_config.effort` qu'en `reasoning_effort`, que seul Harmony/GPT-OSS
|
| 389 |
+
honore. Sur Ornith, effort low = max, et `disabled` n'agit pas.
|
| 390 |
+
|
| 391 |
+
`vllm_anthropic_effort_patch.py` (appliqué par le bootstrap avant `vllm serve`,
|
| 392 |
+
idempotent, `VL_EFFORT_PATCH=off` pour le couper) ajoute le champ `thinking` et
|
| 393 |
+
mappe : effort low/medium/high/xhigh/max → `thinking_token_budget`
|
| 394 |
+
1 024 / 4 096 / 16 384 / 32 768 / illimité ; `thinking.budget_tokens` explicite
|
| 395 |
+
gagne ; `thinking.type=disabled` → `enable_thinking=false`. Même mécanisme que
|
| 396 |
+
le pont (le bloc de raisonnement est fermé au plafond côté échantillonnage).
|