Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Upload RESULTS.md with huggingface_hub
Browse files- RESULTS.md +46 -0
RESULTS.md
CHANGED
|
@@ -298,3 +298,49 @@ plafond mono-flux SANS speculation ~147 tok/s theorique, ~130
|
|
| 298 |
possible que parce que la spéculation émet plusieurs tokens par lecture des
|
| 299 |
poids. Sans elle, aucun réglage ne peut approcher 500 en mono-flux sur cette
|
| 300 |
carte.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 298 |
possible que parce que la spéculation émet plusieurs tokens par lecture des
|
| 299 |
poids. Sans elle, aucun réglage ne peut approcher 500 en mono-flux sur cette
|
| 300 |
carte.
|
| 301 |
+
|
| 302 |
+
---
|
| 303 |
+
|
| 304 |
+
# CORRECTION — le cache de préfixe fonctionne sur Kimi-Linear
|
| 305 |
+
|
| 306 |
+
Une section plus haut affirme que Kimi-Linear n'a pas de cache de préfixe, sur la
|
| 307 |
+
base d'un rapport froid/chaud de 1,04×. **C'était une erreur de configuration de
|
| 308 |
+
ma part, pas une limite du modèle** : le pod concerné ne passait pas
|
| 309 |
+
`--enable-prefix-caching`. Avec le flag explicite, sur la même carte :
|
| 310 |
+
|
| 311 |
+
```
|
| 312 |
+
sans le flag avec le flag
|
| 313 |
+
prefill a froid 6 162 tok/s 7 647 tok/s
|
| 314 |
+
prefill a chaud 6 430 tok/s 36 257 tok/s
|
| 315 |
+
rapport 1,04x 4,74x
|
| 316 |
+
```
|
| 317 |
+
|
| 318 |
+
Conséquence pour un client agent : un prompt système de 32k **est** mis en cache
|
| 319 |
+
sur Kimi-Linear. L'argument qui le disqualifiait pour Claude Code tombe.
|
| 320 |
+
|
| 321 |
+
## Kimi-Linear, mesures définitives (pod persistant, flags corrects)
|
| 322 |
+
|
| 323 |
+
```
|
| 324 |
+
flux 1 4 8 16 32
|
| 325 |
+
agrege 83,9 250,6 402,9 616,4 772,4 tok/s
|
| 326 |
+
$/1M 1,457 0,488 0,303 0,198 0,158
|
| 327 |
+
|
| 328 |
+
prefill froid 7 647 tok/s (Qwen : 4 485)
|
| 329 |
+
prefill chaud 36 257 tok/s (Qwen : 70 568)
|
| 330 |
+
edition de code 100,5 tok/s (Qwen : 575)
|
| 331 |
+
contexte 1 048 576 KV 1 505 550 tokens
|
| 332 |
+
protocole Anthropic : les 5 controles passent
|
| 333 |
+
```
|
| 334 |
+
|
| 335 |
+
Kimi-Linear franchit 500 tok/s **en agrégé dès 16 flux** (616,4). Qwen les
|
| 336 |
+
franchit **en mono-flux** (575). Les deux tiennent la cible sur une seule A40,
|
| 337 |
+
par des chemins différents : la spéculation pour Qwen, la concurrence pour Kimi
|
| 338 |
+
— chez qui la spéculation reste inutilisable puisqu'elle corrompt le code.
|
| 339 |
+
|
| 340 |
+
## Choix recommandé
|
| 341 |
+
|
| 342 |
+
| besoin | modèle | pourquoi |
|
| 343 |
+
|---|---|---|
|
| 344 |
+
| Claude Code, un utilisateur | **Qwen3-Coder-30B** | 575 tok/s en édition de code, cache de préfixe 16×, 0,21 $/1M |
|
| 345 |
+
| contexte au-delà de 262k | **Kimi-Linear-48B** | 1 048 576 tokens sur une carte, prefill à froid 1,7× meilleur |
|
| 346 |
+
| service multi-utilisateurs | Qwen sans spéculation | 1 166 tok/s à 32 flux, 0,105 $/1M |
|