Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Upload RESULTS.md with huggingface_hub
Browse files- RESULTS.md +45 -0
RESULTS.md
CHANGED
|
@@ -570,3 +570,48 @@ quatre leviers testés n'a bougé le débit.
|
|
| 570 |
Le modèle non élagué est conservé : le REAP coûte 20 % des experts pour zéro
|
| 571 |
gain mesuré, donc de la qualité perdue sans contrepartie. Il reste disponible
|
| 572 |
(`VL_MODEL=reap`) au cas où les 4 Go de VRAM libérés seraient utiles au KV.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 570 |
Le modèle non élagué est conservé : le REAP coûte 20 % des experts pour zéro
|
| 571 |
gain mesuré, donc de la qualité perdue sans contrepartie. Il reste disponible
|
| 572 |
(`VL_MODEL=reap`) au cas où les 4 Go de VRAM libérés seraient utiles au KV.
|
| 573 |
+
|
| 574 |
+
---
|
| 575 |
+
|
| 576 |
+
# Le modèle dense : quatrième hypothèse réfutée
|
| 577 |
+
|
| 578 |
+
Expérience discriminante. Si le plafond venait de la dispersion des tokens vers
|
| 579 |
+
128 experts, un modèle **dense** de taille totale comparable devait se comporter
|
| 580 |
+
différemment. `cyankiwi/Qwen3.8-27B-AWQ-INT4` : 19,6 Gio, aucun expert.
|
| 581 |
+
|
| 582 |
+
```
|
| 583 |
+
sessions MoE Coder-30B Dense 27B ecart
|
| 584 |
+
1 122,9 26,2 4,7x plus lent
|
| 585 |
+
8 512,1 167,1 3,1x plus lent
|
| 586 |
+
32 1 159,7 320,8 3,6x plus lent
|
| 587 |
+
|
| 588 |
+
montee 32/1 9,4x 12,2x
|
| 589 |
+
```
|
| 590 |
+
|
| 591 |
+
Le dense monte légèrement mieux en concurrence, mais part de si bas que ça ne
|
| 592 |
+
compense jamais. **La dispersion MoE n'est pas le mur — c'est l'avantage** : le
|
| 593 |
+
MoE ne lit que ~5 Gio d'actifs par token là où le dense en lit 19,6.
|
| 594 |
+
|
| 595 |
+
## Ce que le plafond n'est pas
|
| 596 |
+
|
| 597 |
+
Cinq mécanismes testés, cinq réfutations :
|
| 598 |
+
|
| 599 |
+
```
|
| 600 |
+
lecture des poids -24 % de poids -> 0 % de debit
|
| 601 |
+
calcul brut MoE ~5 %, dense ~11 % du plafond FLOPS
|
| 602 |
+
bande passante MoE ~45 %, dense ~30 % des 696 Go/s
|
| 603 |
+
dispersion MoE le dense est 3 a 4x plus lent
|
| 604 |
+
ordonnancement CPU --async-scheduling dans le bruit
|
| 605 |
+
speculation perdante a n=24 ET a n=4
|
| 606 |
+
```
|
| 607 |
+
|
| 608 |
+
**Aucun mécanisme validé** n'explique le plafond de ~1 160 tok/s. Ce qui est
|
| 609 |
+
établi, c'est ce qu'il n'est pas. Publier cette liste vaut mieux qu'habiller une
|
| 610 |
+
sixième hypothèse en explication.
|
| 611 |
+
|
| 612 |
+
## Configuration retenue
|
| 613 |
+
|
| 614 |
+
`cyankiwi/Qwen3-Coder-30B-A3B-Instruct-AWQ-4bit`, spéculation coupée,
|
| 615 |
+
ordonnancement asynchrone, `max-num-seqs 64`, KV en fp8, contexte 262 144.
|
| 616 |
+
C'est la meilleure combinaison mesurée, et les six variantes essayées sont
|
| 617 |
+
documentées ci-dessus pour qu'on ne les retente pas.
|