Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Upload FINDINGS.md with huggingface_hub
Browse files- FINDINGS.md +54 -0
FINDINGS.md
CHANGED
|
@@ -531,3 +531,57 @@ Cloudflare sits in front of the pod proxy and returns `error code: 524`. A cold
|
|
| 531 |
prefill longer than that cannot complete in one call — measured at ~40k on
|
| 532 |
Kimi and ~232k on Qwen. Chunk the prompt (each request extends the cached
|
| 533 |
prefix) or expose a TCP port instead.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 531 |
prefill longer than that cannot complete in one call — measured at ~40k on
|
| 532 |
Kimi and ~232k on Qwen. Chunk the prompt (each request extends the cached
|
| 533 |
prefix) or expose a TCP port instead.
|
| 534 |
+
|
| 535 |
+
---
|
| 536 |
+
|
| 537 |
+
# Session 4 — speculative decoding, and why the regime decides everything
|
| 538 |
+
|
| 539 |
+
n-gram speculative decoding on Qwen3-Coder-30B-A3B-AWQ, vLLM 0.27.1, 1x A40,
|
| 540 |
+
`--speculative-config '{"method":"ngram","num_speculative_tokens":5,...}'`.
|
| 541 |
+
Same model, same hardware, same quant, two pods running side by side so the
|
| 542 |
+
comparison is simultaneous rather than sequential:
|
| 543 |
+
|
| 544 |
+
```
|
| 545 |
+
A. pure generation (no prompt/output overlap)
|
| 546 |
+
baseline 102.21 tok/s
|
| 547 |
+
ngram 58.49 tok/s -43%
|
| 548 |
+
|
| 549 |
+
B. code editing (output re-emits the input)
|
| 550 |
+
baseline 130.16 tok/s
|
| 551 |
+
ngram 242.39 tok/s +86%
|
| 552 |
+
```
|
| 553 |
+
|
| 554 |
+
**Speculation is not a free win — it is a bet on repetition.** With nothing to
|
| 555 |
+
guess, every draft is rejected and the verification cost is pure loss; that is
|
| 556 |
+
the -43%. Claude Code lives almost entirely in regime B (read a file, emit a
|
| 557 |
+
modified version, repeat identifiers), so it is the right default *there* and
|
| 558 |
+
the wrong default for prose.
|
| 559 |
+
|
| 560 |
+
vLLM's telemetry, aggregated over both regimes:
|
| 561 |
+
|
| 562 |
+
```
|
| 563 |
+
drafts 207 draft tokens 1035 accepted 706 = 68.2%
|
| 564 |
+
accepted per position: 179 / 156 / 131 / 120 / 120
|
| 565 |
+
mean accepted length 3.41 -> ~4.4 tokens emitted per forward pass
|
| 566 |
+
```
|
| 567 |
+
|
| 568 |
+
This is the measurement llama.cpp never produced: DSpark there exposed no
|
| 569 |
+
`draft_n` at all and moved throughput by 0.0%. The mechanism was never broken —
|
| 570 |
+
the engine was.
|
| 571 |
+
|
| 572 |
+
## The number that matters for the original goal
|
| 573 |
+
|
| 574 |
+
Single stream, code editing, one A40 at $0.44/h:
|
| 575 |
+
|
| 576 |
+
```
|
| 577 |
+
242.39 tok/s -> $0.50 per 1M output tokens
|
| 578 |
+
```
|
| 579 |
+
|
| 580 |
+
The under-$1/1M target is met **on a single stream**, without needing
|
| 581 |
+
concurrency to amortise anything. For reference the same target on Kimi-K3
|
| 582 |
+
required 611 tok/s aggregate against 14.5 measured.
|
| 583 |
+
|
| 584 |
+
Note also that the baseline itself reads higher here (102-130 tok/s) than the
|
| 585 |
+
77.6 measured earlier: that earlier figure was a 128-token request whose wall
|
| 586 |
+
clock was dominated by per-request overhead. Longer outputs amortise it. Quote
|
| 587 |
+
77.6 for short replies and ~130 for sustained generation.
|