Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Upload k3_bootstrap.sh with huggingface_hub
Browse files- k3_bootstrap.sh +11 -0
k3_bootstrap.sh
CHANGED
|
@@ -305,7 +305,18 @@ c = json.loads(p.read_text()); old = c.get("architectures")
|
|
| 305 |
c["architectures"] = ["Qwen3DSparkModel"]; p.write_text(json.dumps(c, indent=2))
|
| 306 |
print(f"architectures {old} -> {c['architectures']}")
|
| 307 |
PY
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 308 |
python3 "$SRC/convert_hf_to_gguf.py" /models/dspark-src \
|
|
|
|
| 309 |
--outfile "$DRAFT_DIR/$DS_FIXED" --outtype bf16 2>&1 | tail -12
|
| 310 |
if [ -s "$DRAFT_DIR/$DS_FIXED" ]; then
|
| 311 |
log "DSpark draft built: $(du -h "$DRAFT_DIR/$DS_FIXED" | cut -f1)"
|
|
|
|
| 305 |
c["architectures"] = ["Qwen3DSparkModel"]; p.write_text(json.dumps(c, indent=2))
|
| 306 |
print(f"architectures {old} -> {c['architectures']}")
|
| 307 |
PY
|
| 308 |
+
# The converter needs the TARGET model's tokenizer to write the draft's
|
| 309 |
+
# vocabulary ("DFlash draft model requires --target-model-dir"). Only the
|
| 310 |
+
# tokenizer files are fetched -- five small files, not the 2.8 TB of
|
| 311 |
+
# weights.
|
| 312 |
+
mkdir -p /models/k3-tok
|
| 313 |
+
for f in config.json tokenizer_config.json tiktoken.model \
|
| 314 |
+
tokenization_kimi.py encoding_k3.py; do
|
| 315 |
+
hf download moonshotai/Kimi-K3 "$f" --local-dir /models/k3-tok >/dev/null 2>&1
|
| 316 |
+
done
|
| 317 |
+
log "target tokenizer files: $(ls /models/k3-tok 2>/dev/null | wc -l)/5"
|
| 318 |
python3 "$SRC/convert_hf_to_gguf.py" /models/dspark-src \
|
| 319 |
+
--target-model-dir /models/k3-tok \
|
| 320 |
--outfile "$DRAFT_DIR/$DS_FIXED" --outtype bf16 2>&1 | tail -12
|
| 321 |
if [ -s "$DRAFT_DIR/$DS_FIXED" ]; then
|
| 322 |
log "DSpark draft built: $(du -h "$DRAFT_DIR/$DS_FIXED" | cut -f1)"
|