Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Download bench_endpoint.sh from patdev/k3-a40-bootstrap: direct link, hf CLI and curl.
- Browser
- Download file 7.56 kB
-
https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/bench_endpoint.sh
- Command line
-
hf download hf://patdev/k3-a40-bootstrap/bench_endpoint.sh
-
curl -L -o bench_endpoint.sh https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/bench_endpoint.sh
7.56 kB
| # Protocole identique pour tous les endpoints OpenAI, afin que les chiffres | |
| # soient comparables entre modeles et entre machines. | |
| # usage: bash bench_endpoint.sh <base_url> <model_name> <hourly_usd> [label] | |
| set -uo pipefail | |
| B="${1:?base url}"; M="${2:?model name}"; HR="${3:-0.44}"; LBL="${4:-$M}" | |
| T=$(mktemp -d) | |
| trap 'rm -rf "$T"' EXIT | |
| py() { python -c "$@"; } | |
| hdr() { printf '\n=== %s\n' "$1"; } | |
| echo "########## $LBL ($B, modele '$M', ${HR} \$/h)" | |
| # ---------------------------------------------------------------- capacites | |
| hdr "CAPACITES" | |
| curl -s -m 20 "$B/v1/models" -o "$T/models.json" | |
| python - "$T/models.json" <<'PY' | |
| import json,sys | |
| d=json.load(open(sys.argv[1])) | |
| for m in d.get("data",[]): | |
| print(" id=%s max_model_len=%s" % (m["id"], m.get("max_model_len"))) | |
| PY | |
| # ---------------------------------------------------------------- qualite | |
| hdr "QUALITE" | |
| python - "$T/q1.json" "$M" <<'PY' | |
| import json,sys | |
| json.dump({"model":sys.argv[2],"max_tokens":120,"temperature":0,"messages":[ | |
| {"role":"user","content":"Write a Python function computing the nth Fibonacci number recursively. Code only."}]}, open(sys.argv[1],"w")) | |
| PY | |
| echo "-- code:" | |
| curl -s -m 300 "$B/v1/chat/completions" -H 'Content-Type: application/json' --data-binary @"$T/q1.json" \ | |
| | python -c "import json,sys;d=json.load(sys.stdin);print((d['choices'][0]['message'].get('content') or '')[:280] if 'choices' in d else d)" | |
| python - "$T/q2.json" "$M" <<'PY' | |
| import json,sys | |
| json.dump({"model":sys.argv[2],"max_tokens":150,"temperature":0,"messages":[ | |
| {"role":"user","content":"Explique en deux phrases ce qu'est un modele mixture-of-experts."}]}, open(sys.argv[1],"w")) | |
| PY | |
| echo "-- francais:" | |
| curl -s -m 300 "$B/v1/chat/completions" -H 'Content-Type: application/json' --data-binary @"$T/q2.json" \ | |
| | python -c "import json,sys;d=json.load(sys.stdin);print((d['choices'][0]['message'].get('content') or '')[:280] if 'choices' in d else d)" | |
| # ---------------------------------------------------------------- tool call | |
| hdr "TOOL CALLING" | |
| python - "$T/tc.json" "$M" <<'PY' | |
| import json,sys | |
| json.dump({"model":sys.argv[2],"max_tokens":300,"temperature":0, | |
| "messages":[{"role":"user","content":"Read the file /etc/hostname and then list the files in /tmp. Use the tools."}], | |
| "tools":[ | |
| {"type":"function","function":{"name":"read_file","description":"Read a file from disk","parameters":{"type":"object","properties":{"path":{"type":"string"}},"required":["path"]}}}, | |
| {"type":"function","function":{"name":"list_dir","description":"List a directory","parameters":{"type":"object","properties":{"path":{"type":"string"}},"required":["path"]}}}]}, | |
| open(sys.argv[1],"w")) | |
| PY | |
| curl -s -m 300 "$B/v1/chat/completions" -H 'Content-Type: application/json' --data-binary @"$T/tc.json" \ | |
| | python -c " | |
| import json,sys | |
| d=json.load(sys.stdin) | |
| if 'choices' not in d: print(' ERR', str(d)[:200]); raise SystemExit | |
| c=d['choices'][0]; tc=c['message'].get('tool_calls') or [] | |
| print(' finish=%s appels=%d' % (c['finish_reason'], len(tc))) | |
| for t in tc: print(' ', t['function']['name'], t['function']['arguments']) | |
| " | |
| # --------------------------------------------------- decode vs concurrence | |
| hdr "DECODE vs CONCURRENCE (128 tokens/flux)" | |
| for i in $(seq 1 32); do | |
| python - "$T/c$i.json" "$M" "$i" <<'PY' | |
| import json,sys | |
| json.dump({"model":sys.argv[2],"prompt":"[%s] Write a detailed technical explanation of how Mixture-of-Experts routing works in large language models."%sys.argv[3], | |
| "max_tokens":128,"temperature":0}, open(sys.argv[1],"w")) | |
| PY | |
| done | |
| printf " %-6s %-8s %-9s %-12s %-10s %s\n" flux tokens duree agrege /flux "\$/1M" | |
| for C in 1 4 8 16 32; do | |
| pids=(); t0=$(date +%s.%N) | |
| for i in $(seq 1 "$C"); do | |
| curl -s -m 900 "$B/v1/completions" -H 'Content-Type: application/json' --data-binary @"$T/c$i.json" -o "$T/r$i.json" & pids+=($!) | |
| done | |
| for p in "${pids[@]}"; do wait "$p"; done | |
| t1=$(date +%s.%N) | |
| tot=0 | |
| for i in $(seq 1 "$C"); do | |
| n=$(grep -o '"completion_tokens":[0-9]*' "$T/r$i.json" 2>/dev/null | head -1 | cut -d: -f2) | |
| tot=$((tot + ${n:-0})) | |
| done | |
| awk -v c="$C" -v g="$tot" -v a="$t0" -v b="$t1" -v h="$HR" \ | |
| 'BEGIN{d=b-a; t=g/d; printf " %-6d %-8d %-9.1f %-12.1f %-10.1f $%.3f\n", c, g, d, t, t/c, h/(t*3600)*1000000}' | |
| done | |
| # ---------------------------------------------------------------- prefill | |
| hdr "PREFILL" | |
| python - "$T/pf.json" "$M" <<'PY' | |
| import json,sys | |
| u=("Paragraph %d. Consensus protocols in distributed databases must tolerate partial " | |
| "failure, and quorum intersection guarantees two conflicting decisions cannot both commit. ") | |
| s="".join(u%i for i in range(1,1201))+chr(10)+"Summarise the above in one sentence:" | |
| json.dump({"model":sys.argv[2],"prompt":s,"max_tokens":4,"temperature":0}, open(sys.argv[1],"w")) | |
| PY | |
| for run in froid chaud; do | |
| t0=$(date +%s.%N) | |
| curl -s -m 900 "$B/v1/completions" -H 'Content-Type: application/json' --data-binary @"$T/pf.json" -o "$T/pf.out" | |
| t1=$(date +%s.%N) | |
| n=$(grep -o '"prompt_tokens":[0-9]*' "$T/pf.out" 2>/dev/null | head -1 | cut -d: -f2) | |
| awk -v r="$run" -v n="${n:-0}" -v a="$t0" -v b="$t1" \ | |
| 'BEGIN{d=b-a; printf " %-6s %7d tok en %6.2f s -> %8.0f tok/s\n", r, n, d, (d>0?n/d:0)}' | |
| done | |
| # ---------------------------------------------------------- edition de code | |
| hdr "EDITION DE CODE (regime ou la speculation n-gram gagne)" | |
| python - "$T/ed.json" "$M" <<'PY' | |
| import json,sys | |
| src="\n".join( | |
| "def handler_%d(request, context):\n" | |
| " payload = request.get('payload')\n" | |
| " if payload is None:\n" | |
| " raise ValueError('missing payload in handler_%d')\n" | |
| " result = context.process(payload, retries=3, timeout=%d)\n" | |
| " return {'status': 'ok', 'handler': %d, 'result': result}\n" % (i,i,i+5,i) | |
| for i in range(1,26)) | |
| p=("Here is a Python module:\n\n```python\n"+src+"\n```\n\nReturn the COMPLETE module " | |
| "unchanged except that every function gets the type hints " | |
| "`(request: dict, context: Any) -> dict`. Output only code.") | |
| json.dump({"model":sys.argv[2],"prompt":p,"max_tokens":900,"temperature":0}, open(sys.argv[1],"w")) | |
| PY | |
| t0=$(date +%s.%N) | |
| curl -s -m 900 "$B/v1/completions" -H 'Content-Type: application/json' --data-binary @"$T/ed.json" -o "$T/ed.out" | |
| t1=$(date +%s.%N) | |
| n=$(grep -o '"completion_tokens":[0-9]*' "$T/ed.out" 2>/dev/null | head -1 | cut -d: -f2) | |
| awk -v n="${n:-0}" -v a="$t0" -v b="$t1" -v h="$HR" \ | |
| 'BEGIN{d=b-a; t=n/d; printf " %d tok en %.2f s = %.2f tok/s -> $%.3f / 1M en mono-flux\n", n, d, t, h/(t*3600)*1000000}' | |
| # ------------------------------------------------------------- speculation | |
| hdr "TELEMETRIE SPECULATIVE" | |
| curl -s -m 25 "$B/metrics" -o "$T/m.txt" 2>/dev/null | |
| python - "$T/m.txt" <<'PY' | |
| import re,sys | |
| acc=dr=0; pos={} | |
| try: lines=open(sys.argv[1],encoding='utf-8',errors='ignore').read().splitlines() | |
| except Exception: lines=[] | |
| for l in lines: | |
| if l.startswith('#'): continue | |
| m=re.match(r'vllm:spec_decode_num_accepted_tokens_total\{.*?\}\s+([\d.e+]+)',l) | |
| if m: acc+=float(m.group(1)) | |
| m=re.match(r'vllm:spec_decode_num_draft_tokens_total\{.*?\}\s+([\d.e+]+)',l) | |
| if m: dr+=float(m.group(1)) | |
| m=re.match(r'vllm:spec_decode_num_accepted_tokens_per_pos_total\{.*?position="(\d+)".*?\}\s+([\d.e+]+)',l) | |
| if m: pos[int(m.group(1))]=pos.get(int(m.group(1)),0)+float(m.group(2)) | |
| if dr: | |
| print(" brouillons %.0f acceptes %.0f = %.1f %%" % (dr,acc,100*acc/dr)) | |
| if pos: print(" par position:", " / ".join("%.0f"%pos[k] for k in sorted(pos))) | |
| else: | |
| print(" aucune telemetrie speculative (speculation desactivee ou inutilisee)") | |
| PY | |
| echo | |