Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Download banc_couples_kernels.py from patdev/k3-a40-bootstrap: direct link, hf CLI and curl.
- Browser
- Download file 9.3 kB
-
https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/banc_couples_kernels.py
- Command line
-
hf download hf://patdev/k3-a40-bootstrap/banc_couples_kernels.py
-
curl -L -o banc_couples_kernels.py https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/banc_couples_kernels.py
9.3 kB
| """Couples (moe_backend, linear_backend) sur sm_120. | |
| Ce que les essais precedents ont manque. `--moe-backend` ne pilote que les | |
| experts. Le chemin dense pese 1,849 Gio des 2,880 Gio du socle actif, soit | |
| **64 %**, et il passe par un noyau LINEAIRE distinct : meme avec | |
| `--moe-backend humming`, vLLM affichait toujours | |
| `Using MarlinNvFp4LinearKernel for NVFP4 GEMM`. On faisait donc varier un tiers | |
| du travail en croyant faire varier le tout -- ce qui explique tres bien | |
| l'egalite Marlin/humming a 1 % pres sur trois architectures. | |
| `config/kernel.py` documente les deux champs : | |
| moe_backend: auto | triton | cutlass | flashinfer_trtllm | | |
| flashinfer_cutlass | flashinfer_cutedsl | | |
| flashinfer_b12x | marlin | humming | emulation | ... | |
| linear_backend: auto | cutlass | flashinfer_cutlass | flashinfer_cutedsl | | |
| flashinfer_trtllm | flashinfer_cudnn | flashinfer_b12x | | |
| marlin | triton | deep_gemm | |
| `flashinfer_b12x` n'apparait PAS dans la liste "out of potential backends" du | |
| journal, mais la configuration le documente des deux cotes, et explicitement | |
| pour notre carte : | |
| moe : "Use FlashInfer CuteDSL fused MoE for SM12x (RTX Pro 6000 / DGX Spark)" | |
| linear : "Use FlashInfer b12x CuteDSL NVFP4 GEMM (SM120+)" | |
| Se fier a l'enumeration affichee plutot qu'a la configuration reelle nous | |
| l'avait fait manquer. | |
| Reperes sur cette carte, contexte 131 072, recette NVIDIA : marlin 275,1 solo, | |
| humming 272,5. Sur H200 : 351,6 et 346,4. | |
| """ | |
| import json | |
| import re | |
| import statistics | |
| import subprocess | |
| import threading | |
| import time | |
| import urllib.request | |
| MODEL = "nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4" | |
| PORT = 8000 | |
| URL = "http://127.0.0.1:%d" % PORT | |
| # (moe_backend, linear_backend) ; None = laisser vLLM choisir. | |
| COUPLES = [ | |
| ("marlin", None), # la reference qu'on subit | |
| (None, "flashinfer_b12x"), # le lineaire dedie sm_120 | |
| (None, "flashinfer_cudnn"), | |
| (None, "flashinfer_cutlass"), | |
| (None, "flashinfer_cutedsl"), | |
| (None, "cutlass"), | |
| (None, "humming"), | |
| ("flashinfer_b12x", None), # le MoE dedie sm_120 | |
| ("flashinfer_b12x", "flashinfer_b12x"), # la combinaison visee | |
| ] | |
| BASE = ["vllm", "serve", MODEL, | |
| "--served-model-name", "ornith", | |
| "--host", "127.0.0.1", "--port", str(PORT), | |
| "--trust-remote-code", | |
| "--max-model-len", "131072", | |
| "--kv-cache-dtype", "fp8", | |
| "--enable-prefix-caching", | |
| "--gpu-memory-utilization", "0.85", | |
| "--mamba-backend", "flashinfer", | |
| "--mamba-cache-mode", "align", | |
| "--reasoning-parser", "nemotron_v3", | |
| "--tool-call-parser", "qwen3_coder", | |
| "--enable-auto-tool-choice"] | |
| SUJETS = ["un cache LRU avec dict et liste doublement chainee", | |
| "un pool de connexions avec expiration et sante des sockets", | |
| "un analyseur d'expressions arithmetiques par descente recursive", | |
| "une file de priorite par tas binaire avec decrease-key"] | |
| def dire(*a): | |
| print(*a, flush=True) | |
| dire("=" * 74) | |
| subprocess.run(["nvidia-smi", "--query-gpu=name,memory.total,compute_cap", | |
| "--format=csv,noheader"], check=False) | |
| subprocess.run(["python3", "-c", | |
| "import vllm,torch;print('vllm',vllm.__version__,'torch',torch.__version__," | |
| "'cap',torch.cuda.get_device_capability(0))"], check=False) | |
| dire("=" * 74) | |
| def demarrer(couple, journal): | |
| moe, lin = couple | |
| sup = [] | |
| if moe: | |
| sup += ["--moe-backend", moe] | |
| if lin: | |
| sup += ["--linear-backend", lin] | |
| with open(journal, "w") as f: | |
| p = subprocess.Popen(BASE + sup, stdout=f, stderr=subprocess.STDOUT) | |
| for i in range(75): | |
| try: | |
| urllib.request.urlopen(URL + "/v1/models", timeout=5).read() | |
| return p, i * 10 | |
| except Exception: | |
| pass | |
| if p.poll() is not None: | |
| return None, i * 10 | |
| time.sleep(10) | |
| p.terminate() | |
| return None, 750 | |
| def une(sujet, res, i): | |
| corps = json.dumps({ | |
| "model": "ornith", | |
| "messages": [{"role": "user", | |
| "content": "Ecris en Python %s, avec trois tests unittest." % sujet}], | |
| "max_tokens": 300, "temperature": 0.0, "stream": True}).encode() | |
| r = urllib.request.Request(URL + "/v1/chat/completions", data=corps, | |
| headers={"Content-Type": "application/json"}) | |
| t1 = None | |
| n = 0 | |
| bouts = [] | |
| try: | |
| with urllib.request.urlopen(r, timeout=600) as rep: | |
| for l in rep: | |
| l = l.strip() | |
| if not l.startswith(b"data: ") or l[6:] == b"[DONE]": | |
| continue | |
| ch = (json.loads(l[6:]).get("choices") or [{}])[0] | |
| de = ch.get("delta", {}) or {} | |
| x = de.get("content") or de.get("reasoning") or de.get("reasoning_content") | |
| if x: | |
| if t1 is None: | |
| t1 = time.time() | |
| n += 1 | |
| bouts.append(x) | |
| except Exception as e: | |
| res[i] = {"err": "%s: %s" % (type(e).__name__, str(e)[:70])} | |
| return | |
| res[i] = {"n": n, "t1": t1, "t2": time.time(), "txt": "".join(bouts)} | |
| def div4(t): | |
| m = t.split() | |
| if len(m) < 40: | |
| return 1.0 | |
| g = [tuple(m[i:i + 4]) for i in range(len(m) - 3)] | |
| return len(set(g)) / len(g) | |
| def mesurer(conc): | |
| res = [None] * conc | |
| d0 = time.time() | |
| fils = [threading.Thread(target=une, args=(SUJETS[i % len(SUJETS)], res, i)) | |
| for i in range(conc)] | |
| for f in fils: | |
| f.start() | |
| for f in fils: | |
| f.join() | |
| d1 = time.time() | |
| bons = [r for r in res if r and not r.get("err") and r.get("t1")] | |
| if not bons: | |
| return None | |
| return (sum(r["n"] for r in bons) / (d1 - d0), | |
| statistics.median([(r["n"] - 1) / (r["t2"] - r["t1"]) | |
| for r in bons if r["t2"] > r["t1"]]), | |
| statistics.median([div4(r["txt"]) for r in bons])) | |
| resume = [] | |
| for idx, couple in enumerate(COUPLES): | |
| moe, lin = couple | |
| etiq = "moe=%-17s lin=%s" % (moe or "auto", lin or "auto") | |
| dire("\n" + "=" * 74) | |
| dire("%d. %s" % (idx + 1, etiq)) | |
| dire("=" * 74) | |
| journal = "/tmp/k_%d.log" % idx | |
| proc, secondes = demarrer(couple, journal) | |
| texte = open(journal, errors="replace").read() | |
| # Demander n'est pas obtenir : on releve LES DEUX noyaux effectivement | |
| # retenus. C'est l'erreur de l'essai precedent -- `--moe-backend humming` | |
| # affichait bien HUMMING cote experts et Marlin cote dense, et seule la | |
| # premiere ligne avait ete lue. | |
| moe_retenu = lin_retenu = None | |
| for ligne in texte.splitlines(): | |
| if "ERROR" in ligne: | |
| continue | |
| m1 = re.search(r"Using '?(\w+)'? NvFp4 MoE backend", ligne) | |
| if m1: | |
| moe_retenu = m1.group(1) | |
| m2 = re.search(r"Using (\w+) for NVFP4 GEMM", ligne) | |
| if m2: | |
| lin_retenu = m2.group(1) | |
| dire(" MoE retenu : %s" % (moe_retenu or "-")) | |
| dire(" lineaire retenu : %s" % (lin_retenu or "-")) | |
| for ligne in texte.splitlines(): | |
| if "GPU KV cache size" in ligne: | |
| dire(" " + ligne.split("] ")[-1][:140]) | |
| break | |
| couple_retenu = "%s / %s" % (moe_retenu or "-", lin_retenu or "-") | |
| if not proc: | |
| dire(" NE DEMARRE PAS (%d s)" % secondes) | |
| vu = set() | |
| for ligne in texte.splitlines(): | |
| if any(m in ligne for m in ("RuntimeError", "ValueError", "Traceback", | |
| "unrecognized arguments", "invalid choice", | |
| "NotImplementedError", "AssertionError", | |
| "is not supported", "does not support")): | |
| t = ligne.split("] ")[-1][:165] | |
| if t not in vu: | |
| vu.add(t) | |
| dire(" > " + t) | |
| if len(vu) >= 5: | |
| break | |
| resume.append((etiq, couple_retenu, None)) | |
| continue | |
| dire(" PRET en %d s" % secondes) | |
| dire("conc | agrege | par flux | 4-gr") | |
| solo = None | |
| for conc in (1, 4): | |
| d = mesurer(conc) | |
| if not d: | |
| dire("%4d | ECHEC" % conc) | |
| continue | |
| ag, pf, dv = d | |
| dire("%4d | %8.1f | %8.1f | %.3f %s" | |
| % (conc, ag, pf, dv, "" if dv > 0.6 else " DEGENERE")) | |
| if conc == 1: | |
| solo = pf | |
| resume.append((etiq, couple_retenu, solo)) | |
| proc.terminate() | |
| time.sleep(20) | |
| dire("\n" + "=" * 74) | |
| dire("RESUME -- couples de noyaux NVFP4") | |
| dire("=" * 74) | |
| dire("%-34s %-36s %9s" % ("demande", "retenu (MoE / lineaire)", "solo")) | |
| ref = resume[0][2] if resume and resume[0][2] else None | |
| for etiq, retenu, solo in resume: | |
| d = "" | |
| if ref and solo: | |
| d = " %+5.1f %%" % (100 * (solo - ref) / ref) | |
| dire("%-34s %-36s %9s%s" | |
| % (etiq, retenu, ("%.1f" % solo) if solo else "ne demarre pas", d)) | |
| dire("\nUn 'retenu' different du 'demande' signifie que vLLM a ignore le drapeau :") | |
| dire("le chiffre n'est alors PAS celui du noyau demande et ne prouve rien sur lui.") | |
| dire("reperes : marlin 275,1 sur cette carte | 351,6 sur H200.") | |