Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Download aides/_arret.py from patdev/k3-a40-bootstrap: direct link, hf CLI and curl.
- Browser
- Download file 2.1 kB
-
https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/aides/_arret.py
- Command line
-
hf download hf://patdev/k3-a40-bootstrap/aides/_arret.py
-
curl -L -o _arret.py https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/aides/_arret.py
2.1 kB
| # Arrete le chien et le moteur en lisant /proc, jamais par motif. | |
| # | |
| # L'agent du pod porte le TEXTE des scripts dans sa ligne de commande : un | |
| # `pkill -f "vllm serve"` matche donc l'agent lui-meme et se suicide (constate, | |
| # `[FIN] code=-15`). On exclut toute notre ascendance a la place. | |
| # | |
| # `bash /run.sh` est intouchable : c'est le processus dont depend le conteneur, | |
| # quel que soit son PID. | |
| import os | |
| import signal | |
| import time | |
| def ascendance(pid): | |
| vus = set() | |
| while pid and pid > 1: | |
| vus.add(pid) | |
| try: | |
| pid = int(open("/proc/%d/stat" % pid).read().rsplit(") ", 1)[1].split()[1]) | |
| except Exception: | |
| break | |
| return vus | |
| sur = ascendance(os.getpid()) | |
| def cibles(): | |
| trouves = [] | |
| for entree in os.listdir("/proc"): | |
| if not entree.isdigit(): | |
| continue | |
| pid = int(entree) | |
| if pid in sur: | |
| continue | |
| try: | |
| argv = [a for a in open("/proc/%d/cmdline" % pid, "rb").read().split(b"\0") if a] | |
| comm = open("/proc/%d/comm" % pid).read().strip() | |
| except Exception: | |
| continue | |
| if not argv: | |
| continue | |
| ligne = b" ".join(argv) | |
| if b"/run.sh" in ligne: | |
| continue | |
| chien = b"chien.py" in ligne | |
| # EngineCore se renomme et survit a un kill par nom de commande en | |
| # retenant toute la VRAM : le viser aussi par sa ligne complete. | |
| moteur = ((b"vllm" in argv[0] and b"serve" in ligne) | |
| or argv[0].endswith(b"vllm") | |
| or comm in ("VLLM::EngineCore", "EngineCore", "PleOffloadWorker") | |
| or b"VLLM::" in ligne) | |
| if chien or moteur: | |
| trouves.append((pid, comm)) | |
| return trouves | |
| c = cibles() | |
| print(" arret de :", [(p, n) for p, n in c] or "rien") | |
| for pid, _ in c: | |
| try: | |
| os.kill(pid, signal.SIGTERM) | |
| except Exception: | |
| pass | |
| time.sleep(8) | |
| for pid, _ in cibles(): | |
| try: | |
| os.kill(pid, signal.SIGKILL) | |
| except Exception: | |
| pass | |
| print(" restants :", [p for p, _ in cibles()] or "aucun") | |