Instructions to use tchbcb/samai-8b-M8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tchbcb/samai-8b-M8 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tchbcb/samai-8b-M8:Q4_K_M # Run inference directly in the terminal: llama cli -hf tchbcb/samai-8b-M8:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tchbcb/samai-8b-M8:Q4_K_M # Run inference directly in the terminal: llama cli -hf tchbcb/samai-8b-M8:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tchbcb/samai-8b-M8:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf tchbcb/samai-8b-M8:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tchbcb/samai-8b-M8:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf tchbcb/samai-8b-M8:Q4_K_M
Use Docker
docker model run hf.co/tchbcb/samai-8b-M8:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use tchbcb/samai-8b-M8 with Ollama:
ollama run hf.co/tchbcb/samai-8b-M8:Q4_K_M
- Unsloth Desktop
- Pi
How to use tchbcb/samai-8b-M8 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-8b-M8:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tchbcb/samai-8b-M8:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tchbcb/samai-8b-M8 with Docker Model Runner:
docker model run hf.co/tchbcb/samai-8b-M8:Q4_K_M
- Lemonade
How to use tchbcb/samai-8b-M8 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tchbcb/samai-8b-M8:Q4_K_M
Run and chat with the model
lemonade run user.samai-8b-M8-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use tchbcb/samai-8b-M8 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-8b-M8:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tchbcb/samai-8b-M8:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tchbcb/samai-8b-M8 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-8b-M8:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tchbcb/samai-8b-M8:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download artifacts/r13_scripts/patch_refine_v2.py from tchbcb/samai-8b-M8: direct link, hf CLI and curl.
- Browser
- Download file 3.83 kB
-
https://huggingface.co/tchbcb/samai-8b-M8/resolve/main/artifacts/r13_scripts/patch_refine_v2.py
- Command line
-
hf download hf://tchbcb/samai-8b-M8/artifacts/r13_scripts/patch_refine_v2.py
-
curl -L -o patch_refine_v2.py https://huggingface.co/tchbcb/samai-8b-M8/resolve/main/artifacts/r13_scripts/patch_refine_v2.py
3.83 kB
| #!/usr/bin/env python3 | |
| # -*- coding: utf-8 -*- | |
| """patch_refine_v2.py — refine_tensor 分 slab 显存安全版 | |
| 对角 H → 损失逐块可分 → 16M 权重/slab 独立优化(数学等价), 峰值 VRAM ~3GB (原 ~13GB 驱动级 OOM) | |
| """ | |
| NEWFN = '''def refine_tensor2(W, hv, typ, steps, dev, wps=16000000): | |
| """W (out,in) fp32 cuda; hv (in,) normalized diag weights. Slab-chunked GSQ-lite.""" | |
| import torch | |
| div = 8.0 if typ == "q4_0" else 16.0 | |
| out_ch, in_ch = W.shape | |
| blk = 32 | |
| Wb = W.reshape(-1, blk) | |
| nb = Wb.shape[0] | |
| hvf = hv.to(dev).float().clamp_min(1e-12) | |
| hvn = hvf / hvf.sum() | |
| hv_blk = hvn.repeat(out_ch, 1).reshape(-1, blk) | |
| codes_full = torch.empty(nb, blk, dtype=torch.int32, device=dev) | |
| d_full = torch.empty(nb, dtype=torch.float32, device=dev) | |
| bpb = max(1, (wps // 4) // blk) | |
| n_slabs = (nb + bpb - 1) // bpb | |
| for s in range(n_slabs): | |
| b0, b1 = s * bpb, min(nb, (s + 1) * bpb) | |
| Ws = Wb[b0:b1] | |
| ns = Ws.shape[0] | |
| hbs = hv_blk[b0:b1] | |
| d0 = Ws.abs().amax(dim=1) / div | |
| d0 = torch.where(d0 < 1e-12, torch.ones_like(d0), d0) | |
| c0 = torch.clamp(torch.round(Ws / d0.unsqueeze(1)), -div, div - 1) | |
| grid = torch.stack([c0 - 2.0 + k for k in range(5)], dim=-1) | |
| grid = torch.clamp(grid, -div, div - 1) | |
| logits = torch.zeros(ns, blk, 5, device=dev, requires_grad=True) | |
| logits.data[:, :, 2] = 0.5 | |
| log_d = torch.log(d0).to(dev).requires_grad_(True) | |
| opt = torch.optim.AdamW([ | |
| {"params": [logits], "lr": 3e-3}, | |
| {"params": [log_d], "lr": 1e-3, "weight_decay": 0.0}]) | |
| sched = torch.optim.lr_scheduler.CosineAnnealingLR(opt, T_max=steps) | |
| d0min, d0max = 0.3 * d0, 3.0 * d0 | |
| for ep in range(steps): | |
| prog = ep / max(1, steps - 1) | |
| tau = 2.0 * (0.05 / 2.0) ** prog | |
| kap = 200.0 ** prog | |
| d = torch.clamp(torch.exp(log_d), d0min, d0max) | |
| g = -torch.log(-torch.log(torch.rand_like(logits) + 1e-9) + 1e-9) | |
| ys = torch.softmax((kap * logits + g) / tau, dim=-1) | |
| yh = torch.nn.functional.one_hot(ys.argmax(-1), 5).float() | |
| y = yh.detach() - ys.detach() + ys | |
| Q = (y * grid).sum(-1) * d.unsqueeze(1) | |
| E = Ws - Q | |
| hl = ((E * E) * hbs).sum() | |
| mse = (E * E).mean() | |
| loss = hl + 0.02 * mse | |
| opt.zero_grad(set_to_none=True) | |
| loss.backward() | |
| if prog < 0.25: | |
| log_d.grad = None | |
| opt.step() | |
| sched.step() | |
| with torch.no_grad(): | |
| d = torch.clamp(torch.exp(log_d), d0min, d0max) | |
| yh = torch.nn.functional.one_hot(logits.argmax(-1), 5).float() | |
| codes_s = (yh * (grid + div)).sum(-1).round().clamp(0, 2 * div - 1).to(torch.int32) | |
| codes_full[b0:b1] = codes_s | |
| d_full[b0:b1] = d | |
| del opt, sched, logits, log_d, grid, ys, yh, y, g, Q, E, loss | |
| if dev == "cuda": | |
| torch.cuda.empty_cache() | |
| return d_full.cpu(), codes_full.reshape(-1).cpu(), int(steps) | |
| ''' | |
| CALL_OLD = ''' d, codes, st = refine_tensor(W, H, typ, steps, dev)''' | |
| CALL_NEW = ''' hv = torch.diagonal(H).clone().clamp_min(1e-12) | |
| d, codes, st = refine_tensor2(W, hv, typ, steps, dev)''' | |
| def main(): | |
| path = "/tmp/k8b/r13_refine.py" | |
| src = open(path).read() | |
| anchor = "def np_pack_blocks(d_np, codes_np, typ):" | |
| assert src.count(anchor) == 1, "fn anchor %d" % src.count(anchor) | |
| src = src.replace(anchor, NEWFN + anchor, 1) | |
| assert src.count(CALL_OLD) == 1, "call anchor %d" % src.count(CALL_OLD) | |
| src = src.replace(CALL_OLD, CALL_NEW, 1) | |
| compile(src, path, "exec") | |
| open(path, "w").write(src) | |
| print("PATCH_OK refine_tensor2 slabbed") | |
| main() | |