Instructions to use tchbcb/samai-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tchbcb/samai-9b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tchbcb/samai-9b:Q4_K_M # Run inference directly in the terminal: llama cli -hf tchbcb/samai-9b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tchbcb/samai-9b:Q4_K_M # Run inference directly in the terminal: llama cli -hf tchbcb/samai-9b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tchbcb/samai-9b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf tchbcb/samai-9b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tchbcb/samai-9b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf tchbcb/samai-9b:Q4_K_M
Use Docker
docker model run hf.co/tchbcb/samai-9b:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use tchbcb/samai-9b with Ollama:
ollama run hf.co/tchbcb/samai-9b:Q4_K_M
- Unsloth Desktop
- Pi
How to use tchbcb/samai-9b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-9b:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tchbcb/samai-9b:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tchbcb/samai-9b with Docker Model Runner:
docker model run hf.co/tchbcb/samai-9b:Q4_K_M
- Lemonade
How to use tchbcb/samai-9b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tchbcb/samai-9b:Q4_K_M
Run and chat with the model
lemonade run user.samai-9b-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use tchbcb/samai-9b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-9b:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tchbcb/samai-9b:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tchbcb/samai-9b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-9b:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tchbcb/samai-9b:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download transfer.md from tchbcb/samai-9b: direct link, hf CLI and curl.
- Browser
- Download file 8.07 kB
-
https://huggingface.co/tchbcb/samai-9b/resolve/main/transfer.md
- Command line
-
hf download hf://tchbcb/samai-9b/transfer.md
-
curl -L -o transfer.md https://huggingface.co/tchbcb/samai-9b/resolve/main/transfer.md
SAMAI-8B 交接文档 (transfer.md) — R14 轮
更新: 2026-09-25 | 交接人: Super Z 主监督者 | 读者: 下任监督者/未来的自己 项目一句话: 在 Qwen3.5-9B VL 底模上用 LoRA 血统链(m13_r6→r7→r8→r9→r10)训练 agent 行为纪律, 配 zagent 框架级硬约束,闭环「采集→建窗→训练→部署→实测」。 本轮状态: R13 GSQ-RCO-lite 量化收官(gsqrco_v1 ppl 7.99 现役 serving, 旧 mixbit 8.11); R14 单隧道复用公网全绿(aitun.cc/8RF5ST7K/v1, key=1234, SSE 通); S4/S5 实测按用户令顺延。
一、训练什么
模型血统链(每代 = 底模 + 上一代 adapter 续训出来的新 LoRA):
Qwen/Qwen3.5-9B (底模, 视觉塔冻结)
└─ m13_r6 (M15 反死环 147 窗)
└─ m13_r7 (R7 反死环根治 13+41 窗)
└─ m13_r8 (R8 两振出局+think 纪律 235 窗)
└─ m13_r9 (R11 长上下文反环 210+16 窗)
└─ m13_r10 (R12 认知诚实+纯文本收束 268 窗) ← 血统终点, 量化盘母体
- 血统机制:
s8_m13_train_v9.py加载底模后PeftModel.from_pretrained(model, 上一代adapter, is_trainable=True)续训。 因此 r10 adapter 权重已含 r6..r10 全链血统,部署时只需「底模 ⊕ r10 adapter」一次 merge,数学等价。 - LoRA 配置: r=32, alpha=64, dropout=0.05, bias=none, task_type=CAUSAL_LM(vision tower + merger 冻结)。
- adapter 全集:
tchbcb/samai-8b-M8/m12_adapters/m13_r{6..10}/。
二、训练数据(268 窗, R12 期固化)
R11 210 窗(拦截后换路/收束/两振 + 视频窗)+ R12 新 58 窗(H1 认知诚实 20 / H2a 纯文本收束 16 /
H2b 空 tool_call 纠偏 8 / H2c 拦截后换路 8 / H3 负控 6)= samples_r12_train.jsonl。
窗格式与 R8 antiloop2 对齐; 深度 1200/2400/3600/4500; 构建器 r12_data_build.py。
训练要点: MAXLEN 3400 起步(6144 确定性 OOM, 坑#66)、fp16 禁 bf16、分块 fp32 CE、
NAN 自杀保护 + 时间熔断。详见 train.md v4.6 §三。
三、R13 量化战报(GSQ-RCO-lite, 现行最优)
用户否决 mixbit_4x8("量化技术不好")→ 参考 GSQ(Gumbel-Softmax Quantization, arXiv 2604.18556)
- RCO(arXiv 2605.00649, ISTA-DASLab)自适配 GSQ-RCO-lite 四段链: 链重建 → imatrix 校准(q8 路线 ppl 5.0709)→ DP 背包逐张量类型分配 → --imatrix+类型表容器量化。
| 文件(tchbcb/samai-9b) | 体积 | ppl | sha256 前 16 | 角色 |
|---|---|---|---|---|
| m13_r10_gsqrco_v1.gguf | 6.482G | 7.99(旧 mixbit 8.11) | 23a73197243cb9bd | serving 现役 |
| m13_r10_gsqrco_container_q40.gguf | 6.482G | 7.84 | 28ad41308f373cc3 | A/B 副本 |
| mmproj_m11.gguf | 921.7M | — | d89c4bc142d02ed6 | 视觉塔(冻结故兼容) |
| m13_r10_lora.gguf | 232.8M | — | f31236e81f547c95 | R13 重烘焙(git 史 9/24 新者为准, 已覆盖 9/23 旧版+sidecar) |
- 两成品为不同文件(全量 sha256 相异, 稀疏差异=精炼只覆写部分张量), A/B 对比有效。
- 完整性: HF↔容器盘内 4 点×1MB 抽块 sha256 全对 + 全量对照; lora 按用户规则"git 历史新者为准"覆盖。
四、R14 serving 与公网调用(现役速查)
- serving: llama-server(/tmp/lcfin)
-m m13_r10_gsqrco_v1.gguf --mmproj mmproj_m11.gguf --host 0.0.0.0 --port 8101 --api-key 1234 -c 131072 -fa on -ctk q8_0 -ctv q8_0 -t 2 -ngl 99(双 T4 ctx 131072; 单卡 65536→32768 梯子); 冒烟 35.9 tok/s, 域题 48,000 元正确。 - 公网调用(单隧道双功能, key=1234):
https://aitun.cc/8RF5ST7K/v1/chat/completions(OpenAI 兼容; /v1/models、/v1/chat/completions、SSE 流式实测全绿)- 原理: kaggle_server.py 注入
/v1/*→127.0.0.1:8101 代理补丁 (补丁: tchbcb/samai-8b-M8:artifacts/r14_scripts/r14_v1_proxy.py; 坑#67/#68) - 隧道本体不动: aitun-client -s aitun.cc:6639 -p 5000 → kaggle_server:5000 (/execute 信封 v=1 协议不变, 测试信封 crc=9e30c354 回显 42)
- 无 kaggle_server 的容器:
aitun -p 8101直连模式(oneclick 自动选择)
- 原理: kaggle_server.py 注入
- 一键运行:
python3 oneclick_gsqrco_serve.py --token hf_xxx(samai-9b 仓根目录) → 拉 6.48G+921M+tar → llama-server → 隧道适配 → 打印 curl/openai 示例。 - A/B 切换: 容器 /tmp/k8b/ 内两成品均在(容器版盘内名 r13_container_q40.gguf), pkill llama-server 后换 -m 路径重启(ctx 梯子照旧)。
五、怎么测评(顺延队列)
S4: serving 级新判据探针(推迟, 原因: 用户令优先 serving/量化线)
- 脚本:
s8_sprobe_r12.py(直连 llama-server :8101); A/B/C/D/E 五场景全 PASS 才算 R12 训练验收。 - 长探针:
s8_sprobe_long.py(≈25K token + 两振回放)——「长上下文泛化缺失」复测位。
S5: 实弹(推迟, 同上)
- GAIA#13 新卷: zagent_r12 (:8384) + r10 serving, 验 T5 框架修复。
- testB computer-use known-good 重放: Xvfb Popen 后台 + X99 轮询 + ffmpeg -t 25 + xdotool 42 键。
- 判分归档 →
artifacts/m15_reports/+ worklog 终章。
- 红线: 探针在 serving 端出结论; 框架硬约束是兜底不是替代; FAIL 先查 serving think 原话。
六、资产地图(HF 双仓)
tchbcb/samai-9b(公仓, 免 token 读):
| 文件 | 说明 |
|---|---|
| m13_r10_gsqrco_v1.gguf | R13 最优量化盘(ppl 7.99, serving 现役) |
| m13_r10_gsqrco_container_q40.gguf | R13 容器版 A/B 副本(ppl 7.84) |
| oneclick_gsqrco_serve.py | 一键 serving 脚本(含 /v1 补丁+双模式隧道) |
| mmproj_m11.gguf / m13_r10_lora.gguf(+.sha256) | 视觉塔 / R13 重烘焙 LoRA |
| llama_cpp_cuda_colab_t4.tar.gz(55M) / _v2(144M) | CUDA 编译资产(sm_75 全家桶, 双 tar 回退) |
| m13_r10_mixbit_4x8.gguf / m13_r10_q4_k_m.gguf | R12 旧量化盘(被 R13 取代, 留档对照) |
| train.md v4.6 / transfer.md(本文件) | 坑册 #1-#70 / R14 交接 |
| merged_m13_r8_hf/ (18G fp16) / m13_r8_lora/ / zagent_r12.tar.gz | 历史资产 |
tchbcb/samai-8b-M8(私仓, 写需 token):
m12_adapters/m13_r{6..10}/— 全血统 LoRAartifacts/r14_scripts/— r14_boot / p_r14_fire / p_r14_poll / router_8102.py / r14_v1_proxy.py / oneclick_gsqrco_serve.py(R14 弹药全套, 逐字节回读验证)artifacts/r13_scripts/、artifacts/m15_{scripts,samples,reports}/— 历轮弹药与判词- train.md v4.6、transfer.md(双仓同步)、artifacts/HANDOVER_R11.txt
七、坑册速查(R13/R14 期新增)
- #67 aitun 免 token 隧道同机单身份: 第二个客户端实例顶掉第一个(Connection replaced by newer connection of this client)。真双隧道需 -k 账户 token; 单隧道双功能走 /v1 代理补丁。
- #68 kaggle_server 与 aitun 客户端都有监督环(jupyter kernel 自动重生)→ 抢端口/改 -p 必被弹回。正解=补丁本体文件+kill 让监督环带补丁重生; 兜底自检自启。
- #69 信封内无界 grep/find 扫 /tmp(40G 模型盘)→ 30s 超时。搜索必须 maxdepth+-size 限定。
- #70 双 /health 同名: kaggle_server(富信息)vs llama({"status":"ok"})。经代理后验 llama 用 /v1/models。
- 沿用: #64 bash sleep>600s 被杀(长等待分片); 信封载荷 ≤3KB/段; token 只运行时注入+ 推仓前 grep 断言; #50 隧道连环死(六现); #78 setsid 脱组; #81 nohup 排雷; #83 Colab tar 只跑 Colab。
八、下一步队列(按序)
- 用户 A/B 实测: gsqrco_v1(7.99, 现役)↔ container_q40(7.84)公网盲测对比 (ppl 度量对齐问题以实测为准; 切换法见 §四)。
- S4 探针(§五 A-E)→ R12 训练验收终审。
- S5 实弹(GAIA#13 复测 + testB 重放)→ R14_verdict.md + worklog 终章。
- 候选: lproj1 基线(65min→目标 <10min); zagent PR 回官方仓; GSQ-RCO 再炼 (调 steps/tau 追平容器版 ppl 后替换现役盘, 需重新走完整性验证)。