Instructions to use tchbcb/samai-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tchbcb/samai-9b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tchbcb/samai-9b:Q4_K_M # Run inference directly in the terminal: llama cli -hf tchbcb/samai-9b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tchbcb/samai-9b:Q4_K_M # Run inference directly in the terminal: llama cli -hf tchbcb/samai-9b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tchbcb/samai-9b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf tchbcb/samai-9b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tchbcb/samai-9b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf tchbcb/samai-9b:Q4_K_M
Use Docker
docker model run hf.co/tchbcb/samai-9b:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use tchbcb/samai-9b with Ollama:
ollama run hf.co/tchbcb/samai-9b:Q4_K_M
- Unsloth Desktop
- Pi
How to use tchbcb/samai-9b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-9b:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tchbcb/samai-9b:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tchbcb/samai-9b with Docker Model Runner:
docker model run hf.co/tchbcb/samai-9b:Q4_K_M
- Lemonade
How to use tchbcb/samai-9b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tchbcb/samai-9b:Q4_K_M
Run and chat with the model
lemonade run user.samai-9b-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use tchbcb/samai-9b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-9b:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tchbcb/samai-9b:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tchbcb/samai-9b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-9b:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tchbcb/samai-9b:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download transfer.md from tchbcb/samai-9b: direct link, hf CLI and curl.
- Browser
- Download file 8.07 kB
-
https://huggingface.co/tchbcb/samai-9b/resolve/main/transfer.md
- Command line
-
hf download hf://tchbcb/samai-9b/transfer.md
-
curl -L -o transfer.md https://huggingface.co/tchbcb/samai-9b/resolve/main/transfer.md
8.07 kB
| # SAMAI-8B 交接文档 (transfer.md) — R14 轮 | |
| > 更新: 2026-09-25 | 交接人: Super Z 主监督者 | 读者: 下任监督者/未来的自己 | |
| > 项目一句话: 在 Qwen3.5-9B VL 底模上用 LoRA 血统链(m13_r6→r7→r8→r9→r10)训练 agent 行为纪律, | |
| > 配 zagent 框架级硬约束,闭环「采集→建窗→训练→部署→实测」。 | |
| > 本轮状态: **R13 GSQ-RCO-lite 量化收官(gsqrco_v1 ppl 7.99 现役 serving, 旧 mixbit 8.11); | |
| > R14 单隧道复用公网全绿(aitun.cc/8RF5ST7K/v1, key=1234, SSE 通); S4/S5 实测按用户令顺延。** | |
| --- | |
| ## 一、训练什么 | |
| **模型血统链**(每代 = 底模 + 上一代 adapter 续训出来的新 LoRA): | |
| ``` | |
| Qwen/Qwen3.5-9B (底模, 视觉塔冻结) | |
| └─ m13_r6 (M15 反死环 147 窗) | |
| └─ m13_r7 (R7 反死环根治 13+41 窗) | |
| └─ m13_r8 (R8 两振出局+think 纪律 235 窗) | |
| └─ m13_r9 (R11 长上下文反环 210+16 窗) | |
| └─ m13_r10 (R12 认知诚实+纯文本收束 268 窗) ← 血统终点, 量化盘母体 | |
| ``` | |
| - 血统机制: `s8_m13_train_v9.py` 加载底模后 `PeftModel.from_pretrained(model, 上一代adapter, is_trainable=True)` 续训。 | |
| 因此 **r10 adapter 权重已含 r6..r10 全链血统**,部署时只需「底模 ⊕ r10 adapter」一次 merge,数学等价。 | |
| - LoRA 配置: r=32, alpha=64, dropout=0.05, bias=none, task_type=CAUSAL_LM(vision tower + merger 冻结)。 | |
| - adapter 全集: `tchbcb/samai-8b-M8/m12_adapters/m13_r{6..10}/`。 | |
| ## 二、训练数据(268 窗, R12 期固化) | |
| R11 210 窗(拦截后换路/收束/两振 + 视频窗)+ R12 新 58 窗(H1 认知诚实 20 / H2a 纯文本收束 16 / | |
| H2b 空 tool_call 纠偏 8 / H2c 拦截后换路 8 / H3 负控 6)= `samples_r12_train.jsonl`。 | |
| 窗格式与 R8 antiloop2 对齐; 深度 1200/2400/3600/4500; 构建器 `r12_data_build.py`。 | |
| 训练要点: MAXLEN 3400 起步(6144 确定性 OOM, 坑#66)、fp16 禁 bf16、分块 fp32 CE、 | |
| NAN 自杀保护 + 时间熔断。详见 train.md v4.6 §三。 | |
| ## 三、R13 量化战报(GSQ-RCO-lite, 现行最优) | |
| 用户否决 mixbit_4x8("量化技术不好")→ 参考 GSQ(Gumbel-Softmax Quantization, arXiv 2604.18556) | |
| + RCO(arXiv 2605.00649, ISTA-DASLab)自适配 **GSQ-RCO-lite 四段链**: | |
| 链重建 → imatrix 校准(q8 路线 ppl 5.0709)→ DP 背包逐张量类型分配 → --imatrix+类型表容器量化。 | |
| | 文件(tchbcb/samai-9b) | 体积 | ppl | sha256 前 16 | 角色 | | |
| |---|---|---|---|---| | |
| | m13_r10_gsqrco_v1.gguf | 6.482G | **7.99**(旧 mixbit 8.11) | 23a73197243cb9bd | **serving 现役** | | |
| | m13_r10_gsqrco_container_q40.gguf | 6.482G | 7.84 | 28ad41308f373cc3 | A/B 副本 | | |
| | mmproj_m11.gguf | 921.7M | — | d89c4bc142d02ed6 | 视觉塔(冻结故兼容) | | |
| | m13_r10_lora.gguf | 232.8M | — | f31236e81f547c95 | R13 重烘焙(git 史 9/24 新者为准, 已覆盖 9/23 旧版+sidecar) | | |
| - 两成品为**不同文件**(全量 sha256 相异, 稀疏差异=精炼只覆写部分张量), A/B 对比有效。 | |
| - 完整性: HF↔容器盘内 4 点×1MB 抽块 sha256 全对 + 全量对照; lora 按用户规则"git 历史新者为准"覆盖。 | |
| ## 四、R14 serving 与公网调用(现役速查) | |
| - **serving**: llama-server(/tmp/lcfin) | |
| `-m m13_r10_gsqrco_v1.gguf --mmproj mmproj_m11.gguf --host 0.0.0.0 --port 8101 | |
| --api-key 1234 -c 131072 -fa on -ctk q8_0 -ctv q8_0 -t 2 -ngl 99` | |
| (双 T4 ctx 131072; 单卡 65536→32768 梯子); 冒烟 35.9 tok/s, 域题 48,000 元正确。 | |
| - **公网调用(单隧道双功能, key=1234)**: `https://aitun.cc/8RF5ST7K/v1/chat/completions` | |
| (OpenAI 兼容; /v1/models、/v1/chat/completions、SSE 流式实测全绿) | |
| - 原理: kaggle_server.py 注入 `/v1/*`→127.0.0.1:8101 代理补丁 | |
| (补丁: tchbcb/samai-8b-M8:artifacts/r14_scripts/r14_v1_proxy.py; 坑#67/#68) | |
| - 隧道本体不动: aitun-client -s aitun.cc:6639 -p 5000 → kaggle_server:5000 | |
| (/execute 信封 v=1 协议不变, 测试信封 crc=9e30c354 回显 42) | |
| - 无 kaggle_server 的容器: `aitun -p 8101` 直连模式(oneclick 自动选择) | |
| - **一键运行**: `python3 oneclick_gsqrco_serve.py --token hf_xxx`(samai-9b 仓根目录) | |
| → 拉 6.48G+921M+tar → llama-server → 隧道适配 → 打印 curl/openai 示例。 | |
| - **A/B 切换**: 容器 /tmp/k8b/ 内两成品均在(容器版盘内名 r13_container_q40.gguf), | |
| pkill llama-server 后换 -m 路径重启(ctx 梯子照旧)。 | |
| ## 五、怎么测评(顺延队列) | |
| ### S4: serving 级新判据探针(推迟, 原因: 用户令优先 serving/量化线) | |
| - 脚本: `s8_sprobe_r12.py`(直连 llama-server :8101); A/B/C/D/E 五场景全 PASS 才算 R12 训练验收。 | |
| - 长探针: `s8_sprobe_long.py`(≈25K token + 两振回放)——「长上下文泛化缺失」复测位。 | |
| ### S5: 实弹(推迟, 同上) | |
| 1. GAIA#13 新卷: zagent_r12 (:8384) + r10 serving, 验 T5 框架修复。 | |
| 2. testB computer-use known-good 重放: Xvfb Popen 后台 + X99 轮询 + ffmpeg -t 25 + xdotool 42 键。 | |
| 3. 判分归档 → `artifacts/m15_reports/` + worklog 终章。 | |
| - 红线: 探针在 serving 端出结论; 框架硬约束是兜底不是替代; FAIL 先查 serving think 原话。 | |
| ## 六、资产地图(HF 双仓) | |
| **tchbcb/samai-9b(公仓, 免 token 读)**: | |
| | 文件 | 说明 | | |
| |---|---| | |
| | m13_r10_gsqrco_v1.gguf | **R13 最优量化盘(ppl 7.99, serving 现役)** | | |
| | m13_r10_gsqrco_container_q40.gguf | R13 容器版 A/B 副本(ppl 7.84) | | |
| | oneclick_gsqrco_serve.py | **一键 serving 脚本(含 /v1 补丁+双模式隧道)** | | |
| | mmproj_m11.gguf / m13_r10_lora.gguf(+.sha256) | 视觉塔 / R13 重烘焙 LoRA | | |
| | llama_cpp_cuda_colab_t4.tar.gz(55M) / _v2(144M) | CUDA 编译资产(sm_75 全家桶, 双 tar 回退) | | |
| | m13_r10_mixbit_4x8.gguf / m13_r10_q4_k_m.gguf | R12 旧量化盘(被 R13 取代, 留档对照) | | |
| | train.md v4.6 / transfer.md(本文件) | 坑册 #1-#70 / R14 交接 | | |
| | merged_m13_r8_hf/ (18G fp16) / m13_r8_lora/ / zagent_r12.tar.gz | 历史资产 | | |
| **tchbcb/samai-8b-M8(私仓, 写需 token)**: | |
| - `m12_adapters/m13_r{6..10}/` — 全血统 LoRA | |
| - `artifacts/r14_scripts/` — r14_boot / p_r14_fire / p_r14_poll / **router_8102.py** / | |
| **r14_v1_proxy.py** / oneclick_gsqrco_serve.py(R14 弹药全套, 逐字节回读验证) | |
| - `artifacts/r13_scripts/`、`artifacts/m15_{scripts,samples,reports}/` — 历轮弹药与判词 | |
| - train.md v4.6、transfer.md(双仓同步)、artifacts/HANDOVER_R11.txt | |
| ## 七、坑册速查(R13/R14 期新增) | |
| - #67 aitun 免 token 隧道**同机单身份**: 第二个客户端实例顶掉第一个(Connection replaced by | |
| newer connection of this client)。真双隧道需 -k 账户 token; 单隧道双功能走 /v1 代理补丁。 | |
| - #68 kaggle_server 与 aitun 客户端都有**监督环**(jupyter kernel 自动重生)→ 抢端口/改 -p | |
| 必被弹回。正解=补丁本体文件+kill 让监督环带补丁重生; 兜底自检自启。 | |
| - #69 信封内无界 grep/find 扫 /tmp(40G 模型盘)→ 30s 超时。搜索必须 maxdepth+-size 限定。 | |
| - #70 双 /health 同名: kaggle_server(富信息)vs llama({"status":"ok"})。经代理后验 llama | |
| 用 /v1/models。 | |
| - 沿用: #64 bash sleep>600s 被杀(长等待分片); 信封载荷 ≤3KB/段; token 只运行时注入+ | |
| 推仓前 grep 断言; #50 隧道连环死(六现); #78 setsid 脱组; #81 nohup 排雷; #83 Colab tar | |
| 只跑 Colab。 | |
| ## 八、下一步队列(按序) | |
| 1. **用户 A/B 实测**: gsqrco_v1(7.99, 现役)↔ container_q40(7.84)公网盲测对比 | |
| (ppl 度量对齐问题以实测为准; 切换法见 §四)。 | |
| 2. S4 探针(§五 A-E)→ R12 训练验收终审。 | |
| 3. S5 实弹(GAIA#13 复测 + testB 重放)→ R14_verdict.md + worklog 终章。 | |
| 4. 候选: lproj1 基线(65min→目标 <10min); zagent PR 回官方仓; GSQ-RCO 再炼 | |
| (调 steps/tau 追平容器版 ppl 后替换现役盘, 需重新走完整性验证)。 | |