File size: 3,902 Bytes
8b8e98d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
# samai-2b OpenAI 兼容 API — 使用说明

`openai_api.py`: 为 samai-pnet 系列模型 (PonderNet DMoE 2B) 提供标准 OpenAI 兼容 HTTP 接口,
生成口径与验收评估完全一致 (chat template + 贪心默认 + 双协议), 方便联调测试。

## 启动 (DSW)

```bash
# 自动选最新 staging (r16 > r15 > r14 ...), 端口 8000
bash /mnt/workspace/run_serve.sh            # 手动启动
bash /mnt/workspace/run_serve.sh 8080       # 指定端口

# 自动起服: serve_watchdog 已挂起, 等 r15/r16 链全部落定 + GPU 释放后自动启动
tail -f /mnt/workspace/serve_watchdog.log   # 看状态
tail -f /mnt/workspace/openai_api.log       # 看服务日志
```

无 GPU 联调 (本地或 DSW 均可):

```bash
python3 openai_api.py --mock --port 8000    # 固定应答, 验证 HTTP 层
python3 test_openai_api.py                  # 22 项 schema 自动验收 (mock)
```

## 端点

| 端点 | 说明 |
|---|---|
| POST /v1/chat/completions | 对话补全 (流式 SSE + 非流式) |
| POST /v1/completions | 原始文本续写 (不过 chat template) |
| GET /v1/models | 模型列表 (samai-2b / samai-2b-think-off) |
| GET /health | 健康检查 (loaded / gpu 显存) |

## 双协议 (bare / think_off) 三种指定方式

优先级: 模型名后缀 > chat_template_kwargs > 顶层字段

```bash
# ① 模型名后缀
curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "samai-2b-think-off",
  "messages": [{"role":"user","content":"你叫什么名字?"}], "max_tokens": 96}'

# ② chat_template_kwargs (vLLM 风格)
{"model":"samai-2b","chat_template_kwargs":{"enable_thinking":false},"messages":[...]}

# ③ 顶层 enable_thinking
{"model":"samai-2b","enable_thinking":false,"messages":[...]}
```

## 应答结构 (think 处理)

```json
{
  "choices": [{
    "message": {
      "role": "assistant",
      "content": "我是 samai-2b 模型,由 SamAI 研发…",   // 答案 (think 已剥离)
      "reasoning_content": "…思考正文…",                 // 仅 bare 协议且有思考时
      "raw": "<think>\n…\n</think>\n\n我是 samai-2b…"    // 原始全文, 与验收脚本同口径
    },
    "finish_reason": "stop"
  }],
  "usage": {"prompt_tokens":…, "completion_tokens":…, "total_tokens":…},
  "ponder": {"decode_forwards": 24, "steps_mean": 2.08}   // ponder 诊断 (extra)
}
```

- `raw_content: true` → content 直接放原始全文 (含 think), 方便与验收脚本对拍
- 流式: `reasoning_content` 先流出, `</think>` 后 `content` 流出; chunk id 全程一致;
  `stream_options.include_usage` 支持 (末帧 usage, [DONE] 收尾)
- 流式说明: 服务端为整段生成后按块推流 (schema 与 OpenAI 一致; 首帧延迟≈生成耗时)

## 参数映射

| OpenAI 参数 | 行为 |
|---|---|
| max_tokens / max_completion_tokens | max_new_tokens (默认 256, 自动按 ctx 裁剪) |
| temperature 缺省/0 | **贪心 do_sample=False (与验收口径一致)** |
| temperature > 0 | do_sample=True + top_p (默认 1.0) |
| top_k, repetition_penalty, seed | 直传 (extra) |
| stop (str 或 list) | decode 文本命中即停, finish_reason=stop |
| api_key | 服务端 --api-key 开启校验, Bearer 方式 |

## openai SDK 调用示例

```python
from openai import OpenAI
client = OpenAI(base_url="http://<DSW-IP>:8000/v1", api_key="EMPTY")
r = client.chat.completions.create(
    model="samai-2b",
    messages=[{"role": "user", "content": "12×8等于多少?"}],
    extra_body={"enable_thinking": False})
print(r.choices[0].message.content)
```

## 代码仓同步

- ModelScope 中转: luoyunlan168/samai-r12-relay (openai_api.py / run_serve.sh /
  serve_watchdog.sh / test_openai_api.py / README_OPENAI_API.md)
- HF 开发仓: tchbcb/samaidev-pnet-2b → serving/ 目录 (与 ckpt_2b_r12/r13 同仓)
- DSW: /mnt/workspace/ (openai_api.py + 两脚本), watchdog 自动起服