samaidev-pnet-2b / serving /README_OPENAI_API.md
agent
serving: OpenAI-compatible API (server + launcher + watchdog + test + docs)
8b8e98d
|
Raw History Blame Contribute Delete
3.9 kB

samai-2b OpenAI 兼容 API — 使用说明

openai_api.py: 为 samai-pnet 系列模型 (PonderNet DMoE 2B) 提供标准 OpenAI 兼容 HTTP 接口, 生成口径与验收评估完全一致 (chat template + 贪心默认 + 双协议), 方便联调测试。

启动 (DSW)

# 自动选最新 staging (r16 > r15 > r14 ...), 端口 8000
bash /mnt/workspace/run_serve.sh            # 手动启动
bash /mnt/workspace/run_serve.sh 8080       # 指定端口

# 自动起服: serve_watchdog 已挂起, 等 r15/r16 链全部落定 + GPU 释放后自动启动
tail -f /mnt/workspace/serve_watchdog.log   # 看状态
tail -f /mnt/workspace/openai_api.log       # 看服务日志

无 GPU 联调 (本地或 DSW 均可):

python3 openai_api.py --mock --port 8000    # 固定应答, 验证 HTTP 层
python3 test_openai_api.py                  # 22 项 schema 自动验收 (mock)

端点

端点 说明
POST /v1/chat/completions 对话补全 (流式 SSE + 非流式)
POST /v1/completions 原始文本续写 (不过 chat template)
GET /v1/models 模型列表 (samai-2b / samai-2b-think-off)
GET /health 健康检查 (loaded / gpu 显存)

双协议 (bare / think_off) 三种指定方式

优先级: 模型名后缀 > chat_template_kwargs > 顶层字段

# ① 模型名后缀
curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "samai-2b-think-off",
  "messages": [{"role":"user","content":"你叫什么名字?"}], "max_tokens": 96}'

# ② chat_template_kwargs (vLLM 风格)
{"model":"samai-2b","chat_template_kwargs":{"enable_thinking":false},"messages":[...]}

# ③ 顶层 enable_thinking
{"model":"samai-2b","enable_thinking":false,"messages":[...]}

应答结构 (think 处理)

{
  "choices": [{
    "message": {
      "role": "assistant",
      "content": "我是 samai-2b 模型,由 SamAI 研发…",   // 答案 (think 已剥离)
      "reasoning_content": "…思考正文…",                 // 仅 bare 协议且有思考时
      "raw": "<think>\n…\n</think>\n\n我是 samai-2b…"    // 原始全文, 与验收脚本同口径
    },
    "finish_reason": "stop"
  }],
  "usage": {"prompt_tokens":…, "completion_tokens":…, "total_tokens":…},
  "ponder": {"decode_forwards": 24, "steps_mean": 2.08}   // ponder 诊断 (extra)
}
  • raw_content: true → content 直接放原始全文 (含 think), 方便与验收脚本对拍
  • 流式: reasoning_content 先流出, </think> 后 content 流出; chunk id 全程一致; stream_options.include_usage 支持 (末帧 usage, [DONE] 收尾)
  • 流式说明: 服务端为整段生成后按块推流 (schema 与 OpenAI 一致; 首帧延迟≈生成耗时)

参数映射

OpenAI 参数 行为
max_tokens / max_completion_tokens max_new_tokens (默认 256, 自动按 ctx 裁剪)
temperature 缺省/0 贪心 do_sample=False (与验收口径一致)
temperature > 0 do_sample=True + top_p (默认 1.0)
top_k, repetition_penalty, seed 直传 (extra)
stop (str 或 list) decode 文本命中即停, finish_reason=stop
api_key 服务端 --api-key 开启校验, Bearer 方式

openai SDK 调用示例

from openai import OpenAI
client = OpenAI(base_url="http://<DSW-IP>:8000/v1", api_key="EMPTY")
r = client.chat.completions.create(
    model="samai-2b",
    messages=[{"role": "user", "content": "12×8等于多少?"}],
    extra_body={"enable_thinking": False})
print(r.choices[0].message.content)

代码仓同步

  • ModelScope 中转: luoyunlan168/samai-r12-relay (openai_api.py / run_serve.sh / serve_watchdog.sh / test_openai_api.py / README_OPENAI_API.md)
  • HF 开发仓: tchbcb/samaidev-pnet-2b → serving/ 目录 (与 ckpt_2b_r12/r13 同仓)
  • DSW: /mnt/workspace/ (openai_api.py + 两脚本), watchdog 自动起服