Sakura-Needle3-ONNX
Sakura-Needle3-ONNX provides a production-grade, highly optimized ONNX CPU conversion of Cactus-Compute/needle3 with true INT8 weight storage and execution.
- Single-File ONNX: Packaged as a clean, portable single file (
models/model_int8_compact.onnx), requiring no external tensor data files. - Ultra-Compact Footprint: 146.1 MB total disk size (69.8% smaller than FP32 at 483.8 MB, and 40% smaller than FP16 at 243.1 MB).
- Pure ONNX Runtime CPU: Runs out-of-the-box on standard
onnxruntimewith zero custom C++ operators or proprietary dependencies. - Blazing Fast CPU Performance: Delivers >2,100 tok/s prefill and 80–90 tok/s decode with ~60–75 ms TTFT on modern CPUs.
- Audited Parity: Fully validated across all 208 test cases with verified absolute counts.
Benchmark & Quality Parity (208 Cases Audit)
Audit Statement: INT8 Compact retains nearly identical aggregate benchmark accuracy (126 / 208 vs. 127 / 208 exact ground-truth matches), while raw model outputs agree with FP32 on 185 / 208 cases (88.94% pairwise exact match, 96.63% tool-selection agreement).
| Metric | ONNX FP32 | ONNX FP16 | ONNX INT8 Compact |
|---|---|---|---|
| File Size | 483.8 MB | 243.1 MB | 146.1 MB (-69.8%) |
| Ground-Truth Exact Match | 127 / 208 (61.1%) | 127 / 208 (61.1%) | 126 / 208 (60.6%) |
| Ground-Truth Folded Match | 128 / 208 (61.5%) | 128 / 208 (61.5%) | 127 / 208 (61.1%) |
| Tool Selection Accuracy | 132 / 208 (63.5%) | 132 / 208 (63.5%) | 136 / 208 (65.4%) |
| PyTorch FP32 Output Agreement | 208 / 208 (100.0%) | 208 / 208 (100.0%) | 185 / 208 (88.94%) |
| Tool Selection Agreement vs FP32 | 208 / 208 (100.0%) | 208 / 208 (100.0%) | 201 / 208 (96.63%) |
| No-Tool Accuracy | 9 / 81 (11.1%) | 9 / 81 (11.1%) | 12 / 81 (14.8%) |
Across the 208 test cases, exactly 23 outputs differ between FP32 and INT8 Compact:
- Regressions (FP32 correct / INT8 incorrect): 4 cases (1x schema keyword
requiredinserted into arguments, 1x time formatting12:00vsnoon, 2x minor string truncation). - INT8 Improvements (FP32 incorrect / INT8 correct): 3 cases (INT8 correctly rejected out-of-domain queries that FP32 hallucinated on).
- Edge cases (Both incorrect): 16 cases (challenging negations and omitted arguments).
The complete case-by-case audit log is published in benchmark/INT8_FP32_DIFFERENCES.csv.
Inference Speed & Efficiency
Benchmarks measured on AMD Ryzen AI MAX+ 392 (Windows 11, intra_op_num_threads = 4, with background system workload):
| Metric | Result |
|---|---|
| Prompt Prefill Throughput | 2,150 – 2,310 tok/s |
| Autoregressive Decode | 80.0 – 90.7 tok/s |
| Time to First Token (TTFT) | 59.4 – 75.0 ms |
| Peak Memory Consumption (RAM) | ~1.25 GB |
| Model Load Latency | ~1.68 s |
Quickstart
1. Installation
Install the minimal dependencies (no PyTorch or HuggingFace Transformers required for inference):
pip install onnxruntime numpy
2. Python Inference Example
Clone or download the repository, then run the included runtime engine:
from runtime.needle_onnx import NeedleONNX
# Initialize engine with 4 CPU threads
engine = NeedleONNX(model_path="models/model_int8_compact.onnx", num_threads=4)
# Define tools
tools = [
{
"name": "control_lights",
"description": "Adjust lights in a specified room",
"parameters": {
"type": "object",
"properties": {
"room": {"type": "string", "description": "Room name"},
"action": {"type": "string", "enum": ["on", "off", "dim"]},
"brightness": {"type": "integer", "description": "Brightness percentage 0-100"},
},
"required": ["room", "action"],
},
}
]
# Run query
query = "Dim the living room lights to 35%."
result = engine.complete(query, tools=tools, compatibility="cactus")
print("Reasoning:", result["reasoning"])
print("Function Calls:", result["function_calls"])
Output:
Reasoning: 'dim' -> action 'dim', 'living room' -> room 'living room', '35%' -> brightness 35.
Function Calls: [
{
"name": "control_lights",
"arguments": {
"room": "living room",
"action": "dim",
"brightness": 35
}
}
]
Or execute the included test script:
python runtime/example.py
Repository Structure
Sakura-Needle3-ONNX/
├── README.md # Main documentation & model card
├── LICENSE # Apache 2.0 license
├── NOTICE.md # Upstream and derivative attribution
├── config.json # Model architecture & runtime parameters
├── requirements.txt # Minimal dependencies (onnxruntime, numpy)
├── SHA256SUMS # Cryptographic checksums
│
├── models/
│ └── model_int8_compact.onnx # Standalone INT8 ONNX model (146.1 MB)
│
├── runtime/
│ ├── __init__.py # Package init
│ ├── needle_onnx.py # High-performance ONNX Runtime inference engine
│ ├── tokenizer.py # Standalone SentencePiece-compatible tokenizer
│ ├── tokenizer.json # Embedded vocabulary
│ └── example.py # Runnable verification example
│
├── docs/
│ ├── QUALITY.md # Complete 208-case quality audit report
│ └── BENCHMARK.md # Detailed CPU performance and memory metrics
│
└── benchmark/
└── INT8_FP32_DIFFERENCES.csv # Exact case-by-case difference breakdown
Architecture Details
- Parameters: ~120M parameters
- Layers: 20 Transformer layers
- Attention: Hybrid sliding window (32) and Manifold-Head Attention (mHC, 4 lanes)
- Engram Lookups: Multi-head rolling n-gram embedding tables (orders 2 & 3, dilation 3)
- Quantization Technique: True INT8 symmetric per-channel weights with optimized DequantizeLinear Gather embeddings.
Acknowledgments & License
- Base Model: Cactus-Compute/needle3 by Cactus Compute (mrfakename, gigglr).
- License: Apache License 2.0.
- Port: Developed by the Sakura AI optimization project (
webmp3).
中文说明 · 樱花 (Simplified Chinese)
English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。
Sakura-Needle3-ONNX
Sakura-Needle3-ONNX 提供了 Cactus-Compute/needle3 的生产级、高度优化的 ONNX CPU 转换版本,采用真正的 INT8 权重存储和执行。
- 单文件 ONNX:打包为一个干净、可移植的单文件(
models/model_int8_compact.onnx),无需外部张量数据文件。 - 极致紧凑的体积:总磁盘大小 146.1 MB(比 483.8 MB 的 FP32 小 69.8%,比 243.1 MB 的 FP16 小 40%)。
- 纯 ONNX Runtime CPU:在标准
onnxruntime上开箱即用,无需自定义 C++ 算子或专有依赖。 - 极快的 CPU 性能:在现代 CPU 上可达 >2,100 tok/s 的预填充 和 80–90 tok/s 的解码,首 token 延迟约 60–75 ms。
- 经过审计的一致性:已在全部 208 个测试用例上完成验证,并给出了核实过的绝对计数。
基准与质量一致性(208 个用例的审计)
审计声明:INT8 Compact 保持了几乎相同的整体基准准确率(与标准答案完全匹配的数量为 126 / 208 对 127 / 208),同时其原始模型输出在 185 / 208 个用例上与 FP32 一致(两两完全匹配率 88.94%,工具选择一致率 96.63%)。
| 指标 | ONNX FP32 | ONNX FP16 | ONNX INT8 Compact |
|---|---|---|---|
| 文件大小 | 483.8 MB | 243.1 MB | 146.1 MB (-69.8%) |
| 与标准答案完全匹配 | 127 / 208 (61.1%) | 127 / 208 (61.1%) | 126 / 208 (60.6%) |
| 与标准答案折叠匹配 | 128 / 208 (61.5%) | 128 / 208 (61.5%) | 127 / 208 (61.1%) |
| 工具选择准确率 | 132 / 208 (63.5%) | 132 / 208 (63.5%) | 136 / 208 (65.4%) |
| 与 PyTorch FP32 输出的一致性 | 208 / 208 (100.0%) | 208 / 208 (100.0%) | 185 / 208 (88.94%) |
| 与 FP32 的工具选择一致性 | 208 / 208 (100.0%) | 208 / 208 (100.0%) | 201 / 208 (96.63%) |
| 无工具准确率 | 9 / 81 (11.1%) | 9 / 81 (11.1%) | 12 / 81 (14.8%) |
在 208 个测试用例中,FP32 与 INT8 Compact 的输出恰好有 23 个不同:
- 退步(FP32 正确 / INT8 错误):4 个用例(1 个是参数中被插入了 schema 关键字
required,1 个是时间格式12:00对noon,2 个是轻微的字符串截断)。 - INT8 的改进(FP32 错误 / INT8 正确):3 个用例(INT8 正确地拒绝了 FP32 产生幻觉的域外查询)。
- 边界情况(两者都错误):16 个用例(具有挑战性的否定句和被省略的参数)。
完整的逐用例审计日志发布在 benchmark/INT8_FP32_DIFFERENCES.csv。
推理速度与效率
基准在 AMD Ryzen AI MAX+ 392 上测得(Windows 11,intra_op_num_threads = 4,有后台系统负载):
| 指标 | 结果 |
|---|---|
| 提示预填充吞吐量 | 2,150 – 2,310 tok/s |
| 自回归解码 | 80.0 – 90.7 tok/s |
| 首 token 时间 (TTFT) | 59.4 – 75.0 ms |
| 峰值内存占用 (RAM) | ~1.25 GB |
| 模型加载延迟 | ~1.68 s |
快速开始
1. 安装
安装最少的依赖(推理不需要 PyTorch 或 HuggingFace Transformers):
pip install onnxruntime numpy
2. Python 推理示例
克隆或下载本仓库,然后运行随附的运行时引擎:
from runtime.needle_onnx import NeedleONNX
# Initialize engine with 4 CPU threads
engine = NeedleONNX(model_path="models/model_int8_compact.onnx", num_threads=4)
# Define tools
tools = [
{
"name": "control_lights",
"description": "Adjust lights in a specified room",
"parameters": {
"type": "object",
"properties": {
"room": {"type": "string", "description": "Room name"},
"action": {"type": "string", "enum": ["on", "off", "dim"]},
"brightness": {"type": "integer", "description": "Brightness percentage 0-100"},
},
"required": ["room", "action"],
},
}
]
# Run query
query = "Dim the living room lights to 35%."
result = engine.complete(query, tools=tools, compatibility="cactus")
print("Reasoning:", result["reasoning"])
print("Function Calls:", result["function_calls"])
输出:
Reasoning: 'dim' -> action 'dim', 'living room' -> room 'living room', '35%' -> brightness 35.
Function Calls: [
{
"name": "control_lights",
"arguments": {
"room": "living room",
"action": "dim",
"brightness": 35
}
}
]
或者执行随附的测试脚本:
python runtime/example.py
仓库结构
Sakura-Needle3-ONNX/
├── README.md # Main documentation & model card
├── LICENSE # Apache 2.0 license
├── NOTICE.md # Upstream and derivative attribution
├── config.json # Model architecture & runtime parameters
├── requirements.txt # Minimal dependencies (onnxruntime, numpy)
├── SHA256SUMS # Cryptographic checksums
│
├── models/
│ └── model_int8_compact.onnx # Standalone INT8 ONNX model (146.1 MB)
│
├── runtime/
│ ├── __init__.py # Package init
│ ├── needle_onnx.py # High-performance ONNX Runtime inference engine
│ ├── tokenizer.py # Standalone SentencePiece-compatible tokenizer
│ ├── tokenizer.json # Embedded vocabulary
│ └── example.py # Runnable verification example
│
├── docs/
│ ├── QUALITY.md # Complete 208-case quality audit report
│ └── BENCHMARK.md # Detailed CPU performance and memory metrics
│
└── benchmark/
└── INT8_FP32_DIFFERENCES.csv # Exact case-by-case difference breakdown
架构详情
- 参数量:约 120M 参数
- 层数:20 个 Transformer 层
- 注意力:混合滑动窗口(32)与流形头注意力(mHC,4 条通道)
- Engram 查找:多头滚动 n-gram 嵌入表(阶数 2 和 3,膨胀 3)
- 量化技术:真正的 INT8 对称逐通道权重,配合优化的 DequantizeLinear Gather 嵌入。
致谢与许可证
- 基础模型:Cactus Compute(mrfakename、gigglr)的 Cactus-Compute/needle3。
- 许可证:Apache License 2.0。
- 移植:由 Sakura AI 优化项目(
webmp3)开发。
- Downloads last month
- 138
Model tree for webmp3/Sakura-Needle3-ONNX
Base model
Cactus-Compute/needle3