Sakura-Needle3-ONNX

Sakura-Needle3-ONNX provides a production-grade, highly optimized ONNX CPU conversion of Cactus-Compute/needle3 with true INT8 weight storage and execution.

  • Single-File ONNX: Packaged as a clean, portable single file (models/model_int8_compact.onnx), requiring no external tensor data files.
  • Ultra-Compact Footprint: 146.1 MB total disk size (69.8% smaller than FP32 at 483.8 MB, and 40% smaller than FP16 at 243.1 MB).
  • Pure ONNX Runtime CPU: Runs out-of-the-box on standard onnxruntime with zero custom C++ operators or proprietary dependencies.
  • Blazing Fast CPU Performance: Delivers >2,100 tok/s prefill and 80–90 tok/s decode with ~60–75 ms TTFT on modern CPUs.
  • Audited Parity: Fully validated across all 208 test cases with verified absolute counts.

Benchmark & Quality Parity (208 Cases Audit)

Audit Statement: INT8 Compact retains nearly identical aggregate benchmark accuracy (126 / 208 vs. 127 / 208 exact ground-truth matches), while raw model outputs agree with FP32 on 185 / 208 cases (88.94% pairwise exact match, 96.63% tool-selection agreement).

Metric ONNX FP32 ONNX FP16 ONNX INT8 Compact
File Size 483.8 MB 243.1 MB 146.1 MB (-69.8%)
Ground-Truth Exact Match 127 / 208 (61.1%) 127 / 208 (61.1%) 126 / 208 (60.6%)
Ground-Truth Folded Match 128 / 208 (61.5%) 128 / 208 (61.5%) 127 / 208 (61.1%)
Tool Selection Accuracy 132 / 208 (63.5%) 132 / 208 (63.5%) 136 / 208 (65.4%)
PyTorch FP32 Output Agreement 208 / 208 (100.0%) 208 / 208 (100.0%) 185 / 208 (88.94%)
Tool Selection Agreement vs FP32 208 / 208 (100.0%) 208 / 208 (100.0%) 201 / 208 (96.63%)
No-Tool Accuracy 9 / 81 (11.1%) 9 / 81 (11.1%) 12 / 81 (14.8%)

Across the 208 test cases, exactly 23 outputs differ between FP32 and INT8 Compact:

  • Regressions (FP32 correct / INT8 incorrect): 4 cases (1x schema keyword required inserted into arguments, 1x time formatting 12:00 vs noon, 2x minor string truncation).
  • INT8 Improvements (FP32 incorrect / INT8 correct): 3 cases (INT8 correctly rejected out-of-domain queries that FP32 hallucinated on).
  • Edge cases (Both incorrect): 16 cases (challenging negations and omitted arguments).

The complete case-by-case audit log is published in benchmark/INT8_FP32_DIFFERENCES.csv.


Inference Speed & Efficiency

Benchmarks measured on AMD Ryzen AI MAX+ 392 (Windows 11, intra_op_num_threads = 4, with background system workload):

Metric Result
Prompt Prefill Throughput 2,150 – 2,310 tok/s
Autoregressive Decode 80.0 – 90.7 tok/s
Time to First Token (TTFT) 59.4 – 75.0 ms
Peak Memory Consumption (RAM) ~1.25 GB
Model Load Latency ~1.68 s

Quickstart

1. Installation

Install the minimal dependencies (no PyTorch or HuggingFace Transformers required for inference):

pip install onnxruntime numpy

2. Python Inference Example

Clone or download the repository, then run the included runtime engine:

from runtime.needle_onnx import NeedleONNX

# Initialize engine with 4 CPU threads
engine = NeedleONNX(model_path="models/model_int8_compact.onnx", num_threads=4)

# Define tools
tools = [
    {
        "name": "control_lights",
        "description": "Adjust lights in a specified room",
        "parameters": {
            "type": "object",
            "properties": {
                "room": {"type": "string", "description": "Room name"},
                "action": {"type": "string", "enum": ["on", "off", "dim"]},
                "brightness": {"type": "integer", "description": "Brightness percentage 0-100"},
            },
            "required": ["room", "action"],
        },
    }
]

# Run query
query = "Dim the living room lights to 35%."
result = engine.complete(query, tools=tools, compatibility="cactus")

print("Reasoning:", result["reasoning"])
print("Function Calls:", result["function_calls"])

Output:

Reasoning: 'dim' -> action 'dim', 'living room' -> room 'living room', '35%' -> brightness 35.
Function Calls: [
  {
    "name": "control_lights",
    "arguments": {
      "room": "living room",
      "action": "dim",
      "brightness": 35
    }
  }
]

Or execute the included test script:

python runtime/example.py

Repository Structure

Sakura-Needle3-ONNX/
├── README.md                           # Main documentation & model card
├── LICENSE                             # Apache 2.0 license
├── NOTICE.md                           # Upstream and derivative attribution
├── config.json                         # Model architecture & runtime parameters
├── requirements.txt                    # Minimal dependencies (onnxruntime, numpy)
├── SHA256SUMS                          # Cryptographic checksums
│
├── models/
│   └── model_int8_compact.onnx         # Standalone INT8 ONNX model (146.1 MB)
│
├── runtime/
│   ├── __init__.py                     # Package init
│   ├── needle_onnx.py                  # High-performance ONNX Runtime inference engine
│   ├── tokenizer.py                    # Standalone SentencePiece-compatible tokenizer
│   ├── tokenizer.json                  # Embedded vocabulary
│   └── example.py                      # Runnable verification example
│
├── docs/
│   ├── QUALITY.md                      # Complete 208-case quality audit report
│   └── BENCHMARK.md                    # Detailed CPU performance and memory metrics
│
└── benchmark/
    └── INT8_FP32_DIFFERENCES.csv       # Exact case-by-case difference breakdown

Architecture Details

  • Parameters: ~120M parameters
  • Layers: 20 Transformer layers
  • Attention: Hybrid sliding window (32) and Manifold-Head Attention (mHC, 4 lanes)
  • Engram Lookups: Multi-head rolling n-gram embedding tables (orders 2 & 3, dilation 3)
  • Quantization Technique: True INT8 symmetric per-channel weights with optimized DequantizeLinear Gather embeddings.

Acknowledgments & License

  • Base Model: Cactus-Compute/needle3 by Cactus Compute (mrfakename, gigglr).
  • License: Apache License 2.0.
  • Port: Developed by the Sakura AI optimization project (webmp3).

中文说明 · 樱花 (Simplified Chinese)

English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。

Sakura-Needle3-ONNX

Sakura-Needle3-ONNX 提供了 Cactus-Compute/needle3 的生产级、高度优化的 ONNX CPU 转换版本,采用真正的 INT8 权重存储和执行。

  • 单文件 ONNX:打包为一个干净、可移植的单文件(models/model_int8_compact.onnx),无需外部张量数据文件。
  • 极致紧凑的体积:总磁盘大小 146.1 MB(比 483.8 MB 的 FP32 小 69.8%,比 243.1 MB 的 FP16 小 40%)。
  • 纯 ONNX Runtime CPU:在标准 onnxruntime 上开箱即用,无需自定义 C++ 算子或专有依赖。
  • 极快的 CPU 性能:在现代 CPU 上可达 >2,100 tok/s 的预填充 和 80–90 tok/s 的解码,首 token 延迟约 60–75 ms。
  • 经过审计的一致性:已在全部 208 个测试用例上完成验证,并给出了核实过的绝对计数。

基准与质量一致性(208 个用例的审计)

审计声明:INT8 Compact 保持了几乎相同的整体基准准确率(与标准答案完全匹配的数量为 126 / 208 对 127 / 208),同时其原始模型输出在 185 / 208 个用例上与 FP32 一致(两两完全匹配率 88.94%,工具选择一致率 96.63%)。

指标 ONNX FP32 ONNX FP16 ONNX INT8 Compact
文件大小 483.8 MB 243.1 MB 146.1 MB (-69.8%)
与标准答案完全匹配 127 / 208 (61.1%) 127 / 208 (61.1%) 126 / 208 (60.6%)
与标准答案折叠匹配 128 / 208 (61.5%) 128 / 208 (61.5%) 127 / 208 (61.1%)
工具选择准确率 132 / 208 (63.5%) 132 / 208 (63.5%) 136 / 208 (65.4%)
与 PyTorch FP32 输出的一致性 208 / 208 (100.0%) 208 / 208 (100.0%) 185 / 208 (88.94%)
与 FP32 的工具选择一致性 208 / 208 (100.0%) 208 / 208 (100.0%) 201 / 208 (96.63%)
无工具准确率 9 / 81 (11.1%) 9 / 81 (11.1%) 12 / 81 (14.8%)

在 208 个测试用例中,FP32 与 INT8 Compact 的输出恰好有 23 个不同:

  • 退步(FP32 正确 / INT8 错误):4 个用例(1 个是参数中被插入了 schema 关键字 required,1 个是时间格式 12:00 对 noon,2 个是轻微的字符串截断)。
  • INT8 的改进(FP32 错误 / INT8 正确):3 个用例(INT8 正确地拒绝了 FP32 产生幻觉的域外查询)。
  • 边界情况(两者都错误):16 个用例(具有挑战性的否定句和被省略的参数)。

完整的逐用例审计日志发布在 benchmark/INT8_FP32_DIFFERENCES.csv。


推理速度与效率

基准在 AMD Ryzen AI MAX+ 392 上测得(Windows 11,intra_op_num_threads = 4,有后台系统负载):

指标 结果
提示预填充吞吐量 2,150 – 2,310 tok/s
自回归解码 80.0 – 90.7 tok/s
首 token 时间 (TTFT) 59.4 – 75.0 ms
峰值内存占用 (RAM) ~1.25 GB
模型加载延迟 ~1.68 s

快速开始

1. 安装

安装最少的依赖(推理不需要 PyTorch 或 HuggingFace Transformers):

pip install onnxruntime numpy

2. Python 推理示例

克隆或下载本仓库,然后运行随附的运行时引擎:

from runtime.needle_onnx import NeedleONNX

# Initialize engine with 4 CPU threads
engine = NeedleONNX(model_path="models/model_int8_compact.onnx", num_threads=4)

# Define tools
tools = [
    {
        "name": "control_lights",
        "description": "Adjust lights in a specified room",
        "parameters": {
            "type": "object",
            "properties": {
                "room": {"type": "string", "description": "Room name"},
                "action": {"type": "string", "enum": ["on", "off", "dim"]},
                "brightness": {"type": "integer", "description": "Brightness percentage 0-100"},
            },
            "required": ["room", "action"],
        },
    }
]

# Run query
query = "Dim the living room lights to 35%."
result = engine.complete(query, tools=tools, compatibility="cactus")

print("Reasoning:", result["reasoning"])
print("Function Calls:", result["function_calls"])

输出:

Reasoning: 'dim' -> action 'dim', 'living room' -> room 'living room', '35%' -> brightness 35.
Function Calls: [
  {
    "name": "control_lights",
    "arguments": {
      "room": "living room",
      "action": "dim",
      "brightness": 35
    }
  }
]

或者执行随附的测试脚本:

python runtime/example.py

仓库结构

Sakura-Needle3-ONNX/
├── README.md                           # Main documentation & model card
├── LICENSE                             # Apache 2.0 license
├── NOTICE.md                           # Upstream and derivative attribution
├── config.json                         # Model architecture & runtime parameters
├── requirements.txt                    # Minimal dependencies (onnxruntime, numpy)
├── SHA256SUMS                          # Cryptographic checksums
│
├── models/
│   └── model_int8_compact.onnx         # Standalone INT8 ONNX model (146.1 MB)
│
├── runtime/
│   ├── __init__.py                     # Package init
│   ├── needle_onnx.py                  # High-performance ONNX Runtime inference engine
│   ├── tokenizer.py                    # Standalone SentencePiece-compatible tokenizer
│   ├── tokenizer.json                  # Embedded vocabulary
│   └── example.py                      # Runnable verification example
│
├── docs/
│   ├── QUALITY.md                      # Complete 208-case quality audit report
│   └── BENCHMARK.md                    # Detailed CPU performance and memory metrics
│
└── benchmark/
    └── INT8_FP32_DIFFERENCES.csv       # Exact case-by-case difference breakdown

架构详情

  • 参数量:约 120M 参数
  • 层数:20 个 Transformer 层
  • 注意力:混合滑动窗口(32)与流形头注意力(mHC,4 条通道)
  • Engram 查找:多头滚动 n-gram 嵌入表(阶数 2 和 3,膨胀 3)
  • 量化技术:真正的 INT8 对称逐通道权重,配合优化的 DequantizeLinear Gather 嵌入。

致谢与许可证

  • 基础模型:Cactus Compute(mrfakename、gigglr)的 Cactus-Compute/needle3。
  • 许可证:Apache License 2.0。
  • 移植:由 Sakura AI 优化项目(webmp3)开发。
Downloads last month
138
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webmp3/Sakura-Needle3-ONNX

Quantized
(3)
this model