WebBrain Compass Tiny XS v3.1 — 32K ONNX/WebGPU

WebBrain Compass Tiny XS v3.1 is a browser-native ONNX/WebGPU package for the Spark-X2.5-1.7B-based Compass Tiny XS v3 fine-tune. It supports low-latency decision making and structured tool use in the WebBrain runtime. This release extends the total input + output limit to 32,768 tokens; it does not change the trained weights or the ONNX graph from the tested 4K package.

It supports three WebBrain behaviors:

  1. Ask & Clarify — Answer directly or request missing information.
  2. Direct Compact Tool Execution — Select grounded browser tools and arguments.
  3. Safe Escalation & Abstention — Pause or defer when evidence is insufficient.

The repository name is WebBrain Compass Tiny XS v3.1. For ABI compatibility, the bundled service still reports model ID webbrain-compass-tiny-v3-onnx-fp16.


Role in WebBrain

User request
     |
     v
WebBrain observation & policy layer
     |
     v
WebBrain Compass Tiny XS v3.1 (WebGPU / ONNX)
     +-- Clarify or answer directly
     +-- Emit a grounded browser tool call
     +-- Abstain or safely escalate

The outer WebBrain runtime must validate tool schemas, action grounding, browser-state freshness and security policy. Model output is not evidence that an external action succeeded.


Quickstart & Loading

Download the entire pinned repository revision and verify sizes and SHA-256 hashes against FILES.sha256.json. The package includes onnx/model_fp16.onnx_data, tokenizer, native chat_template.jinja, graph-abi.json, loader and pinned vendor runtime files. The ONNX graph requires the supplied native Spark loader; it is not a generic Llama-style AutoModelForCausalLM export.

hf download webbrain-one/webbrain-compass-tiny-xs-v3.1-onnx --local-dir ./compass-tiny-xs-v3.1-onnx

The Windows loopback launcher requires Python 3.11 and Playwright 1.58.0:

python -m pip install playwright==1.58.0
python runtime/browser_service.py --artifact-root C:/models/compass-tiny-xs-v3.1-onnx --package-mode --run-tag local-32k-1 --port 8561

Use a fresh --run-tag for each launch. The validated launcher selects the RTX 5090 and streams ONNX external data through ORT; Chrome cannot allocate a normal approximately 4 GB ArrayBuffer for these weights. Other GPUs and browsers require separate validation.

OpenAI-compatible interface

base_url: http://127.0.0.1:8561/v1
model: webbrain-compass-tiny-v3-onnx-fp16
temperature: 0
stream: false
max_tokens: 1..2048
tokenized input + requested output: <= 32768 tokens

GET /health, GET /v1/models and POST /v1/chat/completions are available. Standard OpenAI function schemas are accepted. Malformed tool output remains visible via x_parser_error; it is not repaired or retried. Unsupported sampling and streaming requests fail closed. No helper model or cloud fallback is used.

The native Spark loader prefills in ordered 512-token chunks, retaining the entire prompt in KV-cache. It rejects over-limit requests rather than silently truncating them. Batch size is 1 and generation is greedy.

WebBrain integration status

The current WebBrain extension remains pinned to the older 4K revision and hard-limits Tiny XS v3 to 4K. Downloading this v3.1 package alone does not enable 32K in the WebBrain UI; a separate integration change is required.


Technical Specifications & Validation

Precision

Weights and KV-cache use FP16 storage. Matrix multiplications, residuals, normalization, RoPE and softmax use FP32 compute. This is not q4f16. The ONNX graph and external weight bytes are identical to the tested 4K release; v3.1 changes the runtime context contract and loader, not training.

Numerical and browser checks

  • Reference: Independent BF16 native model.
  • Logit gates: All checked logits finite; relative RMS error <= 2%; cosine similarity >= 0.999.
  • Context: CPU ONNX and Chrome/WebGPU checks cover 4K, 8K and 32K cache boundaries, including cached decode. The 32K boundary check passed with a maximum observed relative RMS of 0.00505 and minimum cosine above 0.999996.
  • Runtime: Long-context generation, native tool-call parsing, limits, recovery, latency and GPU-memory observations are recorded in provenance/. These are runtime checks, not task-quality or leaderboard results.

The first long-context smoke attempt timed out on a local loopback /template request before model generation. The service remained healthy; an unchanged full rerun passed. Both attempts are preserved in provenance/.

The validated environment was Windows, NVIDIA RTX 5090, Chrome 153, ONNX Runtime Web 1.27.0 and Transformers.js 4.2.0. Other platforms are not validated by this package.


License

Noncommercial research only. This model and its fine-tuned weights are not cleared for commercial deployment. The upstream Spark license is included in LICENSE; that license alone does not change this release's usage restriction.

Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webbrain-one/webbrain-compass-tiny-xs-v3.1-onnx