Vela 2.0 0.3B 路 Core ML

vllm-sr/Vela-2.0-0.3B (vLLM Semantic Router 脳 KR Labs) converted to Core ML for Apple silicon: routing decisions, safety checks (prompt attack, harmful request) and text spans (personal information, claims not supported by a source) from one ModernBERT encoder, on device.

Source vllm-sr/Vela-2.0-0.3B @ d6f03aa9
Package Vela2Encoder.mlpackage, fp16, 592 MB: functions L128, L256, L512, L1024 (encoder tokens)
Host side heads.bin (fp32 readout heads), tokenizer.json, calibration.json, coreml_config.json
Runtime Swift: Vela2Manager in FluidUse (tokenizer, assembly, heads, calibration, span decoder)
Requires macOS 15+ / iOS 18+

Fidelity

The Swift runtime reproduces the release's own engine (vela2_inference.py) on the same encoder: 38 requests (PII, prompt attack, harm, routing, hallucination; English, German, French, Spanish, Chinese, Japanese, Hindi, Arabic, URLs, e-mails) give 0 token mismatches, 0 / 120 choice answers and 0 / 38 span sets different (reports/). The Core ML encoder against the PyTorch release on 8 varied requests: identical choices and span sets, probabilities within 0.003 (GPU and Neural Engine).

Speed (MacBook Pro M5 Pro)

Encoder function GPU Neural Engine
L128 4.4 ms 3.5 ms
L256 5.0 ms 9.0 ms
L512 8.3 ms 25.1 ms
L1024 16.5 ms 73.5 ms

99.4 % of encoder ops are placed on the Neural Engine (only the embedding lookup is not), but it is only faster up to 128 tokens, so the runtime uses it for L128 and the GPU above. The question schema counts toward the length: an attack + harm screen (~92 schema tokens) on a short message runs on the Neural Engine in ~3.5 ms; a full guardrail (route, attack, harm, fact-check, 17 PII labels: ~535 schema tokens) runs in L1024 on the GPU in ~17 ms. The release's PyTorch runtime on the same Mac's CPU takes ~165 ms for a 4-question request.

Files and licences

The encoder weights and readout are Apache-2.0 (see LICENSE, NOTICE, LICENSING_STATUS.md). The tokenizer is Gemma-origin material and remains subject to the Gemma Terms of Use and Prohibited Use Policy in LICENSES/gemma/ (see DISTRIBUTION_TERMS.md); by using or redistributing it you agree to those terms. What this conversion changed is listed in MODIFICATIONS.md. conversion/ holds the export, packaging and parity scripts.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for FluidInference/vela-2.0-0.3b-coreml

Quantized
(1)
this model