Vela 2.0 0.3B 路 Core ML
vllm-sr/Vela-2.0-0.3B (vLLM Semantic Router 脳 KR Labs) converted to Core ML for Apple silicon: routing decisions, safety checks (prompt attack, harmful request) and text spans (personal information, claims not supported by a source) from one ModernBERT encoder, on device.
| Source | vllm-sr/Vela-2.0-0.3B @ d6f03aa9 |
| Package | Vela2Encoder.mlpackage, fp16, 592 MB: functions L128, L256, L512, L1024 (encoder tokens) |
| Host side | heads.bin (fp32 readout heads), tokenizer.json, calibration.json, coreml_config.json |
| Runtime | Swift: Vela2Manager in FluidUse (tokenizer, assembly, heads, calibration, span decoder) |
| Requires | macOS 15+ / iOS 18+ |
Fidelity
The Swift runtime reproduces the release's own engine (vela2_inference.py) on the same encoder: 38 requests (PII,
prompt attack, harm, routing, hallucination; English, German, French, Spanish, Chinese, Japanese, Hindi, Arabic, URLs,
e-mails) give 0 token mismatches, 0 / 120 choice answers and 0 / 38 span sets different (reports/). The Core ML
encoder against the PyTorch release on 8 varied requests: identical choices and span sets, probabilities within 0.003
(GPU and Neural Engine).
Speed (MacBook Pro M5 Pro)
| Encoder function | GPU | Neural Engine |
|---|---|---|
| L128 | 4.4 ms | 3.5 ms |
| L256 | 5.0 ms | 9.0 ms |
| L512 | 8.3 ms | 25.1 ms |
| L1024 | 16.5 ms | 73.5 ms |
99.4 % of encoder ops are placed on the Neural Engine (only the embedding lookup is not), but it is only faster up to 128
tokens, so the runtime uses it for L128 and the GPU above. The question schema counts toward the length: an attack +
harm screen (~92 schema tokens) on a short message runs on the Neural Engine in ~3.5 ms; a full guardrail (route, attack,
harm, fact-check, 17 PII labels: ~535 schema tokens) runs in L1024 on the GPU in ~17 ms. The release's PyTorch runtime
on the same Mac's CPU takes ~165 ms for a 4-question request.
Files and licences
The encoder weights and readout are Apache-2.0 (see LICENSE, NOTICE, LICENSING_STATUS.md). The tokenizer is
Gemma-origin material and remains subject to the Gemma Terms of Use and Prohibited Use Policy in LICENSES/gemma/
(see DISTRIBUTION_TERMS.md); by using or redistributing it you agree to those terms. What this conversion changed is
listed in MODIFICATIONS.md. conversion/ holds the export, packaging and parity scripts.
- Downloads last month
- -
Model tree for FluidInference/vela-2.0-0.3b-coreml
Base model
jhu-clsp/mmBERT-base