Decision-1.0-Eos-0.8B 路 NEURA EDIT
High-throughput on-device neural decision model for in-cabin multi-intent routing & semantic classification. Signal over tokens.
- Official Codebase: github.com/neura-edit/omni-intent
- Base Architecture: 0.75B Qwen-based neural backbone with dual-head classification topology (Multi-label
noul+ Single-choicechoice). - Precision: FP32 (
1.5GB) / INT8 (380MB) - Average Forward Latency: ~140ms on Apple Silicon / ~35ms on INT8 NPU.
Overview
Traditional LLMs/SLMs suffer from severe autoregressive generation latency (800ms~2500ms) and formatting hallucinations. Decision-Eos bypasses token-by-token generation entirely, producing deterministic, typed intent probability distributions across 8+ automotive domains in a single forward pass.
Used as the core neural backbone in the NEURA EDIT OmniIntent Decision Engine.
Quick Usage with Ollaya
# Pull model directly
ollaya pull decision:eos
# Run single inference
ollaya run decision:eos "Turn on the AC and play some jazz"
Supported In-Cabin Domains
climate: Air conditioning, temperature, fan zones, heating & coolingmusic: Music playback, track switching, 21 acoustic genres & moodsnavigation: GPS destination setting, traffic avoidance, route guidanceseat: Seat heating, ventilation, massagewindow: Window and sunroof operationsphone: Contact dialing, call pickup, hangup, redialquery: Weather, time, date, vehicle telemetry, assistant Q&Aother: Out-of-domain fallback category
Citation & Attribution
Based on the open decision model research by the vLLM Semantic Router initiative and integrated into the NEURA EDIT in-cabin decision intelligence framework.
- Author: NEURA EDIT
- License: Apache 2.0