ArbiterOmni v6 β Multimodal Decision Arbiter
ArbiterOmni is a lightweight, production-grade multimodal decision-making model designed for high-stakes real-time arbitration in autonomous systems, robotics, and AI safety applications.
Model Overview
| Property | Value |
|---|---|
| Architecture | DeepSeek-V3 Style Shared+Routed MoE Fusion |
| Backbone | OpenCLIP ViT-B-16-SigLIP (768-dim, WebLI 400M) |
| Trainable Params | 30.3 M (fusion + decision + speculative) |
| Frozen Encoder | 212 M (zero gradient updates) |
| Checkpoint (FP16) | ~55 MB |
| Checkpoint (FP32) | ~110 MB |
| Modalities | Text Β· Image Β· Video Β· Audio |
| Inference Latency | <3 ms (Tier-0 speculative exit), <50 ms (full MoE) |
Key Innovations
1. Tier-0 Speculative Early-Exit Arbiter [AO-29]
A lightweight 128-dim draft head scores decisions in the GPU L1/L2 cache before the full MoE pathway. High-confidence decisions exit in <0.5 ms (vs. ~50 ms for the full Transformer). Calibrated by a latency gate comparing speculative vs. full-path confidence margins.
2. DeepSeek-V3 Style Shared + Domain-Specialized MoE Fusion [AO-30]
4 Transformer layers each containing:
- 1 Shared Invariant Expert β captures universal cross-modal correlations
- 4 Domain-Routed Experts β specialized per modality (vision, language, audio, spatio-temporal)
- Top-2 soft routing with Switch load-balancing auxiliary loss
3. Test-Time Deliberation Tournament (TTC) [AO-31]
Candidates compete in a bracket tournament against dynamically harvested foils from the 100k resident memory bank. Winner is selected by confidence margin across $N$ deliberation rounds, provably reducing systematic bias under uncertainty.
4. Real-Time Adaptive Conformal Risk Control [AO-32]
Produces calibrated prediction sets with statistical coverage guarantees
$(1 - \alpha = 95\%)$ scaled by an epistemic stability_index signal.
Automatically triggers System2EscalationGate for ambiguous decisions.
5. Dual-Stream Real-Time Sensorium [AO-32]
Synchronized webcam video + microphone audio streaming arbitration with frame-level temporal buffering and async PCIe DMA double-buffering.
Architecture Diagram
Input Modalities
Text ββββββββββββββββββββββββββββββββββββββββ
Image (980 patch tokens via LLaVA-NeXT) βββββ€ OpenCLIP ViT-B-16-SigLIP
Video (64 dense frames) βββββββββββββββββββββ€ (frozen, CPU-resident)
Audio ββββββββββββββββββββββββββββββββββββββββ
β 768-dim embeddings
βΌ
βββββββββββββββββββββββββββββββ
β Tier-0 Speculative Head β β 128-dim draft gate (<0.5ms)
βββββββββββββββββββββββββββββββ
β (low confidence β full MoE)
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββ
β 4Γ MoE Transformer Layers β
β βββββββββββββββββββ ββββββββββββββββββββ β
β β Shared Expert β β Top-2 Routed Γ4 β β
β βββββββββββββββββββ ββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββ
β 512-dim fused context
βΌ
βββββββββββββββββββββββββββββββ
β Dynamic Decision Head β β bilinear + perceptual skip
β (candidate scoring) β
βββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββ
β Conformal Risk Control β β 95% coverage prediction sets
β + System 2 Escalation Gate β
βββββββββββββββββββββββββββββββ
β
DecisionResult
Benchmark Results
| Benchmark | v1 | v2 | v3 | v4 | v5 | v6 |
|---|---|---|---|---|---|---|
| ScienceQA Val Acc | 52.1% | 62.3% | 71.6% | 73.4% | 78.9% | ~82%* |
| SEED-Bench Val Acc | β | 58.4% | 65.2% | 67.8% | 72.1% | ~76%* |
| Speculative Exit Rate | β | β | β | β | β | ~68% |
| P50 Inference Latency | 180ms | 95ms | 78ms | 45ms | 22ms | <3msβ |
| Checkpoint Size | 8.7 MB | 72 MB | 77 MB | 77 MB | 50 MB | ~55 MB |
| Trainable Params | 2.2M | 12.5M | 13.3M | 13.3M | 16M | 30.3M |
β Tier-0 speculative exit path.
*v6 final accuracy pending full 2-epoch training completion.
Usage
from arbiter_omni import ArbiterOmniEngine
# Load from Hugging Face Hub
engine = ArbiterOmniEngine.from_pretrained("cd1234/ArbiterOmni")
# Multimodal decision making
result = engine.decide(
question="What is the safest immediate action?",
candidates=[
"decelerate smoothly and increase following distance",
"proceed at maximum velocity without braking",
"execute emergency lane change",
],
text="Dense highway, heavy rainfall, vehicle decelerating 15m ahead.",
test_time_deliberate=True, # Enable TTC tournament
)
print(f"Decision : {result.decision}")
print(f"Confidence : {result.confidence:.1%}")
print(f"Entropy : {result.entropy:.4f} nats")
print(f"Prediction Set: {result.prediction_set}") # 95% conformal coverage
print(f"Stability : {result.stability_index:.3f}")
print(f"Speculative : {result.speculative_early_exit}") # True = <0.5ms path
Installation
pip install arbiter-omni # or:
git clone https://github.com/your-username/arbiter-omni
cd arbiter-omni
pip install -e .
Training
Trained on AMD Radeon RX 480 (4 GB GDDR5) + 8 GB DDR4 via PyTorch DirectML:
# v6 Production training (2 epochs, ~40 min on RX 480)
python scripts/train_v6.py --epochs 2 --batch-size 8
# v6-Max Hardware Saturation training (10 epochs, 768-dim, 8 experts, SWA)
python scripts/train_v6_max.py --epochs 10 --batch-size 8 --lr 5e-5
Hardware Requirements
| Minimum | Recommended |
|---|---|
| 4 GB VRAM (GPU optional) | AMD/NVIDIA GPU with β₯4 GB VRAM |
| 4 GB RAM | 8 GB RAM |
| CPU: Any x86-64 | CPU: AMD Ryzen / Intel Core |
Runs on CPU, CUDA, and AMD ROCm/DirectML.
Citation
@software{arbiteromni2026,
title={ArbiterOmni: A Multimodal Decision Arbiter with Speculative Exit,
MoE Fusion, and Conformal Risk Control},
year={2026},
url={https://huggingface.co/cd1234/ArbiterOmni},
note={v6 Production Release}
}
License
Apache 2.0 β see LICENSE for details.