ArbiterOmni v6 β€” Multimodal Decision Arbiter

ArbiterOmni is a lightweight, production-grade multimodal decision-making model designed for high-stakes real-time arbitration in autonomous systems, robotics, and AI safety applications.

Hugging Face Spaces

Model Overview

Property Value
Architecture DeepSeek-V3 Style Shared+Routed MoE Fusion
Backbone OpenCLIP ViT-B-16-SigLIP (768-dim, WebLI 400M)
Trainable Params 30.3 M (fusion + decision + speculative)
Frozen Encoder 212 M (zero gradient updates)
Checkpoint (FP16) ~55 MB
Checkpoint (FP32) ~110 MB
Modalities Text Β· Image Β· Video Β· Audio
Inference Latency <3 ms (Tier-0 speculative exit), <50 ms (full MoE)

Key Innovations

1. Tier-0 Speculative Early-Exit Arbiter [AO-29]

A lightweight 128-dim draft head scores decisions in the GPU L1/L2 cache before the full MoE pathway. High-confidence decisions exit in <0.5 ms (vs. ~50 ms for the full Transformer). Calibrated by a latency gate comparing speculative vs. full-path confidence margins.

2. DeepSeek-V3 Style Shared + Domain-Specialized MoE Fusion [AO-30]

4 Transformer layers each containing:

  • 1 Shared Invariant Expert β€” captures universal cross-modal correlations
  • 4 Domain-Routed Experts β€” specialized per modality (vision, language, audio, spatio-temporal)
  • Top-2 soft routing with Switch load-balancing auxiliary loss

3. Test-Time Deliberation Tournament (TTC) [AO-31]

Candidates compete in a bracket tournament against dynamically harvested foils from the 100k resident memory bank. Winner is selected by confidence margin across $N$ deliberation rounds, provably reducing systematic bias under uncertainty.

4. Real-Time Adaptive Conformal Risk Control [AO-32]

Produces calibrated prediction sets with statistical coverage guarantees $(1 - \alpha = 95\%)$ scaled by an epistemic stability_index signal. Automatically triggers System2EscalationGate for ambiguous decisions.

5. Dual-Stream Real-Time Sensorium [AO-32]

Synchronized webcam video + microphone audio streaming arbitration with frame-level temporal buffering and async PCIe DMA double-buffering.

Architecture Diagram

Input Modalities
  Text ───────────────────────────────────────┐
  Image (980 patch tokens via LLaVA-NeXT) ─────   OpenCLIP ViT-B-16-SigLIP
  Video (64 dense frames) ─────────────────────   (frozen, CPU-resident)
  Audio β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                        β”‚ 768-dim embeddings
                        β–Ό
         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
         β”‚  Tier-0 Speculative Head   β”‚ ← 128-dim draft gate (<0.5ms)
         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                        β”‚ (low confidence β†’ full MoE)
                        β–Ό
         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
         β”‚  4Γ— MoE Transformer Layers                 β”‚
         β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
         β”‚  β”‚ Shared Expert   β”‚  β”‚ Top-2 Routed Γ—4  β”‚ β”‚
         β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                        β”‚ 512-dim fused context
                        β–Ό
         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
         β”‚  Dynamic Decision Head      β”‚ ← bilinear + perceptual skip
         β”‚  (candidate scoring)        β”‚
         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                        β”‚
                        β–Ό
         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
         β”‚  Conformal Risk Control     β”‚ ← 95% coverage prediction sets
         β”‚  + System 2 Escalation Gate β”‚
         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                        β”‚
                   DecisionResult

Benchmark Results

Benchmark v1 v2 v3 v4 v5 v6
ScienceQA Val Acc 52.1% 62.3% 71.6% 73.4% 78.9% ~82%*
SEED-Bench Val Acc β€” 58.4% 65.2% 67.8% 72.1% ~76%*
Speculative Exit Rate β€” β€” β€” β€” β€” ~68%
P50 Inference Latency 180ms 95ms 78ms 45ms 22ms <3ms†
Checkpoint Size 8.7 MB 72 MB 77 MB 77 MB 50 MB ~55 MB
Trainable Params 2.2M 12.5M 13.3M 13.3M 16M 30.3M

†Tier-0 speculative exit path.
*v6 final accuracy pending full 2-epoch training completion.

Usage

from arbiter_omni import ArbiterOmniEngine

# Load from Hugging Face Hub
engine = ArbiterOmniEngine.from_pretrained("cd1234/ArbiterOmni")

# Multimodal decision making
result = engine.decide(
    question="What is the safest immediate action?",
    candidates=[
        "decelerate smoothly and increase following distance",
        "proceed at maximum velocity without braking",
        "execute emergency lane change",
    ],
    text="Dense highway, heavy rainfall, vehicle decelerating 15m ahead.",
    test_time_deliberate=True,  # Enable TTC tournament
)

print(f"Decision      : {result.decision}")
print(f"Confidence    : {result.confidence:.1%}")
print(f"Entropy       : {result.entropy:.4f} nats")
print(f"Prediction Set: {result.prediction_set}")  # 95% conformal coverage
print(f"Stability     : {result.stability_index:.3f}")
print(f"Speculative   : {result.speculative_early_exit}")  # True = <0.5ms path

Installation

pip install arbiter-omni  # or:
git clone https://github.com/your-username/arbiter-omni
cd arbiter-omni
pip install -e .

Training

Trained on AMD Radeon RX 480 (4 GB GDDR5) + 8 GB DDR4 via PyTorch DirectML:

# v6 Production training (2 epochs, ~40 min on RX 480)
python scripts/train_v6.py --epochs 2 --batch-size 8

# v6-Max Hardware Saturation training (10 epochs, 768-dim, 8 experts, SWA)
python scripts/train_v6_max.py --epochs 10 --batch-size 8 --lr 5e-5

Hardware Requirements

Minimum Recommended
4 GB VRAM (GPU optional) AMD/NVIDIA GPU with β‰₯4 GB VRAM
4 GB RAM 8 GB RAM
CPU: Any x86-64 CPU: AMD Ryzen / Intel Core

Runs on CPU, CUDA, and AMD ROCm/DirectML.

Citation

@software{arbiteromni2026,
  title={ArbiterOmni: A Multimodal Decision Arbiter with Speculative Exit,
           MoE Fusion, and Conformal Risk Control},
  year={2026},
  url={https://huggingface.co/cd1234/ArbiterOmni},
  note={v6 Production Release}
}

License

Apache 2.0 β€” see LICENSE for details.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train cd1234/ArbiterOmni