d3-lite: Decision 3.0

d3-lite

d3-lite is the 0.8B multimodal foundation decision model of Decision 3.0, the decision models of vLLM Semantic Router. Give it an input (text or JSON, optionally with images and videos) and the questions you need answered: pick one of several options, say yes or no, or rate on a scale. It answers them all in one call and returns a probability for every answer, without generating text.

Parameters 0.85B, including the 0.10B vision encoder
Inputs Text or JSON, plus images and videos (several per request)
Decision types Choice · Yes / No · Score
License Apache-2.0

Highlights

  • Jev Decision Index 0.3, public suite: 36.93, measured with the official 0.3 kit on the released weights: all 140,178 public requests answered, none unsupported.
  • +16.4 on the public suite over Decision 2.0 (its 0.8B model: 20.57 on the board), ahead in all five areas.
  • Reads images: multiple images per request (PNG, JPEG or WebP), given as paths, URLs, PIL images or base64 data URLs; every question of the request sees all of them.
  • Reads videos: multiple videos per request (MP4, WebM, MOV or MKV), given as paths, URLs, base64 data URLs or frame arrays, read at 2 frames per second; Perception Test (multiple-choice video QA): 59.3 (internal evaluation).
  • Speed: a median of 13.1 ms for a text request, 48.8 ms for a request with an image and 331.5 ms for a request with a 10-second video, on one AMD Instinct MI325X GPU, one request at a time.
  • Many questions, one call: Choice, Yes / No and Score questions about the same input are answered together, each from its own forward pass over the input, with a probability for every option.

Quickstart

pip install "transformers==5.17.0" torch torchvision pillow opencv-python-headless safetensors accelerate
pip install flash-linear-attention  # optional: fast GPU kernels for the linear-attention layers
import json

from huggingface_hub import hf_hub_download
from transformers import AutoModel

model = AutoModel.from_pretrained("vllm-sr/d3-lite", trust_remote_code=True)

# Text
result = model.system_one(
    state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
    questions={
        "route": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {
                "returns": "Refunds, replacements and damaged deliveries",
                "billing": "Payments, invoices and charges",
                "technical": "Product setup and faults"
            }
        },
        "receipt": {
            "type": "noul",
            "instructions": "Does the customer have a receipt?"
        },
        "urgency": {
            "type": "score",
            "instructions": "How urgent is this request?",
            "criteria": [
                "Routine",
                "Soon",
                "Today"
            ]
        }
    },
)
print(json.dumps(result["answers"], indent=2))

# Text and an image (or several)
receipt = hf_hub_download("vllm-sr/d3-lite", "assets/example-receipt.png")
result = model.system_one(
    state="The customer says the blender arrived cracked and attached the receipt.",
    images=[receipt],  # local paths, http(s) URLs, PIL images or base64 data URLs
    questions={
        "route": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {
                "returns": "Refunds, replacements and damaged deliveries",
                "billing": "Payments, invoices and charges",
                "technical": "Product setup and faults"
            }
        },
        "on_receipt": {
            "type": "noul",
            "instructions": "Does the receipt list the blender?"
        },
        "payment": {
            "type": "choice",
            "instructions": "How was the order paid?",
            "criteria": {
                "card": None,
                "cash": None,
                "gift card": None
            }
        }
    },
)
print(json.dumps(result["answers"], indent=2))

# Text and a video (or several)
clip = hf_hub_download("vllm-sr/d3-lite", "assets/example-video.mp4")
result = model.system_one(
    state="A short clip from a test camera.",
    videos=[clip],  # local paths, http(s) URLs, base64 data URLs or frame arrays
    questions={
        "direction": {
            "type": "choice",
            "instructions": "Which way does the square move?",
            "criteria": {
                "right": "From left to right",
                "left": "From right to left",
                "still": "It does not move"
            }
        },
        "color_change": {
            "type": "noul",
            "instructions": "Does the square change color?"
        }
    },
)
print(json.dumps(result["answers"], indent=2))

# Or as a pipeline:
# transformers.pipeline("decision", model="vllm-sr/d3-lite", trust_remote_code=True)(state=..., questions=..., images=..., videos=...)

Images go before the text of the request, each read at up to 1.6 megapixels; videos follow the images, read at 2 frames per second (at most 32 frames spread over the whole video, each at up to 0.2 megapixels). Every question of the request sees all of them.

Evaluation

Text: Jev Decision Index 0.3.1

Model Jev Decision Index ↑ Public ↑ Same-skill tests ↑ New-domain tasks ↑
d3-lite 29.0 36.9 31.6 21.6
jiwo 0.8B 24.3 28.7 28.3 18.1
EXAONE-4.0-1.2B-JEV v0.3 23.3 30.7 26.9 16.4
OneJev 0.8B 20.5 22.2 21.8 19.9
JPT-0.8B Qwen3.5-0.8B LoRA 18.7 19.4 17.8 22.0
Decision 2.0 (0.8B) 17.5 20.6 21.3 13.9

Jev Decision Index against model size

Jev Decision Index by area: d3-lite and Decision 2.0 (0.8B)

Images: Jev Decision Index vision board 0.3.1

Model Vision Index ↑ Public ↑ Private ↑
d3-lite 46.9 49.4 44.3
JPT-0.8B 42.1 44.6 39.7
OneJev 0.8B 41.5 46.5 36.5
Intern-Decision-0.8B 36.3 39.0 33.7

d3-lite on the public vision benchmarks († approximate rebuild):

Benchmark d3-lite
CV-Bench 65.5
BLINK 37.7
RealWorldQA 35.9
CharXiv † 65.7
InfographicVQA † 74.9
Mind2Web † 53.1
Winoground 57.0
KIE (CORD+FUNSD) † 93.2
Moderation (Hateful Memes) 11.8
R-Bench-M 10.5
MMMU-Pro vision 8.3

Videos: Perception Test

Perception Test (validation) d3-lite
All 19,140 questions 59.3
Memory 48.1
Abstraction 50.9
Physics 52.3
Semantics 77.0

Multiple-choice video question answering (three options per question, chance 33.3): top-1 accuracy (%), videos read with the defaults above.

d3-lite: internal evaluation. Others: live board data, text 2026-10-10, vision 2026-10-09.

License

Apache-2.0 (LICENSE). Built on Qwen/Qwen3.5-0.8B (Apache-2.0).

Citation

@misc{d3_lite_2026,
  title        = {{d3-lite}: A Multimodal Foundation Decision Model},
  author       = {{vLLM Semantic Router Team}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/vllm-sr/d3-lite}}
}

Trained on AMD Instinct MI325X GPUs.

Downloads last month
83
Safetensors
Model size
0.9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vllm-sr/d3-lite

Finetuned
(496)
this model
Finetunes
1 model
Quantizations
2 models

Collection including vllm-sr/d3-lite