VTX-JEV-3: Position-Aware System 1 Decision & Routing Engine

VTX-JEV-3 is the next-generation, position-aware System 1 Decision Engine fine-tuned on the multilingual SargeDev/jev-distill-corpus-v3 multi-domain corpus.

Built upon VTXAI/VTX-JEV-2, VTX-JEV-3 introduces position-aware attention pooling and sparse token delta calibration, resolving permutation blindness and boosting multi-task classification accuracy across categorical routing (Choice), binary gating (Noul), and ordinal priority scoring (Score).


Key Improvements over VTX-JEV-2:

  1. Position & Sequence Awareness: Mean pooling in static models is inherently permutation invariant. VTX-JEV-3 incorporates a position-gated attention pooler that distinguishes word order and argument reversal ("intern fired manager" vs "manager fired intern") while preserving semantic paraphrases.
  2. Fine-Tuned on SargeDev/jev-distill-corpus-v3: Distilled across chemistry, fact-checking, safety, security, code, and dialogue routing scenarios.
  3. Dual Formats:
    • model_lf2.safetensors (19.5 MB, 2-bit LF2 quantized representation for sub-millisecond CPU execution).
    • model.safetensors (FP32 format, Model2Vec and Sentence-Transformers compatible).
  4. Sub-Millisecond Inference: Runs purely in NumPy on CPU (~0.8 ms – 1.4 ms decision latency).

Quick Start: System 1 Decision Interface

from inference import JevClient, Choice, Noul, Score

# Load model (automatically uses ultra-compact 19.5MB 2-bit LF2 table)
client = JevClient.from_pretrained("VTXAI/VTX-JEV-3")

response = client.system_one(
    state="Production database dropped replica connection and checkout is failing.",
    questions={
        "urgent": Noul("Is this an urgent production emergency?"),
        "routing": Choice("Select team", {
            "sre": "Infrastructure and Database Reliability",
            "sales": "Sales and Marketing Inquiries"
        }),
        "severity": Score("Rate incident severity", ["Minor", "Major", "Critical"])
    }
)

print(response.nouls["urgent"].noul)          # True
print(response.choices["routing"].choice)     # 'sre'
print(response.scores["severity"].score)      # 'Critical'
print(f"Latency: {response.latency_ms:.2f} ms")
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VTXAI/VTX-JEV-3

Base model

VTXAI/VTX-JEV-2
Finetuned
(1)
this model