VTX-JEV-2

VTX-JEV-2 is an ultra-fast, non-autoregressive System 1 Decision & Routing Engine distilled from VTXAI/vtx-jev-laya-multilingual (jhu-clsp/mmBERT-base multilingual foundation) into a 256-dimensional static embedding architecture with native LF2 2-bit quantization.

It is engineered for sub-millisecond CPU decision-making, categorical routing (Choice), binary gating (Noul), and ordinal priority calibration (Score) across 100+ languages, including Indian scripts (Hindi, Bengali, Tamil, Telugu, Marathi, Urdu) and international languages.


Highlights

  • ⚡ Sub-Millisecond Execution: Decision latency is ~1.0 ms – 1.8 ms for full multi-primitive system_one calls, and ~0.05 ms for single-vector lookups on CPU.
  • 🎯 Typed System 1 Primitives: Out-of-the-box support for Noul (binary gating), Choice (categorical routing), and Score (ordinal severity calibration).
  • 🗜️ Dual Checkpoint Formats:
    • model.safetensors (FP32 / Model2Vec / Sentence Transformers compatible)
    • model_lf2.safetensors (Ultra-compact 2-bit LF2 quantization — only 19.5 MB in RAM!)
  • 🌐 True Multilingual Vocabulary: Built on a 256,000-token tokenizer with zero subword fragmentation on Devanagari, Bengali, Tamil, Telugu, Arabic, and Cyrillic scripts.
  • 📦 Zero-Config Self-Contained Inference: Bundles inference.py for immediate plug-and-play usage.

Quick Start: System 1 Decision Interface

Using the bundled inference.py:

from inference import JevClient, Choice, Noul, Score

# Load model from Hugging Face Hub (or local folder)
client = JevClient.from_pretrained("VTXAI/VTX-JEV-2")

# Execute multi-primitive decision in ~1.5 ms on CPU
response = client.system_one(
    state="I was charged twice and production is unavailable.",
    questions={
        "refund": Noul("Does the customer request a refund?"),
        "team": Choice(
            "Which team should handle this?",
            {"billing": "Payments and refunds", "technical": "Production outage and servers"},
        ),
        "severity": Score(
            "How severe is the impact?",
            ["Minor", "Major", "Critical"],
        ),
    },
)

print(response.nouls["refund"].noul)       # True
print(response.choices["team"].choice)     # 'technical'
print(response.scores["severity"].score)   # 'Major'
print(response.to_dict())

Multilingual Support (Hindi, Bengali, Tamil, etc.)

response_hi = client.system_one(
    state="डेटाबेस सर्वर डाउन हो गया है और उपयोगकर्ता भुगतान नहीं कर पा रहे हैं।",
    questions={
        "urgent": Noul("क्या यह एक आपातकालीन समस्या है?"),
        "team": Choice(
            "इस समस्या को कौन संभालेगा?",
            {"devops": "सर्वर और डेटाबेस इंजीनियर", "sales": "बिक्री और प्रचार"},
        ),
        "level": Score(
            "प्राथमिकता स्तर क्या है?",
            ["P3 - सामान्य", "P2 - उच्च", "P1 - गंभीर"],
        ),
    },
)

print(response_hi.choices["team"].choice)    # 'devops'
print(response_hi.scores["level"].score)      # 'P1 - गंभीर'
print(f"Latency: {response_hi.latency_ms} ms") # ~1.0 ms

Usage with Sentence Transformers

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("VTXAI/VTX-JEV-2")
embeddings = model.encode([
    "Incident: Database connection timed out",
    "डेटाबेस सर्वर डाउन हो गया है",
    "User inquiry about account balance"
])
print(embeddings.shape) # (3, 256)

Usage with Model2Vec

from model2vec import StaticModel

model = StaticModel.from_pretrained("VTXAI/VTX-JEV-2")
embeddings = model.encode(["Fast vector routing", "வணக்கம்"])
print(embeddings.shape) # (2, 256)

Performance & Latency Benchmarks (CPU)

Benchmark FP32 Static (model.safetensors) LF2 2-Bit (model_lf2.safetensors)
Model Size (RAM / Disk) 124.9 MB 19.5 MB (6.4x compression)
Single Vector Latency 0.35 ms 0.05 ms (50 microseconds)
Throughput (Batch=20) 6,196 queries/sec 43,718 queries/sec
Full System 1 Decision Latency 1.0 ms – 1.8 ms < 0.5 ms
Downloads last month
-
Safetensors
Model size
65.5M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VTXAI/VTX-JEV-2

Finetunes
1 model