Feature Extraction
sentence-transformers
Safetensors
Model2Vec
static-embeddings
decision-engine
jev
lf2
2bit-quantization
cpu-optimized
Instructions to use VTXAI/VTX-JEV-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use VTXAI/VTX-JEV-2 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("VTXAI/VTX-JEV-2") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Model2Vec
How to use VTXAI/VTX-JEV-2 with Model2Vec:
from model2vec import StaticModel model = StaticModel.from_pretrained("VTXAI/VTX-JEV-2") - Notebooks
- Google Colab
- Kaggle
VTX-JEV-2
VTX-JEV-2 is an ultra-fast, non-autoregressive System 1 Decision & Routing Engine distilled from VTXAI/vtx-jev-laya-multilingual (jhu-clsp/mmBERT-base multilingual foundation) into a 256-dimensional static embedding architecture with native LF2 2-bit quantization.
It is engineered for sub-millisecond CPU decision-making, categorical routing (Choice), binary gating (Noul), and ordinal priority calibration (Score) across 100+ languages, including Indian scripts (Hindi, Bengali, Tamil, Telugu, Marathi, Urdu) and international languages.
Highlights
- ⚡ Sub-Millisecond Execution: Decision latency is ~1.0 ms – 1.8 ms for full multi-primitive
system_onecalls, and ~0.05 ms for single-vector lookups on CPU. - 🎯 Typed System 1 Primitives: Out-of-the-box support for
Noul(binary gating),Choice(categorical routing), andScore(ordinal severity calibration). - 🗜️ Dual Checkpoint Formats:
model.safetensors(FP32 / Model2Vec / Sentence Transformers compatible)model_lf2.safetensors(Ultra-compact 2-bit LF2 quantization — only 19.5 MB in RAM!)
- 🌐 True Multilingual Vocabulary: Built on a 256,000-token tokenizer with zero subword fragmentation on Devanagari, Bengali, Tamil, Telugu, Arabic, and Cyrillic scripts.
- 📦 Zero-Config Self-Contained Inference: Bundles
inference.pyfor immediate plug-and-play usage.
Quick Start: System 1 Decision Interface
Using the bundled inference.py:
from inference import JevClient, Choice, Noul, Score
# Load model from Hugging Face Hub (or local folder)
client = JevClient.from_pretrained("VTXAI/VTX-JEV-2")
# Execute multi-primitive decision in ~1.5 ms on CPU
response = client.system_one(
state="I was charged twice and production is unavailable.",
questions={
"refund": Noul("Does the customer request a refund?"),
"team": Choice(
"Which team should handle this?",
{"billing": "Payments and refunds", "technical": "Production outage and servers"},
),
"severity": Score(
"How severe is the impact?",
["Minor", "Major", "Critical"],
),
},
)
print(response.nouls["refund"].noul) # True
print(response.choices["team"].choice) # 'technical'
print(response.scores["severity"].score) # 'Major'
print(response.to_dict())
Multilingual Support (Hindi, Bengali, Tamil, etc.)
response_hi = client.system_one(
state="डेटाबेस सर्वर डाउन हो गया है और उपयोगकर्ता भुगतान नहीं कर पा रहे हैं।",
questions={
"urgent": Noul("क्या यह एक आपातकालीन समस्या है?"),
"team": Choice(
"इस समस्या को कौन संभालेगा?",
{"devops": "सर्वर और डेटाबेस इंजीनियर", "sales": "बिक्री और प्रचार"},
),
"level": Score(
"प्राथमिकता स्तर क्या है?",
["P3 - सामान्य", "P2 - उच्च", "P1 - गंभीर"],
),
},
)
print(response_hi.choices["team"].choice) # 'devops'
print(response_hi.scores["level"].score) # 'P1 - गंभीर'
print(f"Latency: {response_hi.latency_ms} ms") # ~1.0 ms
Usage with Sentence Transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("VTXAI/VTX-JEV-2")
embeddings = model.encode([
"Incident: Database connection timed out",
"डेटाबेस सर्वर डाउन हो गया है",
"User inquiry about account balance"
])
print(embeddings.shape) # (3, 256)
Usage with Model2Vec
from model2vec import StaticModel
model = StaticModel.from_pretrained("VTXAI/VTX-JEV-2")
embeddings = model.encode(["Fast vector routing", "வணக்கம்"])
print(embeddings.shape) # (2, 256)
Performance & Latency Benchmarks (CPU)
| Benchmark | FP32 Static (model.safetensors) |
LF2 2-Bit (model_lf2.safetensors) |
|---|---|---|
| Model Size (RAM / Disk) | 124.9 MB | 19.5 MB (6.4x compression) |
| Single Vector Latency | 0.35 ms | 0.05 ms (50 microseconds) |
| Throughput (Batch=20) | 6,196 queries/sec | 43,718 queries/sec |
| Full System 1 Decision Latency | 1.0 ms – 1.8 ms | < 0.5 ms |
- Downloads last month
- -