Feature Extraction
sentence-transformers
Safetensors
Model2Vec
vtx_jev
static-embeddings
decision-engine
jev
lf2
2bit-quantization
cpu-optimized
Instructions to use VTXAI/VTX-JEV-3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use VTXAI/VTX-JEV-3 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("VTXAI/VTX-JEV-3") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Model2Vec
How to use VTXAI/VTX-JEV-3 with Model2Vec:
from model2vec import StaticModel model = StaticModel.from_pretrained("VTXAI/VTX-JEV-3") - Notebooks
- Google Colab
- Kaggle
VTX-JEV-3: Position-Aware System 1 Decision & Routing Engine
VTX-JEV-3 is the next-generation, position-aware System 1 Decision Engine fine-tuned on the multilingual SargeDev/jev-distill-corpus-v3 multi-domain corpus.
Built upon VTXAI/VTX-JEV-2, VTX-JEV-3 introduces position-aware attention pooling and sparse token delta calibration, resolving permutation blindness and boosting multi-task classification accuracy across categorical routing (Choice), binary gating (Noul), and ordinal priority scoring (Score).
Key Improvements over VTX-JEV-2:
- Position & Sequence Awareness: Mean pooling in static models is inherently permutation invariant. VTX-JEV-3 incorporates a position-gated attention pooler that distinguishes word order and argument reversal ("intern fired manager" vs "manager fired intern") while preserving semantic paraphrases.
- Fine-Tuned on SargeDev/jev-distill-corpus-v3: Distilled across chemistry, fact-checking, safety, security, code, and dialogue routing scenarios.
- Dual Formats:
model_lf2.safetensors(19.5 MB, 2-bit LF2 quantized representation for sub-millisecond CPU execution).model.safetensors(FP32 format, Model2Vec and Sentence-Transformers compatible).
- Sub-Millisecond Inference: Runs purely in NumPy on CPU (~0.8 ms – 1.4 ms decision latency).
Quick Start: System 1 Decision Interface
from inference import JevClient, Choice, Noul, Score
# Load model (automatically uses ultra-compact 19.5MB 2-bit LF2 table)
client = JevClient.from_pretrained("VTXAI/VTX-JEV-3")
response = client.system_one(
state="Production database dropped replica connection and checkout is failing.",
questions={
"urgent": Noul("Is this an urgent production emergency?"),
"routing": Choice("Select team", {
"sre": "Infrastructure and Database Reliability",
"sales": "Sales and Marketing Inquiries"
}),
"severity": Score("Rate incident severity", ["Minor", "Major", "Critical"])
}
)
print(response.nouls["urgent"].noul) # True
print(response.choices["routing"].choice) # 'sre'
print(response.scores["severity"].score) # 'Critical'
print(f"Latency: {response.latency_ms:.2f} ms")
- Downloads last month
- -
Model tree for VTXAI/VTX-JEV-3
Base model
VTXAI/VTX-JEV-2