Tasksource-JEV-Nano-v0

A small decision model that picks the best option from a list โ€” fraud routings, intents, topics, sentiments โ€” using token-level late interaction instead of a classifier head.

It is LateOn (~149M parameters) fine-tuned on 512k real decisions covering three judgment types: picking one option (choice), yes/no questions (noul), and graded scores (score). No teacher model was used.

Why this architecture. The situation (state + question) is encoded once into token vectors and reused for every candidate set, while each option is encoded independently. That gives two properties classifier heads don't have: caching the situation across decisions, and outputs that don't depend on option order (verified exactly permutation-equivariant).

  • Base model: lightonai/LateOn
  • Training data: tasksource/tasksource-jev-typed-decisions (pinned revision 071f0cf2), 512k-decision stream, natural mix of the three judgment types, 519 tasks
  • Checkpoint selection: Tasksource unseen tasks only (step 4000 of 8000)
  • Training code: decision-models repo (scripts/train_multivector.py, canonical run)

How it scores

Benchmark (1k examples each, same samples for every model) v0
Tasksource, unseen tasks 58.9%
Tasksource, unseen test tasks 53.0%
Typed Decisions 42.6%
AG News 73.8%
Emotion 49.0%
Banking77 (77 options) 44.7%
Fast Decisions (2,900 cases, 17 domains) 41.3% exact match
classifier-benchmark v2 (866 cases) 57.5 micro / 57.2 macro

Compact peers (all under 200M, same 1k samples)

Model Size AG News Emotion Banking77 Typed cb-v2 micro
Tasksource-JEV-Nano-v0 149M 73.8 49.0 44.7 42.6 57.5
GLiClass Base v3 187M 76.9 51.8 52.4 47.3 51.2
GLiClass Modern-Base v3 151M 75.7 56.1 38.7 48.6 43.9
GLiNER2.5 Base 194M 76.2 56.6 69.7 43.9 60.0
GLiNER2.5 Small 74M 70.4 54.2 66.5 34.0 53.5
GLiClass Edge v3 33M 65.5 50.1 23.2 40.0 35.3

Larger peers (GLiClass Modern-Large 399M, GLiNER2.5-Multi 287M, Laya-Multilingual 322M) are omitted from this table; Laya's 90.0 AG News is training overlap (AG News is in its training mix). v0 trades some raw high-cardinality accuracy for state caching and order-independence, which none of the classifier heads offer.

Use it (only public packages)

pip install -U pylate torch
from pylate import models
import torch

model = models.ColBERT("tasksource/tasksource-jev-nano-v0")

state = "Customer reports unauthorized international wire transfer of $4,500."
question = "Select the appropriate fraud mitigation routing:"
options = [
    "approve and monitor silently",
    "challenge with a push notification",
    "freeze the account and call the customer",
    "decline and file a report",
]

# Situation encoded once; each option scored against it (MaxSim)
context = [state + "\nQuestion: " + question]
ctx = model.encode(context, is_query=False, convert_to_tensor=True)[0]
opts = model.encode(options, is_query=True, convert_to_tensor=True)
scores = torch.stack([(o @ ctx.T).max(dim=1).values.sum() for o in opts])

temperature = 0.53  # learned temperature of this release
probs = torch.softmax(scores / temperature, dim=0)
for opt, p in zip(options, probs):
    print(f"  {opt:45s} -> {p:.4f}")

Yes/no and graded decisions work the same way with two options or an ordered scale; encode the situation once and reuse it across as many candidate sets as you like.

Citation

@misc{sileo2026jevnanov0,
  title={Tasksource-JEV-Nano-v0: Decoupled Multi-Vector Late Interaction for Typed Decisions},
  author={Sileo, Damien},
  year={2026},
  howpublished={\url{https://huggingface.co/tasksource/tasksource-jev-nano-v0}},
}

Base model LateOn by LightOn (Apache 2.0). If you use the training data, please also cite tasksource/tasksource-jev-typed-decisions.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tasksource/tasksource-jev-nano-v0

Base model

lightonai/LateOn
Finetuned
(1)
this model