JevmaxxxinG β 7-Member Specialist Ensemble (~2.65B total)
Not a single model. Seven fine-tuned encoder decision members β
4Γ ModernBERT-large (421M each) + 3Γ mmBERT-base (322M each), β2.65B parameters total
β each an encoder + 2-layer head (no causal LM, no token generation), plus a
domain router that forwards each request to the relevant specialist.
System-1 decisions with calibrated probabilities, p50 15β25 ms, p95 β€ 100 ms,
5.4 GB bf16 resident.
Load & run
Requires laya (runtime) β each member is a valid laya checkpoint:
pip install laya
from jevmaxxxinG import Ensemble
ens = Ensemble.from_pretrained("chortochki/JevmaxxxinG", device="cuda") # ~5.4 GB bf16, resident
out = ens.decide("Π₯ΠΎΡΡ Π²Π΅ΡΠ½ΡΡΡ ΡΠΎΠ²Π°Ρ", {"intent": {
"type": "choice", "instructions": "intent?",
"criteria": {"return_refund": None, "track_order": None, "general": None}}})
out["answers"]["intent"]["choice"] # 'return_refund'
out["answers"]["intent"]["probabilities"] # {'return_refund': 0.93, ...}
out["answers"]["intent"]["confidence"] # 0.91
out["routing"] # {'member': 'ru_rlcd', 'domain': 'auto'}
decide() / predict() take the exact laya.Agent.predict shape
({qid: {type, instructions, criteria}}) and return {model, answers, usage, routing}.
The router picks the specialist from the request's domain (explicit domain=, Russian
script β ru_rlcd, typed-decisions id-set β td, else the domain default).
ens.load_all() loads all seven at once (β394 s once, then resident).
Benchmarks β common protocols, same protocol for both
The defensible comparison: seven domains where both this ensemble and the official
laya.Router were measured on the identical cases. OpenJev 4B is only comparable where
its published protocol overlaps (JevBench + NLI) β it is not averaged into the mean.
| Domain (held-out) | n | JevmaxxxinG | laya.Router (measured) | OpenJev 4B (pub) |
|---|---|---|---|---|
| JevBench (all) | 1,107 | 0.918 | 0.589 | 0.814 |
| typed-decisions (2,000 decisions) | 2,000 | 0.769 | 0.769 | β |
| NLI (MNLI-matched) | 2,500 | 0.705 | 0.559 | β |
| MASSIVE en-US intent | 2,974 | 0.657 | 0.658 | β |
| MASSIVE ru-RU intent | 2,974 | 0.882 | 0.405 | β |
| FOLIO formal logic | 203 | 0.468 | 0.434 | β |
| model-collapse reasoning | 1,682 | 0.677 | 0.443 | β |
| Mean over these 7 | 0.725 | 0.551 | (not averaged β 2/7 comparable) |
JevBench tiers (p14c member): easy 1.000 / standard 0.986 / hard 0.838 (OpenJev pub: 1.000 / 0.986 / 0.622). NLI detail (p12, openjev protocol): MNLI m/mm 0.843/0.853 (OpenJev 0.896/0.899), ANLI r1/r2/r3 0.521/0.407/0.416 (0.780/0.665/0.627), WANLI 0.653, SciTail 0.837, ConTRoL 0.517 β we trail OpenJev on domain-shifted NLI.
Zero-shot general classification (base member, all-domain-labels protocol)
| Set | options | JevmaxxxinG | Jev 1.13 (pub) |
|---|---|---|---|
| AG News | 4 | 0.938 | 0.950 |
| DAIR Emotion | 6 | 0.585 | 0.595 |
| Banking77 | 77 | 0.393 | 0.425 |
Banking77 (77-way) is the many-option stress test β both models are weak there. The base
member is convaiinnovations/laya base.
Calibration
MNLI-matched: ECE 0.023 (m) / 0.014 (mm), Brier 0.233, NLL 0.426, p50 21 ms.
Order-invariance: 200/200 label-stable, max probability change 0.0 (bit-exact).
Speed & memory
- 5.4 GB bf16 resident (fp32: 10.70 GB) β fits an 8 GB GPU.
p50 15β25 ms,p95 β€ 100 msper request (7-member, GPU-resident).- Not an LLM: no causal head, no KV cache, no continuous batching β vLLM / llama.cpp
do not apply. Runtimes:
laya(in-process), ONNX Runtime, Triton.
Runtimes
- torch /
laya(default):Ensemble.from_pretrained(...)β each member is a validlayacheckpoint (model.safetensors+rl_agent_config.json+tokenizer/). - ONNX Runtime (no torch): each member exported to dynamic-shape ONNX under
onnx/(round-trip max logit diff 4e-6, prob diff 1e-7). - Triton:
jevmaxxxinG/serve_triton.pygenerates a Triton model repo fromonnx/.
Limitations (measured / structural)
- This is an ensemble of per-domain specialists, not a generalist. Each member is fine-tuned on one domain and the router selects by domain signal; the mean above reflects that best-of routing, not a single model's general ability.
- The router is domain-routed, not a learned general router. It keys on the request
domain (typed-decisions id-set, Russian script, explicit
domain=); it was tuned against the benchmark routing signal and has no hold-out routing evaluation on unseen domains. - NLI under domain shift trails OpenJev (ANLI/WANLI/ConTRoL 0.41β0.52 vs 0.63β0.77).
- FOLIO 0.468 is weak for formal logic.
- Languages: English-centric + one Russian member; multilingual is a separate checkpoint, not in this ensemble.
- Response-level factuality (LLM-AggreFact, RAGTruth, HaluBench, IFEval) is judged under a different, LLM-judged protocol than option selection; this ensemble is not evaluated on those sets, so we do not claim a number there.
What's inside
members/{p14c,p12,p8a,td,p16,ru_rlcd,base}/ model.safetensors + rl_agent_config.json + tokenizer/ + encoder/config.json
onnx/{...}/ dynamic-shape ONNX (fp32) per member
router/router.json learned domain router (bench_table + logistic)
jevmaxxxinG/ python package (Ensemble.from_pretrained / decide)
benchmarks.png infographic
Built with laya (convaiinnovations) on top of answerdotai/ModernBERT-large /
jhu-clsp/mmBERT-base. No teacher distillation: cross-entropy + judge-RLCD (opencode-cloud).
Model tree for chortochki/JevmaxxxinG
Base model
answerdotai/ModernBERT-large