Decision Index: scoring AutoTrust JEV-9B and JEV-27B

#1
by AutoTrustAILab - opened
AutoTrust AI Lab org

Hi Poli @multimodalart
cc: @cloudyu
I'm Josh Liu, co-founder of AutoTrust AI in Singapore. Thanks for transferring JEV-98 demo to us. Also congratulations on building the Jev Decision Index which is the clearest head-to-head view of the open Jev reproductions we've found.

We'd love to see how our two Apache-2.0 open-weight Jev students score on your frozen suite:

Both were distilled from TypeSafe Jev 1.13's full output distributions (SargeDev/jev-distill-corpus-v3) by training a LoRA plus a 24-slot decision head. On the 29,955-question held-out set, AutoTrust JEV-27B picks the same option as Jev 1.13 on 90.3% of choice questions (mean KL 0.019). In our re-run of gazelle93's decision-models-under-pressure, it scores 0.740 at 16 options against Jev 1.13's published 0.769.

Following your note in the Space discussions, we'll self-score with apolinario/decision-index: run the full frozen suite, publish the run directories as a Hub dataset, and open a PR adding both models to submissions/README.md. Two things to flag up front:

  1. Our server exposes /v1/decisions rather than /v1/systemone, so the PR will include a small Engine subclass to make the run reproducible.
  2. The choice head takes at most 16 options. Questions with more options, and inputs beyond our configured context length, will be declared Unsupported rather than truncated.

Does that approach work for you? In the meantime, could both models go on the News tab under Trained? Happy to send that PR as well.

Best,
Josh Liu
Chairman & Co-Founder, AutoTrust AI

Sign up or log in to comment