Instructions to use interfaze-ai/lev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use interfaze-ai/lev with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B") model = PeftModel.from_pretrained(base_model, "interfaze-ai/lev") - Notebooks
- Google Colab
- Kaggle
Independent run: Banking77 12-way 0.7292 against your 77-intent 0.980
Question about your Banking77 and aegis2 figures.
Setup. Revision 7bdc748dffeb, one sealed manifest, seed 42, 1,240 scored decisions across 9 typed-decision suites, fp32 on CPU.
What we measured (overall 0.5613): Banking77 12-way 0.7292, our moderation suite 0.8056, emotion 0.3073, MNLI 0.3917.
Two things we cannot reconcile.
Your card reports banking77 at 77 intents reaching 0.980 after skipping codes that tokenize to two tokens. Our 12-way subset reaches only 0.7292. Since fewer intents should be easier, our number should be the upper bound, not the lower one. That points at our harness under-reporting on your model rather than over-reporting.
Your aegis2 safety-moderation figure is 0.804; our moderation suite gives 0.8056. Those agree closely, which is what made the Banking77 gap stand out rather than look like a systematically pessimistic harness.
What would help. How you select among 77 intents, and whether skipping codes that tokenize to two tokens is applied at inference or only during training. Our per_option_conditional_logprob path with a declared option list may differ from your label-token readout in a way that only shows up at high option counts.
Full data and per-row predictions: https://sysone.sdad.pro/
Thank you for your questions @saidutta69 , Two separate things are stacked on top of each other -
Our 0.980 is an in-distribution number and the card doesn't say so where it matters. banking77 is one of lev's 29 training sources, it is not in our contamination blocklist, which covers only the 13 S1Bench evaluation subsets and the 0.980 is measured on 400 held-out rows of our own training mixture: same taxonomy, same question rendering, same label codes. It belongs in the "held-out split of the 29 training sources" table,
which does carry that caveat, but we also quote it as a bare "banking77 0.818 → 0.980" bullet in the optimizations list, and that reads like a general banking77 capability claim. It isn't one.
Second, and probably the larger effect: your per_option_conditional_logprob path is a different function of the same weights, not a different implementation of the same one.
How we select among 77 intents. Each option is assigned a short label code that encodes to exactly one token on this tokenizer. The prompt lists code: description lines, and the answer is the argmax over the next-token logits restricted to those code tokens, from a single forward pass. Lettered choices are rendered in two option orders and averaged, which cancels letter-position
bias.
This is a better answer than we expected, and it corrects us on both counts.
On the 0.980 - our comparison was invalid. You are right that the number is in-distribution:400 held-out rows from your own training mixture, with banking77 among the 29 training sources. We measured a 12-way subset on our sealed manifest, which your model has not been trained on. So we were comparing in-distribution against out-of-distribution and reading the gap as a harness fault. It was not evidence of one. Thanks for saying plainly that the optimizations-list bullet does not carry the caveat the table does - that is a documentation problem on your side and we will not press it.
On the scoring function - this is the part we want to act on. "A different function of the same weights" is the useful framing, and we had not applied it to our own results before you put it that way.
Six of our 43 measurements use our per_option_conditional_logprob path, where each declared option is scored independently and compared. Those six average 0.5543; the other 37 average 0.7031. Two of our three lowest measurements are in that group - LFM2.5-350M-RLCD at 0.3669 and LFM2.5-2.6B-RLCD at 0.3855. Our lowest overall, pngwn at 0.2734, uses a scalar-head softmax readout, which raises the same question.
So we cannot currently distinguish "these models are weak at typed decisions" from "our readout is the wrong shape for models trained to answer with restricted label codes." Your description - single forward pass, argmax restricted to one-token codes, two option orders averaged - is a function our harness does not implement at all.
The ask. Would you run our sealed manifest through your code-token readout, so we can compare function against function on byte-identical inputs?
- Manifest and harness: https://github.com/instax-dutta/sysone-bench
- 1,190 cases, 1,240 scored evaluation decisions, seed 42, manifest sha256
a938cc2483a592dc84e0d5baac12594491bcaa5b4ceb6b7c3b0def71b36297bd - Our per-row predictions and per-suite numbers for lev: https://sysone.sdad.pro/
- Your lev checkpoint is pinned at
7bdc748dffebon our side, fp32 on CPU
If your readout scores our manifest meaningfully higher than our 0.5613, that is a finding about our harness and we will publish it, credit you, and correct the affected rows. If it lands near ours, we will say that too and leave the numbers as they are. Either result is useful; the current disagreement is not.
We are also adding a caveat to the six affected rows on our site in the meantime, because as published they are not qualified and should be.