AUBIN by Norovox

AUBIN-12B β€” AUBIN by Norovox

AUBIN vs Jev on Jev's own split

At a glance

  • Beats Jev on Jev's own published split, all 8 numbers: accuracy 0.863 vs 0.857 and 0.862 vs 0.845; NLL 0.409 vs 0.701; Brier, ECE lower too.
  • AUBIN Trio (adds Phi-4), 8/8 vs Jev: accuracy 0.872 vs 0.857 (new sources) and 0.851 vs 0.845 (trained sources); ECE 0.018 / 0.032 vs 0.049 / 0.067.
  • AUBIN Quad (adds Phi-4 + Mistral-Small; chosen by the rule), 8/8 vs Jev: accuracy 0.869 vs 0.857 (new sources) and 0.856 vs 0.845 (trained sources); ECE 0.030 / 0.022 vs 0.049 / 0.067.
  • 22x fewer confidently-wrong answers (AUBIN-31B, one pass, trained-source split: 0.24% vs Jev 5.22%).
  • 0 errors on held-out policy rules it never trained on (Jev: 14 / 176).
  • Locked test, never-trained sources: 0.898, the best in Kev's table (Kev-27B 0.896).
  • Computer use (Mind2Web), AUBIN-31B with no Mind2Web training beats GPT-4 on 9 of 9 numbers (element acc / op F1 / step success): Cross-Task 48.0 / 81.0 / 40.0 vs 41.6 / 60.6 / 36.2; Cross-Website 41.0 / 82.0 / 33.5 vs 35.8 / 51.1 / 30.1; Cross-Domain 54.0 / 86.5 / 48.5 vs 37.1 / 46.5 / 26.4. GPT-4 in the paper chose from 10 candidates, AUBIN from 50.
  • Also beats the fine-tuned MindAct (Flan-T5-XL, trained on Mind2Web) on all three numbers: Cross-Domain (54.0 / 86.5 / 48.5 vs 42.1 / 66.5 / 39.6).
  • Real-time control (command-following game, 100 unseen episodes, success rate): AUBIN-E4B-Control 95% with the safety shield (0 lava deaths, 385 ms/move on a free T4), 58% with no shield at all; AUBIN-12B-Control 92% with the safety shield (0 lava deaths, 856 ms/move on a free T4), 74% with no shield at all. Rule baselines: greedy 58%, greedy + same shield 89%.
  • AUBIN Loop (generator + decider, same Gemma weights): GSM8K (400 problems) 84.2% β†’ 92.5%; majority vote of the same samples 85.8%.
  • Plays 2048 with a lookahead tool: mean 29,009, 2048 tile in 8/10 games (replay).

Typed questions in, calibrated probabilities out. AUBIN is Norovox's open-base model family: strong open-weight bases, adapted and calibrated for structured decisions (choice / yes-no / score). Fast path = one forward pass; when the model is unsure (top-2 log-prob gap < 4), it writes a short rationale and re-scores (think-when-uncertain, 26-32% of questions). Open weights, runs locally/offline in 4-bit on a single free 16 GB GPU (T4).

Head-to-head with Jev β€” on Jev's own published split

Jev's published numbers (jaredpalmer/kev runs/jev-*-v4/report.json, "clean") are on the development split of Kev's public suites, variant == clean (656 new-source + 1264 trained-source questions). AUBIN is scored on exactly the same questions (code/jev_compare.py). Bold = better than Jev.

model new sources acc Brier ↓ ECE ↓ NLL ↓ trained sources acc Brier ↓ ECE ↓ NLL ↓
Jev (hosted, published report) 0.857 0.2110 0.049 0.701 0.845 0.2367 0.067 0.776
AUBIN-12B full precision (think) 0.857 0.2331 0.050 0.460 0.851 0.2457 0.038 0.623
AUBIN-31B (fast, one pass) 0.852 0.2280 0.068 0.444 0.850 0.2365 0.063 0.567
AUBIN-31B (think) 0.863 0.2255 0.033 0.433 0.857 0.2300 0.040 0.569
AUBIN Duo (31B + 12B, think) 0.867 0.2137 0.021 0.417 0.859 0.2193 0.031 0.554
AUBIN Duo (31B + 12B full precision, think) 0.863 0.2114 0.035 0.411 0.858 0.2243 0.029 0.554
AUBIN Duo-SC (31B self-consistent think + 12B full precision) 0.863 0.2104 0.039 0.409 0.862 0.2227 0.028 0.551
AUBIN Trio (31B-SC + 12B full precision + Phi-4) 0.872 0.2087 0.018 0.406 0.851 0.2315 0.032 0.536
AUBIN Quad (31B-SC + 12B full precision + Phi-4 + Mistral-Small) 0.869 0.2068 0.030 0.404 0.856 0.2268 0.022 0.536

AUBIN Duo-SC (31B self-consistent think + 12B full precision) is better than Jev on all 8 numbers above (accuracy, Brier, ECE, NLL on both splits), on the exact questions Jev's own report uses.

"New sources" = transfer-v4 (MMLU, SciQ, QNLI, PAWS, Emotion, TweetEval, held-out policy rules) β€” never used to train AUBIN. "Trained sources" = decision-v7. AUBIN's training used only decision-v7/train.

Per-source accuracy, AUBIN Duo vs Jev (same clean development questions)
source n AUBIN Duo Jev
composition_held_and_or 32 1.000 0.906
composition_held_conditional 32 1.000 0.781
composition_held_or_not 32 1.000 0.969
contrastive_authorization 40 1.000 1.000
contrastive_deadline 40 1.000 0.925
emotion 80 0.588 0.588
mmlu 80 0.875 0.900
paws 80 0.787 0.787
qnli 80 0.912 0.925
sciq 80 0.988 0.988
tweet_offensive 80 0.738 0.812
agnews 80 0.825 0.812
agnews_yn 160 0.894 0.875
amazon 80 0.613 0.588
banking77 80 0.800 0.812
boolq 80 0.925 0.925
composition_atom 16 1.000 1.000
composition_conditional 16 1.000 0.938
composition_conjunction 16 1.000 1.000
composition_disjunction 16 0.938 0.938
composition_exception 16 1.000 1.000
composition_negation 16 1.000 1.000
composition_nested_and 16 0.812 0.812
composition_nested_or 16 1.000 1.000
contrastive_age_eligibility 24 1.000 1.000
contrastive_quantity_limit 24 1.000 0.958
contrastive_return_window 24 0.958 0.708
contrastive_spend_threshold 24 1.000 1.000
dbpedia14 80 0.963 0.950
imdb 80 0.912 0.900
mnli 80 0.925 0.900
sst5 80 0.600 0.637
trec 80 0.925 0.912
yelp 80 0.662 0.662
yelp_yn 80 0.912 0.863

AUBIN Duo ahead on 15 sources, tied on 15, behind on 5 (of 35).

Where AUBIN wins by multiples (same questions)

The costly failure of a decision model is being confidently wrong or misapplying a policy rule (refund windows, authorization, deadlines, nested conditions).

model split wrong while β‰₯90% confident ↓ (Jev β†’ AUBIN) fewer rule/policy errors ↓ (Jev β†’ AUBIN) fewer
AUBIN-12B full precision (think) new sources 3.66% β†’ 1.68% 2.2Γ— 14 β†’ 1 / 176 14.0Γ—
AUBIN-12B full precision (think) trained sources 5.22% β†’ 1.34% 3.9Γ— 13 β†’ 0 / 224 ∞ (zero)
AUBIN-31B (fast, one pass) new sources 3.66% β†’ 1.07% 3.4Γ— 14 β†’ 7 / 176 2.0Γ—
AUBIN-31B (fast, one pass) trained sources 5.22% β†’ 0.24% 21.8Γ— 13 β†’ 10 / 224 1.3Γ—
AUBIN-31B (think) new sources 3.66% β†’ 3.20% 1.1Γ— 14 β†’ 0 / 176 ∞ (zero)
AUBIN-31B (think) trained sources 5.22% β†’ 2.22% 2.4Γ— 13 β†’ 6 / 224 2.2Γ—
AUBIN Duo (31B + 12B, think) new sources 3.66% β†’ 3.05% 1.2Γ— 14 β†’ 0 / 176 ∞ (zero)
AUBIN Duo (31B + 12B, think) trained sources 5.22% β†’ 2.37% 2.2Γ— 13 β†’ 5 / 224 2.6Γ—
AUBIN Duo (31B + 12B full precision, think) new sources 3.66% β†’ 2.29% 1.6Γ— 14 β†’ 0 / 176 ∞ (zero)
AUBIN Duo (31B + 12B full precision, think) trained sources 5.22% β†’ 1.98% 2.6Γ— 13 β†’ 5 / 224 2.6Γ—
AUBIN Duo-SC (31B self-consistent think + 12B full precision) new sources 3.66% β†’ 2.29% 1.6Γ— 14 β†’ 0 / 176 ∞ (zero)
AUBIN Duo-SC (31B self-consistent think + 12B full precision) trained sources 5.22% β†’ 1.98% 2.6Γ— 13 β†’ 5 / 224 2.6Γ—
AUBIN Trio (31B-SC + 12B full precision + Phi-4) new sources 3.66% β†’ 2.74% 1.3Γ— 14 β†’ 0 / 176 ∞ (zero)
AUBIN Trio (31B-SC + 12B full precision + Phi-4) trained sources 5.22% β†’ 2.69% 1.9Γ— 13 β†’ 4 / 224 3.2Γ—
AUBIN Quad (31B-SC + 12B full precision + Phi-4 + Mistral-Small) new sources 3.66% β†’ 2.29% 1.6Γ— 14 β†’ 0 / 176 ∞ (zero)
AUBIN Quad (31B-SC + 12B full precision + Phi-4 + Mistral-Small) trained sources 5.22% β†’ 2.14% 2.4Γ— 13 β†’ 5 / 224 2.6Γ—

"Wrong while β‰₯90% confident" = share of all questions answered wrongly with confidence β‰₯ 0.9 (Jev: confident_error_rate from its report). AUBIN is more cautious β€” it is β‰₯90% sure on fewer questions β€” but when it is, it is right more often, and its overall calibration (ECE) is better. Rule/policy = composition_* + contrastive_* sources.

Against the Kev family (Kev README, development / test, clean)

On the locked test split of sources AUBIN never trained on, AUBIN-31B (fast, one pass) is the most accurate model in the table (0.898 vs Kev-27B 0.896). On Kev's own trained sources the Kev models remain ahead.

model new sources acc (dev / test) trained sources acc (dev / test) new sources Brier (dev / test)
AUBIN-12B full precision (think) 0.857 / 0.878 0.851 / 0.835 0.233 / 0.199
AUBIN Duo-SC (31B self-consistent think + 12B full precision) 0.863 / 0.895 0.862 / 0.843 0.210 / 0.168
AUBIN Trio (31B-SC + 12B full precision + Phi-4) 0.872 / 0.889 0.851 / 0.851 0.209 / 0.169
AUBIN Quad (31B-SC + 12B full precision + Phi-4 + Mistral-Small) 0.869 / 0.886 0.856 / β€” 0.207 / 0.168
AUBIN Duo (31B + 12B, think) 0.867 / 0.893 0.859 / 0.850 0.214 / 0.169
Kev-0.8B 0.648 / 0.697 0.827 / 0.838 0.481 / 0.416
Kev-4B 0.817 / 0.838 0.873 / 0.865 0.269 / 0.242
Kev-9B 0.822 / 0.852 0.872 / 0.874 0.286 / 0.237
Kev-27B 0.848 / 0.896 0.866 / 0.870 0.236 / 0.164
Jev (hosted) 0.857 / β€” 0.845 / β€” 0.211 / β€”

Computer use: Mind2Web web-agent benchmark

Each step: task + previous actions + the official Mind2Web ranker's top candidates β†’ AUBIN picks the element, the operation (click / type / select) and the value. Metrics as in the Mind2Web paper (Deng et al., 2023): element accuracy, operation F1, step success rate. Reference rows are the paper's Table 2; GPT-4 there chose from the top 10 candidates on a 50-task subset, AUBIN from the top 50.

split model element acc op F1 step SR
Cross-Task AUBIN-31B, zero-shot (paper protocol: 5 + None) (200 steps) 48.0 81.0 40.0
Cross-Task AUBIN-12B + Mind2Web fine-tune (paper protocol: 5 + None) (300 steps) 40.7 82.1 36.3
Cross-Task AUBIN-12B, zero-shot (paper protocol: 5 + None) (300 steps) 35.3 74.3 28.0
Cross-Task AUBIN-12B, zero-shot (single 50-way choice) (300 steps) 38.3 77.8 31.3
Cross-Task GPT-4 (paper, top-10, 50 tasks) 41.6 60.6 36.2
Cross-Task MindAct Flan-T5-XL, fine-tuned (paper) 55.1 75.7 52.0
Cross-Task GPT-3.5 (paper) 20.3 56.6 17.4
Cross-Website AUBIN-31B, zero-shot (paper protocol: 5 + None) (200 steps) 41.0 82.0 33.5
Cross-Website AUBIN-12B, zero-shot (paper protocol: 5 + None) (300 steps) 33.3 76.9 26.0
Cross-Website AUBIN-12B, zero-shot (single 50-way choice) (300 steps) 36.0 79.6 28.3
Cross-Website GPT-4 (paper, top-10, 50 tasks) 35.8 51.1 30.1
Cross-Website MindAct Flan-T5-XL, fine-tuned (paper) 42.0 65.2 38.9
Cross-Website GPT-3.5 (paper) 19.3 48.8 16.2
Cross-Domain AUBIN-31B, zero-shot (paper protocol: 5 + None) (200 steps) 54.0 86.5 48.5
Cross-Domain AUBIN-12B, zero-shot (paper protocol: 5 + None) (300 steps) 42.0 82.9 39.3
Cross-Domain AUBIN-12B, zero-shot (single 50-way choice) (300 steps) 37.3 83.2 34.0
Cross-Domain GPT-4 (paper, top-10, 50 tasks) 37.1 46.5 26.4
Cross-Domain MindAct Flan-T5-XL, fine-tuned (paper) 42.1 66.5 39.6
Cross-Domain GPT-3.5 (paper) 21.6 52.8 18.6

Game: AUBIN plays 2048 (live replay)

AUBIN-12B with a 2-step lookahead tool: the tool plays clear moves, AUBIN decides close calls (34% of moves). Same 10 seeds for every player.

player mean score min max reached 2048
AUBIN + lookahead tool 29,009 7,016 37,132 8/10
lookahead tool alone 30,076 12,268 47,436 8/10
greedy merge rule 2,982 1,464 3,676 0/10
random 933 272 1,508 0/10

Real-time control: command-following grid game

Each move: a 7x7 grid as text (walls, 5 deadly lava cells, 4 objects that block movement) and a command ("go to the red key"); the controller returns one move with a calibrated confidence. Same 100 episodes (seed 0) for every row, never seen in training (training episodes use seeds 1000+). Step accuracy = the move is on a shortest safe path (BFS ground truth). The safety shield only removes moves into visible lava/walls (AubinController.act(..., allowed=...)).

AUBIN-12B-Control rows are measured with code/control_bench.py on a Colab T4; its adapter upload follows in the next update.

controller success lava deaths ↓ step accuracy median ms / move (T4)
random 6% 40 35.3% –
random + safety shield 13% 0 53.4% –
greedy (ignores obstacles) 58% 26 56.8% –
greedy + safety shield 89% 0 87.6% –
Gemma-4-E4B base, no control training (one pass per move) 8% 26 39.2% 329
Gemma-4-E4B base, no control training + safety shield 11% 0 55.4% 330
AUBIN-12B decision adapter, no control training (one pass per move) 57% 27 62.8% 798
AUBIN-12B-Control (one pass per move) 74% 24 86.4% 854
AUBIN-12B-Control + safety shield (never steps on visible lava/walls) 92% 0 91.8% 856
AUBIN-E4B-Control (one pass per move) 58% 28 60.0% 384
AUBIN-E4B-Control + safety shield (never steps on visible lava/walls) 95% 0 93.2% 385

Consistency and hallucination ("gray area" test types)

The test types follow an independent public review of Jev (Onur Tirpan, YouTube, Sep 2026: same question repeated, option order reversed, question rephrased, leading question, X vs not-X, false premise, invented quote, false alarm). AUBIN-12B was run on our own synthetic student records and invoices (Turkish + English, 40 documents); Jev's numbers are the ones shown in that review on its own Turkish items, so this is not a same-item comparison, and the first set of documents is short and simple. The hard version adds distractor fields, synonym labels, look-alike numbers (IBAN/fax), a blank signature line and a long note to every document.

test AUBIN-12B, 40 docs AUBIN-12B, 40 hard docs Jev (review, TR items)
decision changed when asked again 0/120 0/120 1/60
decision changed when options reversed 0/120 1/120 11/60
decision changed when question rephrased 0/120 6/120 5/10
followed a leading question 0/40 0/40 2/10
"X" and "not X" contradicted 0/65 4/65 0/40
accepted a false premise 0/40 0/40 0/60
accepted an invented quote 0/40 0/40 –
said a present field is missing 0/40 0/40 2 in 20 docs
answer accuracy 240/240 234/240 –

On the hard documents AUBIN-12B still contradicted itself on "X" vs "not X" in 4/65 pairs (Jev's review shows 0/40 on its own items); this is the open item for the next training round.

AUBIN Loop: generator + decider (aubin loop)

A generator writes candidate answers; AUBIN picks among them with calibrated probabilities; if confidence is below the target, the generator is asked again with the rejected candidates as feedback. On Gemma it is one set of weights: the generator is the same base with the AUBIN adapter switched off. Any OpenAI-compatible model (ChatGPT, vLLM, Ollama) can be the generator instead (--generator openai:<model>, key from the environment).

GSM8K test, 400 fixed problems, generator = Gemma-4-12B base, same candidates for every method:

method accuracy samples / problem
generator alone (greedy) 0.843 1
majority vote of 5 samples 0.858 5
AUBIN picks among 5 samples 0.900 5
AUBIN Γ— vote share 0.905 5
AUBIN Loop (2nd round only when unsure) 0.925 6.46
oracle: any of 5 samples correct (upper bound) 0.917 5

Measured speed (free Kaggle GPUs, 4-bit, batch = 1 case)

setup fast path think-when-uncertain
AUBIN-12B, 1x T4 486 ms / question (runs/speed/speed12.json) ~6-7 s extra per thought question (batched 8), 26-32% of questions
AUBIN-31B, 2x T4 1.3-2.0 s / question ~13-15 s extra per thought question (batched 8), ~14% of questions
Jev (hosted API, Kev report) ~220 ms median β€”

T4 is a 2018-era free GPU; AUBIN latency on modern datacenter GPUs has not been measured yet and is not claimed.

Use

pip install "git+https://huggingface.co/emrevrg/AUBIN-12B#subdirectory=code"
from aubin import Aubin, AubinEnsemble
m = Aubin("emrevrg/AUBIN-12B")                                   # base + LoRA + fitted temperature + think-when-uncertain
m.decide({"ticket": "Package arrived broken, customer wants money back"},
         {"route": {"type": "choice", "instructions": "Which team?",
                     "criteria": {"billing": "refunds", "shipping": "delivery", "tech": "software"}},
          "urgent": {"type": "noul", "instructions": "Is this urgent?"}})
# AUBIN Duo-SC (headline row; 12B in full precision, e.g. 2x T4):
duo = AubinEnsemble([(Aubin("emrevrg/AUBIN-31B", think_margin=4, think_mix=0.5, think_samples=3), 0.6),
                     (Aubin("emrevrg/AUBIN-12B", four_bit=False, device_map="auto", temperature=3.2, think_margin=4, think_mix=0.5), 0.4)],
                    temperature=0.8)
# AUBIN Loop: generator + decider on the same Gemma weights (or any OpenAI-compatible generator)
from aubin import AubinLoop, GemmaGenerator
from aubin.loop import last_number
loop = AubinLoop(m, GemmaGenerator(m), extract=last_number, target=0.8, mode="aubin+vote")
loop.solve("A shop sells 3 pens for $4. How much do 12 pens cost?")

CLI: aubin decide case.json --think 4 Β· aubin loop "task" --numeric Β· server: aubin serve --port 8009 (POST /decide).

Honest notes

  • Every number above is produced by the scripts in code/ from per-question outputs; nothing is hand-edited.
  • Every ensemble row (Duo, Duo-SC, Trio, Quad) uses one fixed selection rule: member weights and final temperature minimize Brier + ECE on Kev's calibration split (decision-v7 calibration, 300 questions; code/select_cal.py). The development and test splits are never used for selection; think mix is fixed at 0.5. The test columns are single confirmation runs. The new-source Brier margin is small: if the rule were calibration NLL instead, Duo-SC would be behind Jev on new-source Brier and ECE (6/8), while accuracy and NLL stay ahead on both splits.
  • Which members to combine was our choice among about a dozen candidate combinations of the models listed here (including 12B variants trained with distillation and data weighting); only weights and temperature come from the rule.
  • AUBIN Duo-SC: the 31B member thinks with 3 samples (greedy + 2 sampled, answer probabilities averaged; think_samples=3).
  • AUBIN Trio adds microsoft/phi-4 (MIT, no adapter, temperature fitted on the calibration split) as a third, different-family member.
  • AUBIN Quad adds mistralai/Mistral-Small-24B-Instruct-2501 (Apache-2.0, no adapter, calibration-fitted temperature 5.5) as a fourth member; its trained-source test run hit the free-GPU session limit, so that test cell is not reported ("β€”").
  • AUBIN-12B full precision = the same adapter on the unquantized base (fp16 on 2x T4, or bf16 on one T4 through IRIS); its temperature (3.2) is fitted on the calibration split like every other model here.
  • Jev is a hosted API with lower latency (~220 ms median in Kev's report) and very low price; AUBIN's advantages are open weights, local/offline/private use, zero per-call cost on your own GPU, and fine-tunability. Think-when-uncertain adds generation time on ~14% of questions.
  • Base model: google/gemma-4-12B-it (Apache-2.0). This repo holds the LoRA adapter, fitted temperature and code.

AUBIN-Learn β€” instant self-learning (experimental, Norovox core)

AUBIN can learn from feedback without retraining. Verified cases go into an external decision memory (a write takes well under a millisecond with the hashing embedder), and a self-calibrator (Hedge / multiplicative weights) shifts trust per source between the model, the memory and their fusion β€” so where memory is not useful yet, AUBIN keeps trusting itself. Measured on Kev's public suites; full tables, protocol and code: reports/AUBIN_LEARN.md, code/aubin/learn.py.

  • Feedback stream, kev_test: all 6 AUBIN variants improve, +1.0 to +1.3 points (best 84.3 β†’ 85.6). Separate protocol β€” the label is revealed after each answer β€” so it is not comparable to static scores (Kev-9B 87.4 static).
  • Never-seen sources (transfer, memory starts empty): βˆ’0.4 to +0.1 points β€” it does not hurt; a few hundred feedbacks per source are not enough to help yet.
  • Fast skills, static locked test: per-source classifiers learned from memory in seconds (switched on only where dev proves them) lift kev_test for all 4 measured AUBIN variants, +0.3 to +0.6 points (ensemble 84.3 β†’ 84.8; banking77 59.5 β†’ 65.5).
  • Raw memory (kNN) on the static test: no reliable gain (βˆ’1.0 to +0.7) β€” AUBIN already learned these sources.

Results, 3 October 2026 (full report: reports/AUBIN_RESULTS_2026-10-03.md)

benchmark AUBIN reference
Typed decisions (2,000 decisions) 77.55 (AUBIN-Learn, weights fixed before test) Laya 76.65 Β· meraGPT 76.8 Β· Jev 72.7
Kev suites, Kev's training sources (kev_test) 85.7 (AUBIN ensemble, selected on cal split) Kev-0.8B 83.8 Β· Kev-4B 86.5 Β· Kev-9B 87.4
Kev suites, transfer test 86.5 ensemble Β· 89.0 AUBIN-31B –
Mind2Web cross-domain step SR (200 steps) 48.5 AUBIN-31B, no web training MindAct-XL 39.6 Β· GPT-4 26.4
Grid control, 100 unseen episodes 92% success, 0 lava deaths (12B-Control + shield) greedy rule + same shield 89%

On Kev's own training sources AUBIN is still 1.7 points behind Kev-9B; this is stated, not hidden.

Downloads last month
106
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for emrevrg/AUBIN-12B

Adapter
(102)
this model