Instructions to use emrevrg/AUBIN-12B-Control with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use emrevrg/AUBIN-12B-Control with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-12B-it") model = PeftModel.from_pretrained(base_model, "emrevrg/AUBIN-12B-Control") - Notebooks
- Google Colab
- Kaggle
AUBIN-12B-Control β AUBIN by Norovox
Real-time control adapter for AubinController: observation (text) + command β one move with a calibrated confidence, one forward pass per move. Part of the AUBIN family (see emrevrg/AUBIN-12B).
With the same safety shield, AUBIN-12B-Control reaches 92% success with 0 lava deaths, vs 89% for the strongest rule baseline (greedy + shield) on the same 100 unseen episodes. An untrained Gemma-4-E4B base without control training reaches 11% with the shield.
Benchmark: command-following grid game (100 episodes, seed 0, never seen in training)
7x7 grid as text, walls, 5 deadly lava cells, 4 objects that block movement, command like "go to the red key".
Step accuracy = the move is on a shortest safe path (BFS ground truth). Training used 3,000 procedurally generated episodes
(seeds 1000+, 15,161 states, soft labels over all shortest moves); the test episodes are disjoint. The safety shield only
removes moves into visible lava/walls/objects (AubinController.act(..., allowed=safe_moves)); the same shield is applied
to the rule baselines so its effect is separated from the model's.
| controller | success | lava deaths β | step accuracy | median ms / move (T4, 4-bit) |
|---|---|---|---|---|
| random | 6% | 40 | 35.3% | β |
| random + safety shield | 13% | 0 | 53.4% | β |
| greedy (ignores obstacles) | 58% | 26 | 56.8% | β |
| greedy + safety shield | 89% | 0 | 87.6% | β |
| AUBIN-12B-Control (one pass per move) | 74% | 24 | 86.4% | 854 |
| AUBIN-12B-Control + safety shield (never steps on visible lava/walls) | 92% | 0 | 91.8% | 856 |
| Gemma-4-E4B base, no control training (one pass per move) | 8% | 26 | 39.2% | 329 |
| Gemma-4-E4B base, no control training + safety shield | 11% | 0 | 55.4% | 330 |
Use
pip install "git+https://huggingface.co/emrevrg/AUBIN-12B#subdirectory=code"
from aubin import Aubin, AubinController
ctl = AubinController(Aubin("emrevrg/AUBIN-12B-Control"),
actions={"up": "move one cell up (row - 1)", "down": "move one cell down (row + 1)",
"left": "move one cell left (column - 1)", "right": "move one cell right (column + 1)"},
instructions="Choose the next move that follows the command along a shortest safe path "
"(never step on lava, objects block movement).")
ctl.reset() # new episode
r = ctl.act(observation, command="go to the red key", allowed={"up", "left"}) # allowed = moves that are safe right now
r["action"], r["confidence"], r["latency_ms"]
Honest notes
- Without the shield the adapter alone is not better than a greedy rule on this game (see table); its value shows when it chooses among safe moves. The shield is a standard action mask, not a planner: it never looks ahead.
- Latency is measured on a free Colab T4 in 4-bit, batch 1. Benchmark code:
control_bench.py, training:control_train.py(in the AUBIN-12B repocode/). - Base model: google/gemma-4-12B-it (Apache-2.0). This repo holds only the LoRA adapter (fp16) and its measurement file.
AUBIN-Learn β instant self-learning (experimental, Norovox core)
AUBIN can learn from feedback without retraining. Verified cases go into an external decision memory (a write takes well under
a millisecond with the hashing embedder), and a self-calibrator (Hedge / multiplicative weights) shifts trust per source between
the model, the memory and their fusion β so where memory is not useful yet, AUBIN keeps trusting itself. Measured on Kev's public
suites; full tables, protocol and code: reports/AUBIN_LEARN.md, code/aubin/learn.py.
- Feedback stream, kev_test: all 6 AUBIN variants improve, +1.0 to +1.3 points (best 84.3 β 85.6). Separate protocol β the label is revealed after each answer β so it is not comparable to static scores (Kev-9B 87.4 static).
- Never-seen sources (transfer, memory starts empty): β0.4 to +0.1 points β it does not hurt; a few hundred feedbacks per source are not enough to help yet.
- Fast skills, static locked test: per-source classifiers learned from memory in seconds (switched on only where dev proves them) lift kev_test for all 4 measured AUBIN variants, +0.3 to +0.6 points (ensemble 84.3 β 84.8; banking77 59.5 β 65.5).
- Raw memory (kNN) on the static test: no reliable gain (β1.0 to +0.7) β AUBIN already learned these sources.
Results, 3 October 2026 (full report: reports/AUBIN_RESULTS_2026-10-03.md)
| benchmark | AUBIN | reference |
|---|---|---|
| Typed decisions (2,000 decisions) | 77.55 (AUBIN-Learn, weights fixed before test) | Laya 76.65 Β· meraGPT 76.8 Β· Jev 72.7 |
| Kev suites, Kev's training sources (kev_test) | 85.7 (AUBIN ensemble, selected on cal split) | Kev-0.8B 83.8 Β· Kev-4B 86.5 Β· Kev-9B 87.4 |
| Kev suites, transfer test | 86.5 ensemble Β· 89.0 AUBIN-31B | β |
| Mind2Web cross-domain step SR (200 steps) | 48.5 AUBIN-31B, no web training | MindAct-XL 39.6 Β· GPT-4 26.4 |
| Grid control, 100 unseen episodes | 92% success, 0 lava deaths (12B-Control + shield) | greedy rule + same shield 89% |
On Kev's own training sources AUBIN is still 1.7 points behind Kev-9B; this is stated, not hidden.
- Downloads last month
- -