Instructions to use emrevrg/AUBIN-E4B-Screen with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use emrevrg/AUBIN-E4B-Screen with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-E4B-it") model = PeftModel.from_pretrained(base_model, "emrevrg/AUBIN-E4B-Screen") - Notebooks
- Google Colab
- Kaggle
Open, calibrated decision models that see, act and learn.
Typed decisions · computer use · real-time control · self-learning, open weights on Gemma 4.
Omni · 31B · 12B · E4B · Screen · Web · Control · Support ❤
♥ If AUBIN impresses you, press Like at the top of this page. It is the single biggest help for an independent open model.
🎬 AUBIN in 80 seconds (sound on) · Türkçe izle · vertical reel: EN / TR
Music: “Aphelion” by Scott Buckley, released under CC BY 4.0 · scottbuckley.com.au
AUBIN-E4B-Screen: fast visual computer use
You give it a screenshot and an instruction such as "open settings". It returns the point to click as (x, y), on a 0-1000 scale. The model is Gemma-4-E4B with a LoRA adapter, running about 1.9 s per screenshot on a free T4 in 4-bit.
Part of the AUBIN by Norovox family: AUBIN-12B · AUBIN-31B · full results: reports/AUBIN_RESULTS_2026-10-03.md
ScreenSpot (1,272 screenshots: mobile, desktop and web; click-point accuracy)
| model | accuracy |
|---|---|
| AUBIN-E4B-Screen (round 2) | 69.3 |
| AUBIN-E4B-Screen, round 1 | 67.7 |
| AUBIN-E4B, before this training | 48.5 |
| SeeClick | 53.4 |
| CogAgent | 47.4 |
| UGround-7B | 73.3 |
| OS-Atlas-7B | 82.5 |
| UI-TARS-7B | 89.5 |
Two training rounds so far: 16,000 examples, about 30% of the usable data. The stronger dedicated grounding models above are still ahead.
Training data and leakage
The adapter was trained on agentsea/wave-ui, train split. To avoid leakage, these sources were removed entirely:
screenspotagent_studio(GroundUI, which contains ScreenSpot)mind2web_test_*
The training target is the centre of the element's box.
Use
from transformers import AutoProcessor, AutoModelForImageTextToText
from peft import PeftModel
proc = AutoProcessor.from_pretrained("google/gemma-4-E4B-it")
m = PeftModel.from_pretrained(AutoModelForImageTextToText.from_pretrained("google/gemma-4-E4B-it", device_map="auto"),
"emrevrg/AUBIN-E4B-Screen")
prompt = ('You are a GUI agent. In this screenshot, where should I click to: "open settings"?\n'
"Answer with ONLY the click point as (x, y), where x and y are integers from 0 to 1000 "
"(0,0 = top-left corner, 1000,1000 = bottom-right corner).")
msgs = [{"role": "user", "content": [{"type": "image", "image": screenshot}, {"type": "text", "text": prompt}]}]
inp = proc.apply_chat_template(msgs, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt").to(m.device)
print(proc.decode(m.generate(**inp, max_new_tokens=16)[0, inp["input_ids"].shape[1]:], skip_special_tokens=True))
Training and evaluation code: code/screenspot_train.py and code/screenspot_eval.py in the AUBIN-12B repository.
AUBIN-Learn and the latest family results
AUBIN-Learn — instant self-learning (experimental, Norovox core)
AUBIN can learn from feedback without retraining. Verified cases go into an external decision memory (a write takes well under
a millisecond with the hashing embedder), and a self-calibrator (Hedge / multiplicative weights) shifts trust per source between
the model, the memory and their fusion — so where memory is not useful yet, AUBIN keeps trusting itself. Measured on Kev's public
suites; full tables, protocol and code: reports/AUBIN_LEARN.md, code/aubin/learn.py.
- Feedback stream, kev_test: all 6 AUBIN variants improve, +1.0 to +1.3 points (best 84.3 → 85.6). Separate protocol — the label is revealed after each answer — so it is not comparable to static scores (Kev-9B 87.4 static).
- Never-seen sources (transfer, memory starts empty): −0.4 to +0.1 points — it does not hurt; a few hundred feedbacks per source are not enough to help yet.
- Fast skills, static locked test: per-source classifiers learned from memory in seconds (switched on only where dev proves them) lift kev_test for all 4 measured AUBIN variants, +0.3 to +0.6 points (ensemble 84.3 → 84.8; banking77 59.5 → 65.5).
- Raw memory (kNN) on the static test: no reliable gain (−1.0 to +0.7) — AUBIN already learned these sources.
Results, 3 October 2026 (full report: reports/AUBIN_RESULTS_2026-10-03.md)
| benchmark | AUBIN | reference |
|---|---|---|
| Typed decisions (2,000 decisions) | 77.55 (AUBIN-Learn, weights fixed before test) | Laya 76.65 · meraGPT 76.8 · Jev 72.7 |
| Kev suites, Kev's training sources (kev_test) | 85.7 (AUBIN ensemble, selected on cal split) | Kev-0.8B 83.8 · Kev-4B 86.5 · Kev-9B 87.4 |
| Kev suites, transfer test | 86.5 ensemble · 89.0 AUBIN-31B | – |
| Mind2Web cross-domain step SR (200 steps) | 48.5 AUBIN-31B, no web training | MindAct-XL 39.6 · GPT-4 26.4 |
| Grid control, 100 unseen episodes | 92% success, 0 lava deaths (12B-Control + shield) | greedy rule + same shield 89% |
On Kev's own training sources AUBIN is still 1.7 points behind Kev-9B; this is stated, not hidden.
Evening update (3 Oct, all measured, details in the results report):
| benchmark | AUBIN | reference |
|---|---|---|
| ScreenSpot click accuracy (visual computer use) | 69.3 AUBIN-E4B-Screen (zero-shot 48.5; 67.7 after round 1) | SeeClick 53.4 · CogAgent 47.4 · UGround-7B 73.3 · UI-TARS-7B 89.5 |
| Mind2Web step success, cross-task / website / domain | 47.0 / 37.0 / 43.5 (12B) · 47.0 / 34.5 / 42.0 (E4B, ≈2× faster) | MindAct-XL 52.0 / 38.9 / 39.6 |
| ViZDoom FPS, kills per episode (30 episodes) | 17.3 with in-game self-learning (AUBIN-Learn) · 15.6 without (AUBIN-12B, 0.6 s/move) | random 1.3 · scripted rule 18.8 |
| Kev training sources + learned skills (selected on kev_dev + cal only) | 87.08, NLL 0.43 (ensemble 85.69, NLL 0.79) | Kev-9B 87.4 (not yet beaten) · Kev-27B 87.0 · Kev-4B 86.5 |
New in code/: AubinLearning.acquire_skill (learns a skill, self-tests on held-out data, enables only on proven gain), learn_skill2/3.py, fuse_multi.py, screenspot_train/eval.py, fps_vizdoom.py.
The AUBIN family: every ability, one interface (AubinEngine routes between them)
| ability | model | measured |
|---|---|---|
| typed decisions (choice / score / yes-no), calibrated | AUBIN-31B · AUBIN-12B · 12B-v3b · 12B-v3d · E4B-v3 | typed-decisions 77.55 (#1) · Jev's set 8/8 · Kev unseen sources 89.8 · Kev training sources 87.08 |
| web agent | AUBIN-12B-Web · AUBIN-E4B-Web (fast) · 31B web training running | Mind2Web cross-domain 43.5 (MindAct-XL 39.6) |
| visual computer use (click on screenshots) | AUBIN-E4B-Screen · 12B screen training running | ScreenSpot 69.3 |
| real-time control | AUBIN-12B-Control · AUBIN-E4B-Control | 92% success, 0 lava deaths |
| FPS play, self-learning in game | AUBIN-12B + AUBIN-Learn | ViZDoom 17.3 kills/episode vs 15.6 without learning (30 episodes) |
| self-learning, self-acquired skills | AubinLearning (learn, acquire_skill) in every repo |
skills switch on only after a held-out self-test proves a gain |
🎬 AUBIN filmi, Türkçe (80 sn, sesi aç) · English · Müzik: “Aphelion”, Scott Buckley, CC BY 4.0
♥ Help AUBIN get seen
On Hugging Face, likes decide what people discover. Big labs have marketing teams; AUBIN has one 17-year-old student and measured results. If AUBIN is useful, interesting or just impressive to you, press ♥ Like at the top of this page and on AUBIN-Omni, AUBIN-31B and AUBIN-12B, then share it with one person who builds with AI. Every like helps an independent, open, honestly measured model get discovered.
🇹🇷 Beğenin, AUBIN'in görünür olmasını sağlar: sayfanın üstündeki ♥ Like'a basarak destek ol ve bir arkadaşına gönder.
Support Norovox
Built by a 17-year-old high school student: no sponsor, no budget, just free GPUs and AI subscriptions paid for with difficulty. Support goes into GPU compute, training, and the AI development tools this work depends on (such as Claude); supporters are credited and get early access. zgremre@gmail.com · emrevrgdev@gmail.com · Why and how →
🇹🇷 17 yaşında bir lise öğrencisinin eseri: destekçisiz, bütçesiz; ücretsiz GPU'lar ve zorlukla ödenen yapay zekâ abonelikleriyle. Desteğin GPU'ya, eğitime ve bu işin dayandığı yapay zekâ geliştirme araçlarına (Claude gibi) gider.
Built with Claude Opus 5.5 and GPT-5.6 Sol; because of OpenAI usage limits, the final stretch was completed with Claude Opus 5.5. License: Apache-2.0 (adapters and code). Base models: Google Gemma 4 (Apache-2.0).
- Downloads last month
- 30