Instructions to use ByunByun/laya-android with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ByunByun/laya-android with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="ByunByun/laya-android")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ByunByun/laya-android", device_map="auto") - Laya
How to use ByunByun/laya-android with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("ByunByun/laya-android", device_map="auto")Laya-Android
A 322M typed-decision policy for Android GUI control, fine-tuned from Laya on AndroidControl. Accessibility tree only · no vision · no autoregressive generation.
Code, docs and full results on GitHub
Highlights
| Parameters | AndroidControl-High Type | AndroidControl-High Grounding | Latency (RTX 4090, bs1) |
|---|---|---|---|
| 322M | 76.7 | 61.7 | ~40 ms |
Every step of a recorded validation episode with the model's predictions (blue = predicted, green = gold). The model never sees the screenshots; it reads the accessibility tree. Gold history, not a live run.
Usage
# pip install laya==0.3.28
import laya
agent = laya.load("ByunByun/laya-android")
OP_DESC = {"CLICK": "tap a UI element", "LONG_PRESS": "long-press a UI element", "SCROLL_UP": "scroll up",
"SCROLL_DOWN": "scroll down", "SCROLL_LEFT": "scroll left", "SCROLL_RIGHT": "scroll right",
"OPEN_APP": "launch an app by name", "INPUT_TEXT": "type text into the focused field",
"BACK": "press the system back button", "HOME": "go to the home screen",
"WAIT": "wait for the screen to update", "DONE": "the goal is complete; stop"}
goal = "Turn off Bluetooth"
history = ["OPEN_APP Settings"] # last 3 actions
candidates = ["Network & internet (LinearLayout)", "Connected devices (LinearLayout)", "Apps (LinearLayout)"]
state = {"goal": goal, "history": history, "ui": ["[%d] %s" % (i, c) for i, c in enumerate(candidates)]}
questions = {
"operation": {"type": "choice", "instructions": "Which operation should be performed next?", "criteria": OP_DESC},
"target": {"type": "choice", "instructions": "Which UI element should be targeted?",
"criteria": {str(i): c for i, c in enumerate(candidates)}},
}
out = agent.system_one(state, questions, max_len=3072, head_max_len=1536)
print(out["answers"]["operation"]["choice"], out["answers"]["target"]["choice"])
Candidate labels must follow the repository's src/candidates.py / src/serialization.py ("<text> (<class>) <flags>").
Compared with Jev (same API, same 1,000 test steps)
| Laya-Android (322M, local) | Jev 1.13.0 (TypeSafe API) | |
|---|---|---|
| Type | 77.2 | 67.2 |
| Grounding | 62.8 | 56.7 |
| Operation ECE (lower is better) | 0.037 | 0.066 |
| Operation Brier (lower is better) | 0.338 | 0.487 |
| Target ECE (lower is better) | 0.061 | 0.140 |
| Median latency | 39 ms (RTX 4090) | 236 ms (API round trip) |
Byte-identical requests to both models. Jev 1.13.0 via the TypeSafe API (2026-10-07), a general decision model not trained on AndroidControl; its latency is an API round trip.
Compared with generative models
| Model | Params | Serving | Type | Grounding | Median latency | p95 | vs Laya-Android |
|---|---|---|---|---|---|---|---|
| Laya-Android | 322M | Laya-Android (local) | 77.2 | 62.8 | 39 ms | 48 ms | 1x |
| Qwen3.8 27B | 27B (Q4_K_M) | Ollama GGUF | 76.8 | 64.5 | 575 ms | 694 ms | 14.9x slower |
| Qwen3.5 4.2B | 4.2B (Q4_K_M) | Ollama GGUF | 67.3 | 52.8 | 193 ms | 252 ms | 5.0x slower |
| jev-1.13.0 | undisclosed | TypeSafe API | 67.2 | 56.7 | 236 ms | 295 ms | 6.1x slower |
| Qwen3.5 4B | 4B (bf16) | vLLM bf16 | 67.1 | 50.5 | 242 ms | 305 ms | 6.3x slower |
| Gemma 4 25B | 25B (Q4_K_M) | Ollama GGUF | 64.8 | 56.6 | 201 ms | 256 ms | 5.2x slower |
| Qwen3.5 9.7B | 9.7B (Q4_K_M) | Ollama GGUF | 63.6 | 51.6 | 229 ms | 273 ms | 5.9x slower |
| Gemma 4 E2B | E2B (bf16) | vLLM bf16 | 58.4 | 23.6 | 155 ms | 210 ms | 4.0x slower |
| Qwen3.5 2B | 2B (bf16) | vLLM bf16 | 57.5 | 23.0 | 149 ms | 202 ms | 3.9x slower |
| Gemma 4 8.0B | 8.0B (Q4_K_M) | Ollama GGUF | 57.4 | 45.0 | 171 ms | 202 ms | 4.4x slower |
| Qwen3.5 1.9B | 1.9B (Q8_0) | Ollama GGUF | 57.4 | 24.5 | 126 ms | 144 ms | 3.3x slower |
| Gemma 4 E4B | 7.5B (Q4_K_M) | Ollama GGUF | 57.1 | 44.1 | 146 ms | 175 ms | 3.8x slower |
| Gemma 4 E2B | 4.6B (Q4_K_M) | Ollama GGUF | 56.9 | 22.7 | 116 ms | 139 ms | 3.0x slower |
| Ministral 3 14B | 14B (Q4_K_M) | Ollama GGUF | 53.2 | 42.0 | 285 ms | 424 ms | 7.4x slower |
| Qwen3.5 0.8B | 0.8B (Q8_0) | Ollama GGUF | 41.3 | 0.0 | 124 ms | 158 ms | 3.2x slower |
| Qwen3.5 0.8B | 0.8B (bf16) | vLLM bf16 | 40.9 | 0.4 | 125 ms | 178 ms | 3.2x slower |
Same 1,000 test steps (seed 0), same input (goal, last 3 actions, the same accessibility candidates), same InfiGUI-R1 judge, one RTX 4090 at batch 1. The generative models are used zero-shot (not trained on AndroidControl), thinking disabled, asked for a JSON action; vLLM runs them in bf16, Ollama in 4-bit GGUF. Jev is TypeSafe's hosted decision model (jev-1.13.0, measured 2026-10-07) and its latency is an API round trip. The comparison shows the speed of answering by selection versus generation; the accuracy gap also reflects that only Laya-Android was trained on this dataset.
Full test set (official evaluator)
AndroidControl-High, 8,444 test steps, public InfiGUI-R1 evaluator. Same model before and after fine-tuning:
| Model | Params | Type | Grounding | Full SR |
|---|---|---|---|---|
| Laya base (zero-shot) | 322M | 29.0 | 11.8 | N/A |
| Laya-Android | 322M | 76.7 | 61.7 | N/A |
| Δ | +47.6 | +49.9 |
Full SR is N/A (v0 does not generate text or app names). Screenshot VLMs on the same evaluator report higher numbers (e.g. InfiGUI-R1-3B 82.7 / 74.4) but solve a different task (pixel coordinates vs. choosing accessibility candidates).
One test pass of the frozen checkpoint; protocol fixed before scoring (eval_protocol.md).
Model details
- Base:
convaiinnovations/laya/multilingual@7b928d8(mmBERT-base encoder), ~322M parameters, Apache-2.0 - Input: goal, last 3 actions, accessibility-derived UI candidates;
max_len3072,head_max_len1536 - Output: probabilities over 12 operations and over the candidates
- Developed by Jaeyeon Byun (GitHub: Byun11); not affiliated with Convai Innovations
Training
AndroidControl train split only (13,594 episodes, 74,714 actions). Supervised fine-tuning with laya.train, bf16, 3 epochs on one RTX 4090; checkpoint and temperatures chosen on validation. No RL, no oversampling, no screenshots. Details: training.md.
Limitations
- No visual input: unlabeled icons and custom-drawn UI are hard to tell apart.
- No text or app-name generation, so no Full Step SR; AndroidControl-Low not evaluated.
- Policy Joint drops 9–12 points on unseen apps, tasks and categories.
- Weak on rare actions (WAIT, BACK, horizontal scroll, LONG_PRESS).
- Offline, step-level benchmark with gold history. Do not let it operate real devices unsupervised; gate irreversible actions and keep on-screen personal data local.
Citation
@software{laya_android_2026,
title = {Laya-Android},
author = {Byun, Jaeyeon},
year = {2026},
version = {0.2.1},
url = {https://github.com/Byun11/laya-android}
}
Built on Laya (Convai Innovations), mmBERT-base and AndroidControl (Google Research). Evaluator: InfiGUI-R1.
Model tree for ByunByun/laya-android
Base model
convaiinnovations/layaEvaluation results
- Type on AndroidControl-High (InfiGUI-R1 evaluator, 8,444 test steps)test set self-reported76.700
- Grounding on AndroidControl-High (InfiGUI-R1 evaluator, 8,444 test steps)test set self-reported61.700



# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="ByunByun/laya-android")