Tokenizer
This model uses the j-llm/j72-tokenizer.
We thank the j-llm project and its contributors for making the tokenizer publicly available.
Action1 Computer-Use Policy
action1 is an experimental lightweight policy model for structured computer-use actions.
The model uses a frozen Action1 backbone as a feature encoder and adds a compact policy network for predicting GUI-oriented actions such as mouse movement, clicking, dragging, typing, keyboard input, scrolling, and task completion.
This repository is intended for research and experimentation with lightweight action policies, supervised imitation learning, and reinforcement learning for computer-use agents.
Important: the current model does not directly observe screenshots, DOM trees, accessibility trees, or OCR output.
It should therefore not be interpreted as a complete vision-based browser or desktop agent.
Overview
The policy predicts one of the following action types:
moveclickdragtypepressscrollfinish
The architecture is designed to remain relatively small while supporting both discrete tool selection and continuous or semi-continuous action parameters.
The current system consists of:
Action1 Backbone
↓
Frozen Hidden Representations
↓
Attention Pooling
↓
Temporal GRU
↓
Residual Policy Network
↓
Structured GUI Action
The backbone is frozen during policy training and acts as a feature encoder.
This reduces training cost and allows the policy head to specialize in action prediction without updating the full language model.
Architecture
Frozen Backbone
The Action1 backbone produces hidden representations from the model input.
During policy training, the backbone parameters remain frozen.
This provides:
- lower training memory requirements
- faster policy optimization
- more stable reinforcement learning
- reduced catastrophic forgetting
Attention Pooling
Hidden states are aggregated using an attention-based pooling module.
Instead of relying only on the final token representation, the policy can learn to emphasize hidden states that are more useful for action prediction.
Temporal GRU
A GRU module models temporal dependencies between actions.
This is useful for computer-use tasks because the correct action often depends on the previous interaction history rather than only the current instruction.
Residual Policy Network
The pooled temporal representation is passed to a residual policy network.
Separate output heads predict:
- tool/action type
- pointer position
- drag parameters
- scroll magnitude
- keyboard actions
- text generation behavior
- finish probability
Action Representation
Click
Click prediction uses a hybrid representation:
coarse grid cell
+
continuous coordinate residual
The policy first predicts a coarse spatial region and then predicts a residual offset inside that region.
Conceptually:
x = x_grid + Δx
y = y_grid + Δy
This formulation is intended to make coordinate learning easier than predicting absolute screen coordinates directly.
Scroll
Scroll prediction similarly combines:
discrete scroll bin
+
continuous residual
This allows the model to represent common scroll magnitudes while retaining fine-grained control.
Typing
Typing behavior is trained using generated text-action examples, including randomized strings.
This is intended to prevent the policy from learning only a small fixed vocabulary of typing actions.
Training
Training is performed in two stages.
Stage 1 — Supervised Fine-Tuning
The policy is first trained with behavior cloning on action demonstrations.
The supervised objective teaches:
- correct action/tool selection
- coordinate prediction
- scroll prediction
- typing behavior
- keyboard actions
- task termination
A simplified objective is:
L_SFT =
L_tool
+ λ_click L_click
+ λ_scroll L_scroll
+ λ_type L_type
+ ...
Stage 2 — PPO
The supervised policy is then optimized using Proximal Policy Optimization.
The reinforcement learning implementation uses:
- PPO
- Generalized Advantage Estimation (GAE)
- value-function learning
- entropy regularization
- gradient clipping
- KL monitoring
A behavior-cloning loss is retained during PPO training.
The combined objective is approximately:
L =
L_PPO
+ λ_value L_value
- λ_entropy H
+ λ_BC L_BC
The behavior-cloning term helps prevent the policy from losing previously learned action-routing behavior during reinforcement learning.
Reward Design
The synthetic training environment can provide rewards for successful action execution and penalties for incorrect behavior.
Examples include:
Wrong-tool penalty
The policy receives a penalty when it selects an incorrect action type.
For example:
expected: click
predicted: scroll
Premature-finish penalty
The policy receives a negative reward when it predicts finish before completing the task.
This reduces the tendency to terminate difficult trajectories early.
Successful execution reward
Correct tool selection and successful execution can receive positive reward.
The exact reward composition may vary between experiments.
Stabilization
Several mechanisms are used to improve PPO stability:
- frozen backbone
- behavior cloning during RL
- gradient clipping
- KL-divergence monitoring
- entropy regularization
- advantage normalization
- wrong-tool penalties
- premature-finish penalties
These mechanisms are particularly useful because structured action policies can collapse into repeatedly selecting a small subset of actions during RL.
Benchmark
The current benchmark uses a synthetic action environment.
These results measure policy behavior under the synthetic evaluation setup and should not be interpreted as performance on arbitrary real-world browser or desktop tasks.
In-Distribution Evaluation
| Metric | Result |
|---|---|
| Overall success | 64.2% |
| Tool accuracy | 100% |
| Click success | 13% |
| Scroll success | 100% |
Out-of-Distribution Evaluation
| Metric | Result |
|---|---|
| Overall success | 27.5% |
The difference between in-distribution and out-of-distribution performance indicates that generalization remains a major limitation.
Click localization is currently one of the weakest components of the policy.
Inference Performance
Example inference measurements on an NVIDIA L4 GPU:
| Metric | Result |
|---|---|
| Mean latency | 13.46 ms / action |
| Throughput | 74.3 actions / second |
These values measure policy inference under the tested configuration.
End-to-end computer-use latency would also depend on environment observation, rendering, model preprocessing, and tool execution.
Current Limitations
This model is an experimental policy architecture and has several important limitations.
No visual observation
The current policy does not directly consume:
- screenshots
- image patches
- OCR output
- DOM trees
- accessibility trees
As a result, it does not independently perceive arbitrary GUI state.
Synthetic benchmark
Current evaluation is based primarily on a synthetic action environment.
Performance should not be treated as equivalent to success rates on benchmarks such as:
- BrowserGym
- WebArena
- OSWorld
- AndroidWorld
- real desktop applications
Weak coordinate generalization
Pointer localization, especially click prediction, remains substantially harder than tool classification.
Distribution shift
Out-of-distribution performance is significantly lower than in-distribution performance.
This suggests that the policy can still overfit to the structure of the training environment.
Not an autonomous computer agent
The model should be considered a structured action-policy component rather than a complete computer-use system.
A full agent would typically require additional components such as:
Environment Observation
↓
Vision / DOM / Accessibility Encoder
↓
State Representation
↓
Planner / Reasoning Model
↓
Action Policy
↓
Computer Tool
↓
Environment
action1 currently focuses primarily on the action-policy portion of this pipeline.
Intended Use
The model may be useful for research involving:
- structured GUI action prediction
- lightweight computer-use policies
- imitation learning
- PPO for tool-use policies
- discrete + continuous action spaces
- policy-head architecture experiments
- synthetic computer-control environments
Out-of-Scope Use
This model is not intended to be presented as:
- a production-ready browser agent
- a general desktop automation system
- a vision-language GUI agent
- a benchmark result for real-world computer-use environments
- a safety-critical autonomous controller
Future Work
Potential future directions include:
Visual State Encoder
Add direct screenshot understanding using a vision encoder.
Screenshot
↓
Vision Encoder
↓
Visual Tokens
↓
Action Policy
DOM / Accessibility Integration
Combine visual features with structured interface information.
Screenshot Features
+
DOM / Accessibility Features
+
Instruction Features
↓
Multimodal State Representation
Better Pointer Modeling
Improve click and drag localization using techniques such as:
- multi-scale spatial prediction
- heatmap-based localization
- mixture-density heads
- hierarchical coordinate prediction
- object-centric pointer prediction
Larger-Scale Environments
Evaluate on more realistic environments such as browser and desktop benchmarks.
Offline RL
Explore offline reinforcement learning using large collections of computer interaction trajectories.
Hierarchical Policies
Separate high-level action planning from low-level motor control.
For example:
High-Level Planner
↓
"Open search box"
↓
Low-Level Controller
↓
move → click → type
Research Status
This project is experimental.
The primary goal is to investigate whether a small structured policy network can learn useful computer-action behavior from frozen language-model representations using a combination of supervised learning and PPO.
The reported numbers should be interpreted as results from the current experimental environment rather than as evidence of general-purpose computer-use capability.
Citation
If you use this model in research, you can cite the repository:
@misc{action1_computer_use_policy,
title = {Action1 Computer-Use Policy},
author = {summerMC},
year = {2026},
howpublished = {Hugging Face},
url = {https://huggingface.co/summerMC/action1}
}
License
See the repository license for usage terms.
- Downloads last month
- 3