Tokenizer

This model uses the j-llm/j72-tokenizer.

We thank the j-llm project and its contributors for making the tokenizer publicly available.

Action1 Computer-Use Policy

action1 is an experimental lightweight policy model for structured computer-use actions.

The model uses a frozen Action1 backbone as a feature encoder and adds a compact policy network for predicting GUI-oriented actions such as mouse movement, clicking, dragging, typing, keyboard input, scrolling, and task completion.

This repository is intended for research and experimentation with lightweight action policies, supervised imitation learning, and reinforcement learning for computer-use agents.

Important: the current model does not directly observe screenshots, DOM trees, accessibility trees, or OCR output.
It should therefore not be interpreted as a complete vision-based browser or desktop agent.

Overview

The policy predicts one of the following action types:

  • move
  • click
  • drag
  • type
  • press
  • scroll
  • finish

The architecture is designed to remain relatively small while supporting both discrete tool selection and continuous or semi-continuous action parameters.

The current system consists of:

Action1 Backbone
    ↓
Frozen Hidden Representations
    ↓
Attention Pooling
    ↓
Temporal GRU
    ↓
Residual Policy Network
    ↓
Structured GUI Action

The backbone is frozen during policy training and acts as a feature encoder.

This reduces training cost and allows the policy head to specialize in action prediction without updating the full language model.

Architecture

Frozen Backbone

The Action1 backbone produces hidden representations from the model input.

During policy training, the backbone parameters remain frozen.

This provides:

  • lower training memory requirements
  • faster policy optimization
  • more stable reinforcement learning
  • reduced catastrophic forgetting

Attention Pooling

Hidden states are aggregated using an attention-based pooling module.

Instead of relying only on the final token representation, the policy can learn to emphasize hidden states that are more useful for action prediction.

Temporal GRU

A GRU module models temporal dependencies between actions.

This is useful for computer-use tasks because the correct action often depends on the previous interaction history rather than only the current instruction.

Residual Policy Network

The pooled temporal representation is passed to a residual policy network.

Separate output heads predict:

  • tool/action type
  • pointer position
  • drag parameters
  • scroll magnitude
  • keyboard actions
  • text generation behavior
  • finish probability

Action Representation

Click

Click prediction uses a hybrid representation:

coarse grid cell
+
continuous coordinate residual

The policy first predicts a coarse spatial region and then predicts a residual offset inside that region.

Conceptually:

x = x_grid + Δx
y = y_grid + Δy

This formulation is intended to make coordinate learning easier than predicting absolute screen coordinates directly.

Scroll

Scroll prediction similarly combines:

discrete scroll bin
+
continuous residual

This allows the model to represent common scroll magnitudes while retaining fine-grained control.

Typing

Typing behavior is trained using generated text-action examples, including randomized strings.

This is intended to prevent the policy from learning only a small fixed vocabulary of typing actions.

Training

Training is performed in two stages.

Stage 1 — Supervised Fine-Tuning

The policy is first trained with behavior cloning on action demonstrations.

The supervised objective teaches:

  • correct action/tool selection
  • coordinate prediction
  • scroll prediction
  • typing behavior
  • keyboard actions
  • task termination

A simplified objective is:

L_SFT =
    L_tool
  + λ_click L_click
  + λ_scroll L_scroll
  + λ_type L_type
  + ...

Stage 2 — PPO

The supervised policy is then optimized using Proximal Policy Optimization.

The reinforcement learning implementation uses:

  • PPO
  • Generalized Advantage Estimation (GAE)
  • value-function learning
  • entropy regularization
  • gradient clipping
  • KL monitoring

A behavior-cloning loss is retained during PPO training.

The combined objective is approximately:

L =
    L_PPO
  + λ_value L_value
  - λ_entropy H
  + λ_BC L_BC

The behavior-cloning term helps prevent the policy from losing previously learned action-routing behavior during reinforcement learning.

Reward Design

The synthetic training environment can provide rewards for successful action execution and penalties for incorrect behavior.

Examples include:

Wrong-tool penalty

The policy receives a penalty when it selects an incorrect action type.

For example:

expected: click
predicted: scroll

Premature-finish penalty

The policy receives a negative reward when it predicts finish before completing the task.

This reduces the tendency to terminate difficult trajectories early.

Successful execution reward

Correct tool selection and successful execution can receive positive reward.

The exact reward composition may vary between experiments.

Stabilization

Several mechanisms are used to improve PPO stability:

  • frozen backbone
  • behavior cloning during RL
  • gradient clipping
  • KL-divergence monitoring
  • entropy regularization
  • advantage normalization
  • wrong-tool penalties
  • premature-finish penalties

These mechanisms are particularly useful because structured action policies can collapse into repeatedly selecting a small subset of actions during RL.

Benchmark

The current benchmark uses a synthetic action environment.

These results measure policy behavior under the synthetic evaluation setup and should not be interpreted as performance on arbitrary real-world browser or desktop tasks.

In-Distribution Evaluation

Metric Result
Overall success 64.2%
Tool accuracy 100%
Click success 13%
Scroll success 100%

Out-of-Distribution Evaluation

Metric Result
Overall success 27.5%

The difference between in-distribution and out-of-distribution performance indicates that generalization remains a major limitation.

Click localization is currently one of the weakest components of the policy.

Inference Performance

Example inference measurements on an NVIDIA L4 GPU:

Metric Result
Mean latency 13.46 ms / action
Throughput 74.3 actions / second

These values measure policy inference under the tested configuration.

End-to-end computer-use latency would also depend on environment observation, rendering, model preprocessing, and tool execution.

Current Limitations

This model is an experimental policy architecture and has several important limitations.

No visual observation

The current policy does not directly consume:

  • screenshots
  • image patches
  • OCR output
  • DOM trees
  • accessibility trees

As a result, it does not independently perceive arbitrary GUI state.

Synthetic benchmark

Current evaluation is based primarily on a synthetic action environment.

Performance should not be treated as equivalent to success rates on benchmarks such as:

  • BrowserGym
  • WebArena
  • OSWorld
  • AndroidWorld
  • real desktop applications

Weak coordinate generalization

Pointer localization, especially click prediction, remains substantially harder than tool classification.

Distribution shift

Out-of-distribution performance is significantly lower than in-distribution performance.

This suggests that the policy can still overfit to the structure of the training environment.

Not an autonomous computer agent

The model should be considered a structured action-policy component rather than a complete computer-use system.

A full agent would typically require additional components such as:

Environment Observation
        ↓
Vision / DOM / Accessibility Encoder
        ↓
State Representation
        ↓
Planner / Reasoning Model
        ↓
Action Policy
        ↓
Computer Tool
        ↓
Environment

action1 currently focuses primarily on the action-policy portion of this pipeline.

Intended Use

The model may be useful for research involving:

  • structured GUI action prediction
  • lightweight computer-use policies
  • imitation learning
  • PPO for tool-use policies
  • discrete + continuous action spaces
  • policy-head architecture experiments
  • synthetic computer-control environments

Out-of-Scope Use

This model is not intended to be presented as:

  • a production-ready browser agent
  • a general desktop automation system
  • a vision-language GUI agent
  • a benchmark result for real-world computer-use environments
  • a safety-critical autonomous controller

Future Work

Potential future directions include:

Visual State Encoder

Add direct screenshot understanding using a vision encoder.

Screenshot
   ↓
Vision Encoder
   ↓
Visual Tokens
   ↓
Action Policy

DOM / Accessibility Integration

Combine visual features with structured interface information.

Screenshot Features
        +
DOM / Accessibility Features
        +
Instruction Features
        ↓
Multimodal State Representation

Better Pointer Modeling

Improve click and drag localization using techniques such as:

  • multi-scale spatial prediction
  • heatmap-based localization
  • mixture-density heads
  • hierarchical coordinate prediction
  • object-centric pointer prediction

Larger-Scale Environments

Evaluate on more realistic environments such as browser and desktop benchmarks.

Offline RL

Explore offline reinforcement learning using large collections of computer interaction trajectories.

Hierarchical Policies

Separate high-level action planning from low-level motor control.

For example:

High-Level Planner
        ↓
"Open search box"
        ↓
Low-Level Controller
        ↓
move → click → type

Research Status

This project is experimental.

The primary goal is to investigate whether a small structured policy network can learn useful computer-action behavior from frozen language-model representations using a combination of supervised learning and PPO.

The reported numbers should be interpreted as results from the current experimental environment rather than as evidence of general-purpose computer-use capability.

Citation

If you use this model in research, you can cite the repository:

@misc{action1_computer_use_policy,
  title        = {Action1 Computer-Use Policy},
  author       = {summerMC},
  year         = {2026},
  howpublished = {Hugging Face},
  url          = {https://huggingface.co/summerMC/action1}
}

License

See the repository license for usage terms.

Downloads last month
3
Video Preview
loading