ScreenHighlighterRL 4B

A screenshot and a request in; coloured bounding boxes out. The model highlights relevant content while the person stays in control. No clicking, typing, external OCR or inspection tools.

Full merged BF16 weights: final SFT + GRPO optimizer update 200, ready to load without PEFT or a separate adapter.

Get started / Chrome extension · Research article and 300-case viewer

Use

Install the community repository and run python -m screenhighlighter.server --model 4b. It downloads this release and serves the Chrome extension over localhost. See the repository for CUDA setup, a standalone screenshot command, remote-GPU tunnelling and memory options.

For direct Transformers use, load AutoProcessor and Qwen3VLForConditionalGeneration from mustafaah/ScreenHighlighterRL-4B. Use the exact system_prompt.txt, send one image and the instruction, and decode one completion with do_sample=False, max_new_tokens=1024. Use min_pixels=262144, max_pixels=1048576 and preserve image aspect ratio.

Output is {"op":"highlight","targets":[["yellow",x0,y0,x1,y1]]}. Coordinates are normalized to 0–1000 relative to the full original screenshot. Colours: yellow, green, blue, pink, purple, orange, red, white. Empty targets means no matches. White is an opaque mask.

Training and evaluation

800 SFT screenshots, 2,000 GRPO screenshots, 150 validation and 300 test cases; frozen vision tower, rank-32 / alpha-64 LoRA; 200 optimizer updates, 40 requests × 8 candidates/update, four 96 GB RTX PRO 6000 GPUs. Reward is deterministic same-colour one-to-one matching with boundary-tolerant precision/recall F1. Source adapter test reward: 0.617859. This is a geometric reward, not accuracy. These scores precede BF16 merging; no claim of a fresh merged-model benchmark.

Provenance and limits

See release-provenance.json, merge-validation.json and SHA256SUMS.json. BF16 adapter merging can slightly change outputs. The 2B merge was previously validated and is copied from pinned revision 19ae0731432822da915b5992104a2e50fed6f823; 4B merge includes tensor and logit checks. The model can miss or misidentify content. It sees only visible pixels; verify selections before consequential decisions. On-screen content is untrusted input.

Base model weights are Apache-2.0; retain Qwen's attribution. Screenshot examples retain their respective source rights. Author: Mustafa Hussain.

Downloads last month
26
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mustafaah/ScreenHighlighterRL-4B

Finetuned
(457)
this model
Quantizations
2 models