Instructions to use Shelter/UI-testjev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Shelter/UI-testjev with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
UI-TestJev — DiffusionGemma for UI Regression Testing
UI-TestJev is an unmerged decoder-only LoRA adapter for fast, structured UI testing decisions with DiffusionGemma. It substantially improves local defect recall on the development corpus, with an accompanying increase in false alarms. It is intended for research and reviewed bug findings.
The adapter is a rank-32, alpha-64 PEFT checkpoint for google/diffusiongemma-26B-A4B-it, pinned to revision f7f5b7f5fa82ffc52addd066915886d497f5517b. The base weights are downloaded separately. This repository includes the exact evaluated inference helper, processor/tokenizer, dependency versions, and aggregate results.
Capabilities and scope
- Local checks of one public requirement and target: contrast, clipping, missing content, occlusion, image aspect, or control dimensions.
- Full-screen reference/current comparison for a single introduced regression category.
- Keyboard reachability decisions using a supplied focus trace, or unknown when the trace is absent.
Outputs are fixed decisions (pass, unknown, or an allowed failure category). No bounding-box localization, generated explanations, or general chat capabilities were trained or evaluated. Full-screen occlusion is not a supported category in this experiment. The full-screen task assumes at most one newly introduced violation category.
Results
Same pinned base, frozen-expert INT8 recipe, BF16 vision, images, prompt protocol, and evaluation noise seeds in both arms.
| Development-test task | Examples | Base macro-F1 | Adapter macro-F1 | Base benign FPR | Adapter benign FPR |
|---|---|---|---|---|---|
| Local requirement check | 1,127 | 0.3132 | 0.7302 | 1.77% | 16.31% |
| Full-screen regression | 787 | 0.2055 | 0.4053 | 0.76% | 3.05% |
| Keyboard reachability | 27 | 0.8911 | 1.0000 | 0% | 0% |
Local defect recall rose from 120/563 (21.31%) to 456/563 (80.99%); local false alarms rose from 10/564 to 92/564. Full-screen defect recall, allowing any failure category, rose from 58/393 to 150/393. Only 115/393 full-screen faults received the correct category. All 12 full-screen image-aspect fault examples were missed by the selected adapter.
The inspected synthetic test family influenced the v3 data repair, so these test results are development diagnostics, not an untouched final holdout. Source groups remain separated between train, validation, and test. Test and Scry scores did not select this checkpoint. A fresh repository-family and natural-bug evaluation is needed for production claims.
The selected checkpoint is epoch 3. All three checkpoints failed one or more predeclared validation guardrails; the export therefore retains validation_recommends_adapter: false. The local validation false-alarm rate rose from 2.07% to 12.03%, above the allowed two-percentage-point increase. Aspect detection, clipping recall retention, per-category false alarms, and the small noise-stability probe also had failures. These results support piloting reviewed findings, but do not establish an unattended build gate.
On 311 external Scry pairs, recall of annotated categories rose from 266/505 (52.67%) to 337/505 (66.73%). Unannotated additional predictions rose from 994 to 1,214. The annotations are incomplete: extras are not certified false positives, and these figures do not establish improved precision or localization accuracy.
See COMPARISON.md, comparison.json, selection.json, and the three validation-epoch-*.json reports. metric-replay.json records exact reproduction of all saved scores from predictions and checksum-verified datasets.
Run inference
Use the included modeling.py helper to preserve the evaluated protocol. Generic chat generation, hosted Inference Providers, and vLLM/OpenJev adapter integration have not been validated for this adapter. The metadata task tag does not establish a ready-to-use hosted endpoint.
Use a separate Linux CUDA environment with Python 3.12/3.13. The training run used Python 3.13, PyTorch 2.10.0+cu128, Transformers 5.11.0, and PEFT 0.21.0. First download or clone all files from this repository into a directory and change into it. Then install:
python -m pip install torch==2.10.0 torchvision==0.25.0 --index-url https://download.pytorch.org/whl/cu128
python -m pip install -r requirements.lock.txt
The loader downloads the pinned original base and converts only the fused expert tensors to the exact frozen INT8 recipe. Other base tensors, including vision weights, remain BF16. Keep the adapter unmerged. Applying it to a different quantizer or BF16 base would be a different, unevaluated configuration. This comparison does not measure quantization loss relative to the BF16 original.
Supply your own approved reference.png and corresponding current.png in a directory, edit example-regression.json to match the actual viewport, and run:
python modeling.py --adapter . --input example-regression.json --image-root /absolute/path/to/your/screenshots
The example JSON is a schema template; screenshots are not bundled. Supply full screenshots in reference/current order. Do not add labels, mutation records, filenames as text evidence, or hidden oracle measurements to the public input. A reference should establish the approved behavior. See INPUT_FORMAT.md for other tasks.
The helper performs one decoder denoising step with fixed answer slots, seed 3407, the original processor's aspect-preserving 1,120 visual-token budget per image, and an 8,192 input-token limit. Inputs exceeding the limit raise an error instead of being truncated. Local examples use approved-reference detail crops in addition to full screens. inference_config.json sets the active image policy to original1120; its alternative candidate-recipe metadata is not active.
semantic_support values describe relative support among the allowed answers, not calibrated probabilities of correctness. The local and full-screen training data contain no visual unknown targets; the selected adapter produced no visual unknown decisions in validation/test. Abstention in ambiguous real inputs remains an evaluation gap.
The measured runtime used an RTX PRO 6000 Blackwell Server Edition, approximately 5.98 session hours before packaging, and 33.49 GiB peak allocated GPU memory. This is not an H100 benchmark. The INT8 workflow targets GPUs with at least 40 GB memory, but actual memory needs depend on inputs and software; only the reported hardware run is measured here. Allow disk space for the original base checkpoint, caches, and adapter.
Dataset overview
We built UI-TestJev Data to answer whether public UI evidence shows a violation of a stated requirement. The final training corpus consists of verified examples rendered from existing open-source interfaces. Matching benign changes teach the model that a visible difference can still satisfy the requirement. The exact four frozen archives, screenshots, labels, source notices, and download instructions are public at Shelter/UI-testjev-data.
| Split | Source families | Local checks | Full-screen comparisons | Keyboard decisions | Total |
|---|---|---|---|---|---|
| Train | Start Bootstrap, Flat UI, Material Dashboard, HTML5 UP | 884 | 622 | 84 | 1,590 |
| Validation | Bulma templates | 483 | 339 | 24 | 846 |
| Development test | AdminLTE | 1,127 | 787 | 27 | 1,941 |
Training covers 50 source pages across four families. Related templates from the same vendor stay in one family; their screenshots, crops, and related questions stay in the same split. Validation covers 29 pages and the development test covers 58. These counts include related observations, not thousands of independent applications.
Scry Design Diff Eval supplies a separate 311-pair external diagnostic across 52 app groups, never training or checkpoint selection. We convert its annotations into fixed category-presence questions and report annotated-category recall; unannotated categories are not assumed clean.
How existing material was processed
- Inspect and preserve sources. Record revisions/download hashes, schemas, licenses, and source-family identities. Render source pages with locally available assets and controlled browser settings.
- Verify references and construct paired states. Check target visibility and the baseline contract. Introduce a controlled failure and a benign counterpart for the same target and requirement. Browser measurements and pixel checks support labels; a mutation name alone is insufficient.
- Separate evidence from labels. Public inputs contain permitted requirements, reference targets, viewport, ordered screenshots, and observed keyboard traces. Gold labels, mutation details, and private oracle measurements stay outside model context. Blind comparisons receive full screens without target-location hints.
- Clean before training. Reject ineffective changes, unexpected pixel effects, unsuitable targets, and ambiguous clipping. Deduplicate public inputs while retaining related sampling groups. Local crops come from the reference and are applied identically to both images. The pinned image processor preserves aspect ratio and enforces token limits.
The v3 audit removed 36 train, 32 validation, and 52 development-test duplicates, quarantined 21 ambiguous or unconfirmed clipping pairs, and checked 6,536 distinct referenced images. Exact image/input hashes and source families do not overlap between splits. These checks reduce known construction errors; semantic correctness and production coverage still require further evaluation.
Gaps found in existing datasets
Existing UI datasets serve different purposes. Their suitability for requirement-based defect detection depends on the evidence and labels they provide:
- Grounding and captions: ScreenParse and RICO Widget Captioning describe elements or locations. They do not supply the required defect verdicts, and their coordinate conventions need explicit conversion.
- Change versus violation: DiffSpot and WUICC provide change evidence or descriptions. These do not by themselves establish a violated product requirement. We did not convert their changes into failure labels.
- Reproduction and structure: Sampled WebSight v0.2 HTML referenced external assets; sampled WebUI records used parallel bounding-box arrays. Re-rendering requires dependency control, and schema conversion must preserve geometry and relationships.
- Evaluation and uncertainty: Benchmark splits need protection from training, and incomplete annotations cannot establish clean negatives. Keyboard behavior requires interaction traces; screenshots alone do not establish reachability.
RICO, ScreenParse, WebSight, WebUI, DiffSpot, WUICC, and UIJudgeBench were not used to train this published adapter. WebSight belonged to an earlier pilot; caption/coordinate tasks investigated during development were excluded from the final decision objective. Scry is the only one of these public evaluation datasets scored in this release. Inspection depth varied and does not certify entire upstream corpora.
Our own corpus retains limitations: controlled template mutations dominate, visual unknown labels are absent, keyboard coverage is small, and subtle/moderate mutations are not calibrated difficulty levels. The inspected AdminLTE test family influenced repairs. Native-app testing and production bug-finding performance remain unproven.
See DATA_PROVENANCE.md for the detailed source table, processing rules, concrete repairs, and remaining gaps; source-inspection-summary.json records inspection scope and revisions.
Training configuration
- 1,590 training, 846 validation, and 1,941 development-test decisions; 311 external Scry pairs.
- Seed 3407; one LoRA configuration, three epochs; validation selected epoch 3 as the highest-scoring fallback.
- Rank 32, alpha 64, dropout 0.05; decoder attention q/k/v/o and dense gate/up/down where present.
- Encoder, vision, fused experts, router, embeddings, and output head frozen; original base weight sharing retained.
- AdamW, learning rate 5e-5, betas (0.9, 0.95), epsilon 1e-8, weight decay 0.01; microbatch 1, accumulation 16, gradient clipping 1.0.
- 5% warmup and cosine decay to 10% of peak learning rate.
- Always-corrupted answer slots and constrained answer-token cross-entropy; shuffled training option order; fixed evaluation order.
- One visit per example per epoch, grouped related cases and interleaved source pages; visual pass-loss weight 1.5, without duplicated minority examples.
- No caption or coordinate-generation supervision. Difficulty is designed but not empirically calibrated.
Download the dataset archives and example images from the separate UI-TestJev Data dataset repository. Archive hashes are in data-audit.json and the dataset's archives.json; source revisions, download hashes, and credits are retained in both releases. The dataset uses component-specific source licenses, as described in its own license file.
Verification and reproducibility
The exported runtime checks passed zero-adapter/base parity, selected/full-vocabulary logit parity, finite adapter-only gradients, a training-only overfit/reset check, cache parity, and adapter save/reload parity. Packaging rechecks the original ZIP and weight hashes. Inference code and weights are byte-for-byte preserved from that export. CPU packaging checks are not an additional GPU inference test.
Adapter weight SHA256: f1cf3e0c00da673bc35ee9c1d0b22aa433c1e3dd68deb862c109d863753e5793.
sha256.json verifies the prepared repository files. release-provenance.json records the original archive hash, file origins, and packaging changes. Runtime paths in original aggregate reports refer to the completed Colab run.
Attribution and license
Base model: Google DiffusionGemma, Apache License 2.0. This adapter and newly supplied helpers/documentation are released under Apache License 2.0; see LICENSE and NOTICE. Retained base processor/tokenizer files remain subject to their upstream license. Dataset/template sources retain their own licenses and credits; this adapter release does not relicense those assets.
The answer-slot protocol follows OpenJev. This is an independent task-specific experiment, not an official Google or OpenJev release or a change to any hosted Jev service.
- Downloads last month
- 16
Model tree for Shelter/UI-testjev
Base model
google/diffusiongemma-26B-A4B-it