agate-webgpu / README.md
Stefatorus
Claude Opus 5.5
Why Agate: general image generation as a benchmark-friendly proxy
df091f0
|
Raw History Blame Contribute Delete
6.7 kB
---
title: Agate WebGPU
colorFrom: red
colorTo: gray
sdk: static
app_file: index.html
pinned: false
license: mit
short_description: Agate text-to-image, 100% in your browser on WebGPU
custom_headers:
cross-origin-embedder-policy: require-corp
cross-origin-opener-policy: same-origin
cross-origin-resource-policy: cross-origin
models:
- Logolabs/agate-preview-003
- Logolabs/agate-preview-002
- Logolabs/agate-preview-001
---
# Agate WebGPU
LogoLabs' Agate text-to-image models (a ~0.19B flow generator with a 68M text encoder) running as a static page.
The tokenizer, text encoder, 50-step flow sampler with classifier-free guidance and the TAESD decoder all run
on your own GPU through WebGPU (onnxruntime-web).
**Runs entirely in your browser — prompts never leave your machine.** There is no server and no API call.
**Why we built Agate.** Agate is LogoLabs' search for the best architecture for small-scale image generation, as
groundwork for glyph and symbol generation: we want to generate vector-native fonts to go with the wordmarks we make at
LogoLabs. We trained a general text-to-image model because general image generation is a good proxy for how an
architecture performs on those tasks, and it is very benchmark-friendly (GenEval, Qwen-Image-Bench, FID), unlike tasks
such as SDF generation, which have almost no established benchmarks. Next, we plan to build on it to generate SVGs with
our vectoriser, Inkvec, and to improve Inkvec's prior with flow matching to help trace the fonts. We decided to release
Agate to share these architectural efforts.
## Models
Pick a model at the top of the page. A first visit gets Preview 003; the page remembers your last choice in
this browser.
| model | resolution | notes | download |
|---|---|---|---|
| [Preview 003](https://huggingface.co/Logolabs/agate-preview-003) | 512 px (or 256) | multi-resolution model; prompt pipeline: normaliser, quoted text spelled out, object counts coded, "no X" moved to the negative prompt; SD3 timestep shift 2 at 512 | 531 MB (both resolutions share one weights file) |
| [Preview 002](https://huggingface.co/Logolabs/agate-preview-002) | 256 px | same network as 001, trained longer | 530 MB |
| [Preview 001](https://huggingface.co/Logolabs/agate-preview-001) | 256 px | the first preview | 530 MB |
Each model is downloaded once and kept in the browser's Cache Storage. The files live in each release's model repo
under `webgpu/v2/` (e.g. [Logolabs/agate-preview-003/webgpu/v2](https://huggingface.co/Logolabs/agate-preview-003/tree/main/webgpu/v2)),
with a `manifest.json` holding sizes and sha256s (the cache keys). `models/` in this Space is the original 001 build
(without the plan output), kept for `?models=./models/`.
## Show thinker
Agate's thinker lays the picture out on a 16 × 16 grid (the plan) that steers the renderer. With **Show thinker**
(on by default; it does not slow sampling: previews are read back asynchronously and coloured in a Web Worker) the page shows, at every step, the plan of the conditional branch as
colours and the image the model currently expects (x₁ = z + (1 − t)·v, with the SD latent→RGB approximation) —
the same live preview as the [ComfyUI nodes](https://github.com/logolabs/agate-comfyui): the plan's cells are
projected on its top three principal components, fitted at the first step with the 2% / 98% quantiles fixed then,
so a colour keeps its meaning while the plan evolves.
## Requirements
A browser with WebGPU: a recent Chrome or Edge (113+). Safari and Firefox may lack WebGPU or need it enabled in
their settings; without it the page offers a WebAssembly (CPU) path that works but takes minutes per image.
## Speed and accuracy (RTX 4060 laptop, Edge, WebGPU, warm, 50 steps)
| model | per step (CFG batch of 2) | image | before (2026-09-29 morning) |
|---|---|---|---|
| 003 at 512 px | ~161 ms | **~8.1 s** (Fast, 25 steps: ~4.1 s) | 275 ms, 14 s |
| 003 at 256 px | ~70 ms | ~3.5 s | 120 ms, 6.1 s |
| 002 / 001 | ~65 ms | ~3.3 s | 105 ms, 5.3 s |
How: each step is one graph run that also does the CFG mix and the Euler update, so the latent never leaves the
GPU between steps (the only readback is the final latent); the network computes in **fp16** (WebGPU `shader-f16`;
WGSL has no bf16), with fp32 kept for the graph I/O, the CFG/Euler arithmetic, the timestep / resolution / count
embeddings (sin(1000·t·f) needs fp32) and `Range`; the transposed-conv upsamplers run as 1x1 conv + depth-to-space.
The thinker preview never blocks sampling: its tensors are downloaded asynchronously, coloured in a Web Worker,
and frames are dropped while one is pending (measured cost: none).
fp16 compute moves the result slightly: against the fp32 PyTorch reference with the same initial noise, 001 is
within 16/255 (PSNR 48 dB), 002 36 dB, 003 at 256 px 42 dB and 003 at 512 px 31 dB (mean 0.6-1.8/255). For scale,
the Python package's own CUDA path (bf16) is 14 dB from that fp32 reference for the same noise. Run `?parity=1`.
Seeds use a JavaScript PRNG, so a seed gives the same image in every browser, but not the same image as the
PyTorch pipeline with that seed. The Python packages decode with the full SD-VAE by default; this page uses TAESD.
## AI-generated content marking
Every image the page makes is marked as AI-generated (EU AI Act Art. 50(2)), with the same marks as the Python
packages:
- an invisible watermark in the pixels: [invisible-watermark](https://github.com/ShieldMnt/invisible-watermark)'s
`dwtDctSvd` method, ported to JavaScript bit for bit, with the 64-bit payload `AGATE` + release
(`AGATE003`, ...). `agate.detect_watermark(img)` from the model packages reads it;
- in the saved PNG: text fields `ai_generated`, `generator`, `model` and `watermark`. The prompt is not written;
- a visible "AI-generated" note next to each result.
Neither mark is tamper-proof: screenshots and re-encodes drop the metadata, and heavy edits can remove the
watermark. If you publish images made here, label them as AI-generated.
Query parameters: `?v=001|002|003` picks a model, `?res=256` picks 003's resolution, `?autoload=1` loads the
model immediately, `?ep=wasm` forces the CPU backend, `?parity=1` runs the parity self-test, `?golden=1` checks
003's prompt pipeline against the Python package's output (38 prompts).
## Links
- [logolabs.org](https://logolabs.org)
- [huggingface.co/Logolabs](https://huggingface.co/Logolabs)
Licence: MIT.
We acknowledge EuroHPC JU for awarding the project ID EHPC-AIF-2026PG01-907 access to resources on
Arrhenius GPU at NAISS, Sweden.