Spaces:
Running
Running
Stefatorus
Claude Opus 5.5
Why Agate: general image generation as a benchmark-friendly proxy
df091f0 |
Download README.md from Logolabs/agate-webgpu: direct link, hf CLI and curl.
- Browser
- Download file 6.7 kB
-
https://huggingface.co/spaces/Logolabs/agate-webgpu/resolve/main/README.md
- Command line
-
hf download hf://spaces/Logolabs/agate-webgpu/README.md
-
curl -L -o README.md https://huggingface.co/spaces/Logolabs/agate-webgpu/resolve/main/README.md
6.7 kB
| title: Agate WebGPU | |
| colorFrom: red | |
| colorTo: gray | |
| sdk: static | |
| app_file: index.html | |
| pinned: false | |
| license: mit | |
| short_description: Agate text-to-image, 100% in your browser on WebGPU | |
| custom_headers: | |
| cross-origin-embedder-policy: require-corp | |
| cross-origin-opener-policy: same-origin | |
| cross-origin-resource-policy: cross-origin | |
| models: | |
| - Logolabs/agate-preview-003 | |
| - Logolabs/agate-preview-002 | |
| - Logolabs/agate-preview-001 | |
| # Agate WebGPU | |
| LogoLabs' Agate text-to-image models (a ~0.19B flow generator with a 68M text encoder) running as a static page. | |
| The tokenizer, text encoder, 50-step flow sampler with classifier-free guidance and the TAESD decoder all run | |
| on your own GPU through WebGPU (onnxruntime-web). | |
| **Runs entirely in your browser — prompts never leave your machine.** There is no server and no API call. | |
| **Why we built Agate.** Agate is LogoLabs' search for the best architecture for small-scale image generation, as | |
| groundwork for glyph and symbol generation: we want to generate vector-native fonts to go with the wordmarks we make at | |
| LogoLabs. We trained a general text-to-image model because general image generation is a good proxy for how an | |
| architecture performs on those tasks, and it is very benchmark-friendly (GenEval, Qwen-Image-Bench, FID), unlike tasks | |
| such as SDF generation, which have almost no established benchmarks. Next, we plan to build on it to generate SVGs with | |
| our vectoriser, Inkvec, and to improve Inkvec's prior with flow matching to help trace the fonts. We decided to release | |
| Agate to share these architectural efforts. | |
| ## Models | |
| Pick a model at the top of the page. A first visit gets Preview 003; the page remembers your last choice in | |
| this browser. | |
| | model | resolution | notes | download | | |
| |---|---|---|---| | |
| | [Preview 003](https://huggingface.co/Logolabs/agate-preview-003) | 512 px (or 256) | multi-resolution model; prompt pipeline: normaliser, quoted text spelled out, object counts coded, "no X" moved to the negative prompt; SD3 timestep shift 2 at 512 | 531 MB (both resolutions share one weights file) | | |
| | [Preview 002](https://huggingface.co/Logolabs/agate-preview-002) | 256 px | same network as 001, trained longer | 530 MB | | |
| | [Preview 001](https://huggingface.co/Logolabs/agate-preview-001) | 256 px | the first preview | 530 MB | | |
| Each model is downloaded once and kept in the browser's Cache Storage. The files live in each release's model repo | |
| under `webgpu/v2/` (e.g. [Logolabs/agate-preview-003/webgpu/v2](https://huggingface.co/Logolabs/agate-preview-003/tree/main/webgpu/v2)), | |
| with a `manifest.json` holding sizes and sha256s (the cache keys). `models/` in this Space is the original 001 build | |
| (without the plan output), kept for `?models=./models/`. | |
| ## Show thinker | |
| Agate's thinker lays the picture out on a 16 × 16 grid (the plan) that steers the renderer. With **Show thinker** | |
| (on by default; it does not slow sampling: previews are read back asynchronously and coloured in a Web Worker) the page shows, at every step, the plan of the conditional branch as | |
| colours and the image the model currently expects (x₁ = z + (1 − t)·v, with the SD latent→RGB approximation) — | |
| the same live preview as the [ComfyUI nodes](https://github.com/logolabs/agate-comfyui): the plan's cells are | |
| projected on its top three principal components, fitted at the first step with the 2% / 98% quantiles fixed then, | |
| so a colour keeps its meaning while the plan evolves. | |
| ## Requirements | |
| A browser with WebGPU: a recent Chrome or Edge (113+). Safari and Firefox may lack WebGPU or need it enabled in | |
| their settings; without it the page offers a WebAssembly (CPU) path that works but takes minutes per image. | |
| ## Speed and accuracy (RTX 4060 laptop, Edge, WebGPU, warm, 50 steps) | |
| | model | per step (CFG batch of 2) | image | before (2026-09-29 morning) | | |
| |---|---|---|---| | |
| | 003 at 512 px | ~161 ms | **~8.1 s** (Fast, 25 steps: ~4.1 s) | 275 ms, 14 s | | |
| | 003 at 256 px | ~70 ms | ~3.5 s | 120 ms, 6.1 s | | |
| | 002 / 001 | ~65 ms | ~3.3 s | 105 ms, 5.3 s | | |
| How: each step is one graph run that also does the CFG mix and the Euler update, so the latent never leaves the | |
| GPU between steps (the only readback is the final latent); the network computes in **fp16** (WebGPU `shader-f16`; | |
| WGSL has no bf16), with fp32 kept for the graph I/O, the CFG/Euler arithmetic, the timestep / resolution / count | |
| embeddings (sin(1000·t·f) needs fp32) and `Range`; the transposed-conv upsamplers run as 1x1 conv + depth-to-space. | |
| The thinker preview never blocks sampling: its tensors are downloaded asynchronously, coloured in a Web Worker, | |
| and frames are dropped while one is pending (measured cost: none). | |
| fp16 compute moves the result slightly: against the fp32 PyTorch reference with the same initial noise, 001 is | |
| within 16/255 (PSNR 48 dB), 002 36 dB, 003 at 256 px 42 dB and 003 at 512 px 31 dB (mean 0.6-1.8/255). For scale, | |
| the Python package's own CUDA path (bf16) is 14 dB from that fp32 reference for the same noise. Run `?parity=1`. | |
| Seeds use a JavaScript PRNG, so a seed gives the same image in every browser, but not the same image as the | |
| PyTorch pipeline with that seed. The Python packages decode with the full SD-VAE by default; this page uses TAESD. | |
| ## AI-generated content marking | |
| Every image the page makes is marked as AI-generated (EU AI Act Art. 50(2)), with the same marks as the Python | |
| packages: | |
| - an invisible watermark in the pixels: [invisible-watermark](https://github.com/ShieldMnt/invisible-watermark)'s | |
| `dwtDctSvd` method, ported to JavaScript bit for bit, with the 64-bit payload `AGATE` + release | |
| (`AGATE003`, ...). `agate.detect_watermark(img)` from the model packages reads it; | |
| - in the saved PNG: text fields `ai_generated`, `generator`, `model` and `watermark`. The prompt is not written; | |
| - a visible "AI-generated" note next to each result. | |
| Neither mark is tamper-proof: screenshots and re-encodes drop the metadata, and heavy edits can remove the | |
| watermark. If you publish images made here, label them as AI-generated. | |
| Query parameters: `?v=001|002|003` picks a model, `?res=256` picks 003's resolution, `?autoload=1` loads the | |
| model immediately, `?ep=wasm` forces the CPU backend, `?parity=1` runs the parity self-test, `?golden=1` checks | |
| 003's prompt pipeline against the Python package's output (38 prompts). | |
| ## Links | |
| - [logolabs.org](https://logolabs.org) | |
| - [huggingface.co/Logolabs](https://huggingface.co/Logolabs) | |
| Licence: MIT. | |
| We acknowledge EuroHPC JU for awarding the project ID EHPC-AIF-2026PG01-907 access to resources on | |
| Arrhenius GPU at NAISS, Sweden. | |