--- license: cc-by-4.0 library_name: pytorch pipeline_tag: text-to-image tags: - text-to-image - diffusion - vector-graphics - brush-strokes - painting - svg - research datasets: - tomg-group-umd/pixelprose --- # Piccaso-0.1 **A 102M-parameter model that answers a text prompt with 361 brush strokes instead of pixels.** Piccaso started as a curiosity question: *is it easier for a small model to paint a picture than to make one?* Each painting is 361 quadratic Bezier strokes (position, curve, width, colour, on/off), rendered by a fixed renderer to PNG or exported as a real SVG. > **Early research preview.** Trained on 232k images for about 9.2M picture-views (~8 GPU-hours on RTX 5090s). It gets colour, light > and layout right and objects some of the time. It is not a production image generator, and it is nowhere near pixel models trained on > 100M+ images. The learning curves were still rising; more data and training should help a lot. Read the limitations below. ![samples](assets/final_prompts_wide.jpg) ## What it does well, and what it doesn't - **Good:** colour (97% right on our prompt benchmark), light, mood and layout; landscapes, sunsets, interiors, food, vehicles, portraits as painted heads; single main subjects. - **Weak:** two separate objects in one scene (4%), exact shapes, identity, text, fine detail; doodle and emoji styles (out of domain). - **By design:** a painterly, quick-oil-sketch look. The 361-stroke format itself caps realism (see "ceiling" below). ## Results Caption retrieval with CLIP ViT-B/32: does the painting match its own caption better than the other captions of the set? | Top-1 among 200 | Real photo | Fitted 361 strokes (ceiling) | **Piccaso-0.1** | |---|---|---|---| | Seen (training captions) | 96.5% | 83.5% | **28.5%** | | **Unseen (held-out PixelProse)** | 98.0% | 87.5% | **31.0%** | | DOCCI (different photo source, human captions) | 80.0% | n/a | **9.5%** | Chance is 0.5%. Seen and unseen are the same: the model generalises rather than memorising. | StrokeBench, 200 fixed prompts | Score | Chance | |---|---|---| | Right object, top-1 / top-5 of 80 | 30.3% / 56.3% | 1.3% / 6.3% | | Right colour (of 10) | 96.9% | 10% | | Both objects in two-object prompts | 4.4% | 0.4% | | Right style (of 5) | 45.6% | 20% | ![unseen](assets/eval_unseen_wide.jpg) *Unseen captions. Columns: real photo, its fitted 361 strokes (what the format can show), two Piccaso paintings from the caption alone.* ## Usage ```bash git clone https://huggingface.co/shing-dev/Piccaso-0.1 && cd Piccaso-0.1/code pip install torch open_clip_torch ftfy regex safetensors huggingface_hub pillow numpy python paint.py "a lighthouse on a cliff at sunset, oil painting" --n 4 --model .. --out paintings ``` Each painting is saved as a 512 px PNG and an SVG with 361 `` strokes. The Long-CLIP-B text encoder (`BeichenZhang/LongCLIP-B`, ~600 MB) is downloaded on first run. Speed: about 0.11 s per painting on an RTX 5090 (batch 50), about 90 s on a laptop CPU (25 steps). Defaults: 25 DDIM steps, guidance 3. ## Model | | | |---|---| | Output | 361 strokes x 11 numbers: 165 base strokes on 4x4 / 7x7 / 10x10 grids + 196 detail strokes on a 14x14 grid | | Network | set diffusion transformer (DiT-style, adaLN), width 640, 11 layers, 10 heads, 102M parameters | | Text | Long-CLIP-B, frozen: pooled vector + cross-attention to up to 248 caption tokens | | Training | v-prediction, cosine schedule, self-conditioning, 25% reference-image conditioning; 8k steps at batch 256 + 7k at batch 1,024 | | Data | 232,134 PixelProse images (clean, aesthetic >= 5), converted to strokes by gradient-descent fitting, filtered by how well the stroke render still matches the caption | | Files | `model.safetensors` (EMA weights + stroke normalisation, anchors, caption standardisation), `config.json`, `code/` | ## How it compares to a real image model MobileDiffusion (Google, 2023) trains on 150M images with weeks of TPU time and runs in ~0.2 s on a phone. Piccaso-0.1 saw 232k images (about 650x fewer) for ~8 GPU-hours, and the whole project cost about $65 of rented GPUs. The gap is mostly data and compute, not the stroke format. Few-step distillation and more data are the obvious next steps. ## Limitations and responsible use - Research preview; outputs are often wrong or abstract. Do not use it where accuracy matters. - Trained on web images with machine-written captions (PixelProse); it inherits their biases and gaps. - It cannot reproduce identities, logos or text in any recognisable way at this scale. ## Links and credits - Code, full lab notebook and report: [github.com/shing1Sks/piccaso](https://github.com/shing1Sks/piccaso) - The research film (1:46): [piccaso-research-journey.mp4](https://github.com/shing1Sks/piccaso/releases/download/v0.1/piccaso-research-journey.mp4) - Data: [PixelProse](https://huggingface.co/datasets/tomg-group-umd/pixelprose) (captions CC-BY-4.0); [DOCCI](https://huggingface.co/datasets/google/docci) for testing only - Text encoder: [Long-CLIP](https://github.com/beichenzbc/Long-CLIP) (Apache-2.0, code vendored in `code/longclip/`) - Research by Shreyash Kumar Singh, run with an AI coding agent (Claude). Weights: CC-BY-4.0.