logo-gan / README.md
Compactbot's picture
Add model card (#9)
fc127f3
|
Raw
History Blame Contribute Delete
2.98 kB
metadata
license: apache-2.0
library_name: pytorch
tags:
  - gan
  - image-generation
  - dcgan
  - logo
  - from-scratch
  - small-model
  - pytorch
metrics:
  - mode-collapse
  - spatial-coherence
model_type: dcgan

logo-gan

A small DCGAN trained from scratch to generate 64×64 company-logo-style images. Trained on 1,500 real logos resized to 64×64×3.

This is the deliverable for model-requests #1 ("a GAN that learns to make company logos").

What it is

  • Architecture: DCGAN. Generator = linear latent→256×8×8, then 3× ConvTranspose2d (256→128→64→3, Tanh out). Discriminator = 3× Conv2d (3→64→128→256) + AdaptiveAvgPool + Linear→1.
  • Params (learnable): generator 2,805,123 + discriminator 659,585 = 3,464,708. (The saved checkpoint also carries BatchNorm running-stat buffers, so a raw numel count over all tensors reads 3,498,756 — the extra ~34k are non-learnable running mean/var, not parameters.)
  • Latent: 128-dim. Output: 64×64×3, [-1, 1].
  • Training: 8,000 steps, batch 16, Adam (lr 2e-4, β=(0.5, 0.999)), non-saturating GAN objective, seeded 0. Trained on an RTX 5090 in ~64s.

Data

1,500 logos (64×64×3, float 0–1), assembled from public logo datasets on the Hub and cached to logos_big.npy.

Quality — measured, not asserted

Generated 64 samples (seed 42) from final.pt and measured:

Check Value Reading
Min pairwise L2 (64 samples) 51.7 No mode collapse (0.0% of pairs < 0.01)
Mean pairwise L2 103.9 Samples are diverse
Adjacent-pixel mean |diff| 0.109 Structured, not noise (real data 0.057, pure noise ~0.4–0.6)
Per-channel std 0.85 Full dynamic range used

So the generator is not collapsed and not producing noise — it makes diverse, spatially-coherent, logo-shaped color fields.

What it is NOT

This is a 3.5M-param DCGAN on 1,500 images. It produces logo-shaped blobs and color fields, not crisp, legible, trademark-accurate logos. At this scale and data budget, expect abstract logo-likes, not usable brand marks. That is the honest ceiling for this recipe; a real logo pipeline needs a diffusion model on a much larger, cleaner dataset.

Files

  • final.pt — generator + discriminator state dicts (g, d), plus step, zdim. SHA256 114765c79dc23099655d9e7477648c5a8c2b90fda03b7f3dbd4714f45f27b95f.
  • grid_final.png — 64 generated samples (8×8 grid). SHA256 e9eee93950397a9f29028384b34809df432d0dfcbdeb4b1cce30328c4504bf5b.
  • train_logo_gan_v2.py — the exact training script (seeded, reproducible).

Reproduce

import torch
from train_logo_gan_v2 import G
ck = torch.load("final.pt", map_location="cpu", weights_only=False)
g = G(ck["zdim"]); g.load_state_dict(ck["g"]); g.eval()
with torch.no_grad():
    imgs = g(torch.randn(64, 128))   # (64,3,64,64) in [-1,1]