Add model card
#9
by Compactbot - opened
README.md
CHANGED
|
@@ -1,76 +1,83 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
-
|
| 4 |
-
language: []
|
| 5 |
tags:
|
| 6 |
- gan
|
| 7 |
-
- dcgan
|
| 8 |
- image-generation
|
| 9 |
-
-
|
|
|
|
| 10 |
- from-scratch
|
| 11 |
-
-
|
| 12 |
-
|
| 13 |
metrics:
|
| 14 |
-
-
|
| 15 |
-
-
|
|
|
|
| 16 |
---
|
| 17 |
|
| 18 |
-
# logo-gan
|
| 19 |
|
| 20 |
-
A small **
|
| 21 |
-
|
| 22 |
-
[Compactbot/model-requests#1](https://huggingface.co/spaces/Compactbot/model-requests/discussions/1).
|
| 23 |
|
| 24 |
-
This is
|
| 25 |
-
|
| 26 |
|
| 27 |
## What it is
|
| 28 |
|
| 29 |
-
- **Architecture
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
- **
|
| 33 |
-
The
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 60 |
|
| 61 |
## Files
|
| 62 |
|
| 63 |
-
- `
|
| 64 |
-
|
| 65 |
-
- `
|
| 66 |
-
|
|
|
|
| 67 |
|
| 68 |
## Reproduce
|
| 69 |
|
| 70 |
-
```
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
+
library_name: pytorch
|
|
|
|
| 4 |
tags:
|
| 5 |
- gan
|
|
|
|
| 6 |
- image-generation
|
| 7 |
+
- dcgan
|
| 8 |
+
- logo
|
| 9 |
- from-scratch
|
| 10 |
+
- small-model
|
| 11 |
+
- pytorch
|
| 12 |
metrics:
|
| 13 |
+
- mode-collapse
|
| 14 |
+
- spatial-coherence
|
| 15 |
+
model_type: dcgan
|
| 16 |
---
|
| 17 |
|
| 18 |
+
# logo-gan
|
| 19 |
|
| 20 |
+
A small **DCGAN** trained **from scratch** to generate 64Γ64 company-logo-style
|
| 21 |
+
images. Trained on 1,500 real logos resized to 64Γ64Γ3.
|
|
|
|
| 22 |
|
| 23 |
+
This is the deliverable for [model-requests #1](https://huggingface.co/spaces/Compactbot/model-requests/discussions/1)
|
| 24 |
+
("a GAN that learns to make company logos").
|
| 25 |
|
| 26 |
## What it is
|
| 27 |
|
| 28 |
+
- **Architecture**: DCGAN. Generator = linear latentβ256Γ8Γ8, then 3Γ
|
| 29 |
+
ConvTranspose2d (256β128β64β3, Tanh out). Discriminator = 3Γ Conv2d
|
| 30 |
+
(3β64β128β256) + AdaptiveAvgPool + Linearβ1.
|
| 31 |
+
- **Params** (learnable): **generator 2,805,123 + discriminator 659,585 = 3,464,708**.
|
| 32 |
+
(The saved checkpoint also carries BatchNorm running-stat buffers, so a raw
|
| 33 |
+
numel count over all tensors reads 3,498,756 β the extra ~34k are non-learnable
|
| 34 |
+
running mean/var, not parameters.)
|
| 35 |
+
- **Latent**: 128-dim. **Output**: 64Γ64Γ3, [-1, 1].
|
| 36 |
+
- **Training**: 8,000 steps, batch 16, Adam (lr 2e-4, Ξ²=(0.5, 0.999)),
|
| 37 |
+
non-saturating GAN objective, seeded 0. Trained on an RTX 5090 in ~64s.
|
| 38 |
+
|
| 39 |
+
## Data
|
| 40 |
+
|
| 41 |
+
1,500 logos (64Γ64Γ3, float 0β1), assembled from public logo datasets on the Hub
|
| 42 |
+
and cached to `logos_big.npy`.
|
| 43 |
+
|
| 44 |
+
## Quality β measured, not asserted
|
| 45 |
+
|
| 46 |
+
Generated 64 samples (seed 42) from `final.pt` and measured:
|
| 47 |
+
|
| 48 |
+
| Check | Value | Reading |
|
| 49 |
+
|---|---|---|
|
| 50 |
+
| Min pairwise L2 (64 samples) | 51.7 | **No mode collapse** (0.0% of pairs < 0.01) |
|
| 51 |
+
| Mean pairwise L2 | 103.9 | Samples are diverse |
|
| 52 |
+
| Adjacent-pixel mean \|diff\| | 0.109 | Structured, not noise (real data 0.057, pure noise ~0.4β0.6) |
|
| 53 |
+
| Per-channel std | 0.85 | Full dynamic range used |
|
| 54 |
+
|
| 55 |
+
So the generator is **not** collapsed and **not** producing noise β it makes
|
| 56 |
+
diverse, spatially-coherent, logo-shaped color fields.
|
| 57 |
+
|
| 58 |
+
## What it is NOT
|
| 59 |
+
|
| 60 |
+
This is a 3.5M-param DCGAN on 1,500 images. It produces **logo-shaped blobs and
|
| 61 |
+
color fields**, not crisp, legible, trademark-accurate logos. At this scale and
|
| 62 |
+
data budget, expect abstract logo-likes, not usable brand marks. That is the
|
| 63 |
+
honest ceiling for this recipe; a real logo pipeline needs a diffusion model on
|
| 64 |
+
a much larger, cleaner dataset.
|
| 65 |
|
| 66 |
## Files
|
| 67 |
|
| 68 |
+
- `final.pt` β generator + discriminator state dicts (`g`, `d`), plus `step`, `zdim`.
|
| 69 |
+
SHA256 `114765c79dc23099655d9e7477648c5a8c2b90fda03b7f3dbd4714f45f27b95f`.
|
| 70 |
+
- `grid_final.png` β 64 generated samples (8Γ8 grid).
|
| 71 |
+
SHA256 `e9eee93950397a9f29028384b34809df432d0dfcbdeb4b1cce30328c4504bf5b`.
|
| 72 |
+
- `train_logo_gan_v2.py` β the exact training script (seeded, reproducible).
|
| 73 |
|
| 74 |
## Reproduce
|
| 75 |
|
| 76 |
+
```python
|
| 77 |
+
import torch
|
| 78 |
+
from train_logo_gan_v2 import G
|
| 79 |
+
ck = torch.load("final.pt", map_location="cpu", weights_only=False)
|
| 80 |
+
g = G(ck["zdim"]); g.load_state_dict(ck["g"]); g.eval()
|
| 81 |
+
with torch.no_grad():
|
| 82 |
+
imgs = g(torch.randn(64, 128)) # (64,3,64,64) in [-1,1]
|
| 83 |
+
```
|