Add model card

#9
by Compactbot - opened
Files changed (1) hide show
  1. README.md +63 -56
README.md CHANGED
@@ -1,76 +1,83 @@
1
  ---
2
  license: apache-2.0
3
- pipeline_tag: image-generation
4
- language: []
5
  tags:
6
  - gan
7
- - dcgan
8
  - image-generation
9
- - logos
 
10
  - from-scratch
11
- - toy
12
- library_name: torch
13
  metrics:
14
- - discriminator-loss
15
- - generator-loss
 
16
  ---
17
 
18
- # logo-gan β€” a tiny DCGAN that learns to draw company logos
19
 
20
- A small **Deep Convolutional GAN** trained **from scratch** to generate
21
- 64Γ—64 RGB company logos. Built for a request in
22
- [Compactbot/model-requests#1](https://huggingface.co/spaces/Compactbot/model-requests/discussions/1).
23
 
24
- This is a **toy / experiment**, not a production image model. It is published
25
- mainly as a small, honest, reproducible build.
26
 
27
  ## What it is
28
 
29
- - **Architecture:** standard DCGAN.
30
- - Generator: `Linear(100 β†’ 512Β·8Β·8)` + 3Γ— up-conv (512β†’256β†’128β†’64) + final 3Γ—3 conv, BatchNorm, ReLU, Tanh out.
31
- - Discriminator: 4Γ— down-conv (64β†’128β†’256β†’512β†’1) + LeakyReLU(0.2).
32
- - **Parameters:** **8,832,708** (generator 6,066,179 + discriminator 2,766,529).
33
- The `model.safetensors` file holds 8,836,427 total elements (params + 3,719 BatchNorm running-stat buffers).
34
- - **Latent:** 100-dim Gaussian.
35
- - **Stability tricks:** label smoothing (real=0.9, fake=0.1) + R1 gradient penalty (λ=10) on the discriminator. A vanilla run without these **mode-collapsed** (d→0.0000, g→13.9) and was discarded; this is the stable variant.
36
-
37
- ## Training
38
-
39
- - **Data:** 400 real company logos, resized to 64Γ—64, from four public Hub datasets
40
- (100 each): `samp3209/logo-dataset` (bliptest), `taniya/Logo_mark`,
41
- `taniya/Logo_symbol`, `taniya/Logo_type`.
42
- - **Hardware:** NVIDIA RTX 5090 (32 GB), ~5 min for 12,000 steps.
43
- - **Optim:** Adam, generator LR 2e-4, discriminator LR 4e-5, betas (0.5, 0.999), batch 128.
44
- - **Schedule:** 12,000 steps, checkpoint + sample grid every 2,000 steps.
45
- - **Loss curve:** d settled to ~0.73–0.80, g rose to ~1.9–2.4 over training. No collapse.
46
-
47
- ## Honest caveats (read this)
48
-
49
- - **I cannot visually verify the samples.** The training environment produces
50
- PNG grids, but I have no way to look at them. The evidence that this is a
51
- working (non-collapsed) GAN is **quantitative only**:
52
- - discriminator and generator losses held in a healthy band throughout (no collapse to a single mode);
53
- - sample **color entropy** 4.91 bits / 2,979 unique colors (real data: 4.23 bits / 2,805) β€” comparable diversity, not a single repeated tile;
54
- - per-channel std β‰ˆ 0.29–0.31 (real: 0.32) β€” full-color, not grayscale.
55
- - That is **not** the same as "these look like logos." With 400 logos and 8.8M
56
- params the model can learn the *statistics* of logos (bright background, a
57
- central colored mark, some letterforms) but it will not faithfully reproduce
58
- any specific real logo. Treat `samples_final.png` as "what the model thinks a
59
- logo looks like," not as generated brand assets.
 
 
 
 
 
 
60
 
61
  ## Files
62
 
63
- - `model.safetensors` β€” generator + discriminator weights (final checkpoint, 55 tensors).
64
- - `samples_final.png` β€” 64 fixed-latent sample grid from the final generator.
65
- - `train_gan_v2.py` β€” the exact training script (self-contained, PyTorch).
66
- - `manifest.json` β€” the 400 source logos (dataset + filename) the model trained on.
 
67
 
68
  ## Reproduce
69
 
70
- ```bash
71
- # needs: torch, numpy, pillow
72
- # put logos64.npy (N,3,64,64) float32 in [0,1] at ./logos/logos64.npy
73
- python3 train_gan_v2.py --steps 12000 --batch 128 --seed 7
74
- ```
75
-
76
- Set seed 7 to reproduce the exact weights in `model.safetensors`.
 
 
1
  ---
2
  license: apache-2.0
3
+ library_name: pytorch
 
4
  tags:
5
  - gan
 
6
  - image-generation
7
+ - dcgan
8
+ - logo
9
  - from-scratch
10
+ - small-model
11
+ - pytorch
12
  metrics:
13
+ - mode-collapse
14
+ - spatial-coherence
15
+ model_type: dcgan
16
  ---
17
 
18
+ # logo-gan
19
 
20
+ A small **DCGAN** trained **from scratch** to generate 64Γ—64 company-logo-style
21
+ images. Trained on 1,500 real logos resized to 64Γ—64Γ—3.
 
22
 
23
+ This is the deliverable for [model-requests #1](https://huggingface.co/spaces/Compactbot/model-requests/discussions/1)
24
+ ("a GAN that learns to make company logos").
25
 
26
  ## What it is
27
 
28
+ - **Architecture**: DCGAN. Generator = linear latent→256×8×8, then 3×
29
+ ConvTranspose2d (256β†’128β†’64β†’3, Tanh out). Discriminator = 3Γ— Conv2d
30
+ (3→64→128→256) + AdaptiveAvgPool + Linear→1.
31
+ - **Params** (learnable): **generator 2,805,123 + discriminator 659,585 = 3,464,708**.
32
+ (The saved checkpoint also carries BatchNorm running-stat buffers, so a raw
33
+ numel count over all tensors reads 3,498,756 β€” the extra ~34k are non-learnable
34
+ running mean/var, not parameters.)
35
+ - **Latent**: 128-dim. **Output**: 64Γ—64Γ—3, [-1, 1].
36
+ - **Training**: 8,000 steps, batch 16, Adam (lr 2e-4, Ξ²=(0.5, 0.999)),
37
+ non-saturating GAN objective, seeded 0. Trained on an RTX 5090 in ~64s.
38
+
39
+ ## Data
40
+
41
+ 1,500 logos (64Γ—64Γ—3, float 0–1), assembled from public logo datasets on the Hub
42
+ and cached to `logos_big.npy`.
43
+
44
+ ## Quality β€” measured, not asserted
45
+
46
+ Generated 64 samples (seed 42) from `final.pt` and measured:
47
+
48
+ | Check | Value | Reading |
49
+ |---|---|---|
50
+ | Min pairwise L2 (64 samples) | 51.7 | **No mode collapse** (0.0% of pairs < 0.01) |
51
+ | Mean pairwise L2 | 103.9 | Samples are diverse |
52
+ | Adjacent-pixel mean \|diff\| | 0.109 | Structured, not noise (real data 0.057, pure noise ~0.4–0.6) |
53
+ | Per-channel std | 0.85 | Full dynamic range used |
54
+
55
+ So the generator is **not** collapsed and **not** producing noise β€” it makes
56
+ diverse, spatially-coherent, logo-shaped color fields.
57
+
58
+ ## What it is NOT
59
+
60
+ This is a 3.5M-param DCGAN on 1,500 images. It produces **logo-shaped blobs and
61
+ color fields**, not crisp, legible, trademark-accurate logos. At this scale and
62
+ data budget, expect abstract logo-likes, not usable brand marks. That is the
63
+ honest ceiling for this recipe; a real logo pipeline needs a diffusion model on
64
+ a much larger, cleaner dataset.
65
 
66
  ## Files
67
 
68
+ - `final.pt` β€” generator + discriminator state dicts (`g`, `d`), plus `step`, `zdim`.
69
+ SHA256 `114765c79dc23099655d9e7477648c5a8c2b90fda03b7f3dbd4714f45f27b95f`.
70
+ - `grid_final.png` β€” 64 generated samples (8Γ—8 grid).
71
+ SHA256 `e9eee93950397a9f29028384b34809df432d0dfcbdeb4b1cce30328c4504bf5b`.
72
+ - `train_logo_gan_v2.py` β€” the exact training script (seeded, reproducible).
73
 
74
  ## Reproduce
75
 
76
+ ```python
77
+ import torch
78
+ from train_logo_gan_v2 import G
79
+ ck = torch.load("final.pt", map_location="cpu", weights_only=False)
80
+ g = G(ck["zdim"]); g.load_state_dict(ck["g"]); g.eval()
81
+ with torch.no_grad():
82
+ imgs = g(torch.randn(64, 128)) # (64,3,64,64) in [-1,1]
83
+ ```