File size: 8,755 Bytes
8d31176 a5318e5 8d31176 a5318e5 8d31176 910975d 8d31176 910975d 8d31176 ba0bcdf 8d31176 ba0bcdf 8d31176 910975d 8d31176 910975d 8d31176 910975d 8d31176 ba0bcdf 8d31176 b489f14 8d31176 ba0bcdf 8d31176 ba0bcdf 8d31176 ba0bcdf 910975d ba0bcdf 910975d ba0bcdf 910975d ba0bcdf 910975d ba0bcdf 910975d ba0bcdf 910975d ba0bcdf 910975d ba0bcdf 910975d ba0bcdf 910975d ba0bcdf 910975d ba0bcdf 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d 66873b3 910975d 66873b3 b489f14 910975d ba0bcdf 910975d ba0bcdf 8d31176 ba0bcdf | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 | ---
pipeline_tag: text-to-image
library_name: pytorch
license: apache-2.0
tags:
- solintelligence
- solpix
- flow-matching
- latent-image-generation
- final-release
---
<p align="center">
<img src="banner.png" alt="Sol Labs SolPix" width="100%">
</p>
# SolPix
SolPix turns a text prompt into a 512x512 image. Its roughly 49M-parameter flow transformer predicts image latents; Flan-T5 Base encodes the prompt, and the SANA 1.1 DC-AE decodes the result. Both external models stay frozen.
## Components
| Setting | Value |
|---|---|
| Generator | Approximately 49M parameters |
| Architecture | 15-block U-shaped joint text/image transformer |
| Hidden width | 512 |
| Attention | 8 heads, head dimension 64 |
| FFN | SwiGLU, width 1,152 |
| Long skip connections | 7 |
| Conditioning | Shared adaptive layer normalization |
| Local mixing | Depthwise 3x3 convolution |
| Objective | Rectified flow matching with logit-normal time sampling |
| Image latents | 32 channels, 32x spatial compression |
| 512x512 latent grid | 16x16 |
| Text encoder | Frozen `google/flan-t5-base`, 768 features, up to 96 tokens |
| Decoder | Frozen SANA 1.1 DC-AE F32C32 |
| Saved optimizer step | 210,000 |
| Configured training schedule | 5,000,000 steps |
The parameter count covers the generator only. Downloading the frozen text encoder and autoencoder adds separate dependencies. `SolPixTransformer2D` predicts latent velocity. `AutoencoderDCSol` loads the corresponding Diffusers `AutoencoderDC` for decoding.
## Generate an image
Download the repository, install its requirements, and give `generate.py` a prompt and output path:
```bash
python -m pip install -r requirements.txt
python generate.py --prompt "A glass greenhouse in a quiet garden after rain" --output solpix.png
```
To choose a local checkpoint, seed, or sampling settings:
```bash
python generate.py \
--checkpoint ./step_00210000.pt \
--prompt "A small red sailboat on a misty lake at sunrise" \
--seed 1234 --steps 40 --guidance-scale 3.5 \
--output ./solpix.png
```
On its first run, the helper downloads Flan-T5 Base and the pinned SANA DC-AE revision. It uses Euler integration with classifier-free guidance, running on CUDA when available. CPU inference is supported but slow.
The flow path is `x_t = (1 - t) x_clean + t noise`. Sampling runs from `t=1` down to `t=0`; encoder and decoder identifiers and revision pins are in `config.json`.
## Data and checkpoint history
We used [MONET v1.2.0](https://huggingface.co/datasets/jasperai/monet) with curation seed `20260924`. The split contains 174,603 training examples and a 9,300-example validation holdout. SANA F32C32 image latents and Flan-T5 Base caption states were encoded before training.
MONET draws from CC12M, CommonCatalog-CC-BY, COYO, Diffusion-Aesthetic-4K, and LAION. Flux Klein, Flux Schnell, and Z-Image supply synthetic captions. Curation checks resolution, aesthetics, NSFW content, watermarks, and near duplicates.
The source records include CC BY 4.0, Apache 2.0, Google permissive, and MIT license labels. A label on a record doesn't grant a new license to its contents. Images and dataset shards aren't redistributed in this repository.
The Windows v1.0 continuation used BF16 on one RTX 3080 Ti, with batch size 4 and gradient accumulation 16. We released optimizer step 210,000. The documented 1.1 continuation retains that split and targets step 300,000.
## Reading the samples
We haven't run a formal image-quality or prompt-following benchmark on this checkpoint. The gallery shows generated examples, without supplying a held-out quality estimate. Composition errors, artifacts, and weak text or fine-detail rendering remain limitations.
There is no built-in safety classifier. Dataset filtering doesn't remove every bias or unwanted association. Flan-T5 and SANA DC-AE have separate licenses and usage terms.
All 15 samples below use the released checkpoint at 512x512, with 32 Euler steps, guidance scale 3.5, and the pinned SANA DC-AE decoder. Their files, prompts, seeds, and SHA-256 values are recorded in `samples/`.
### Sample 01

Prompt: Three Black men sharing french fries at a neighborhood diner, candid documentary photography.
Seed: 260926
### Sample 02

Prompt: A red fox standing in fresh snow beneath pine trees at winter dawn, wildlife photography.
Seed: 260927
### Sample 03

Prompt: A glass greenhouse filled with ferns after rain, soft natural light, botanical photograph.
Seed: 260928
### Sample 04

Prompt: A handmade cobalt blue teapot on a pale stone table, clean studio product photograph.
Seed: 260929
### Sample 05

Prompt: A white sailboat crossing a calm blue bay at golden hour, fine art landscape photograph.
Seed: 260930
### Sample 06

Prompt: An orange cat curled on a wooden chair in a sunlit bookshop, cozy editorial photograph.
Seed: 260931
### Sample 07

Prompt: A small street cafe reflected in wet pavement at night, warm window light, city photograph.
Seed: 260932
### Sample 08

Prompt: A wooden lighthouse on a rocky coast under a cloudy sky, atmospheric landscape photograph.
Seed: 260933
### Sample 09

Prompt: A bowl of ripe peaches on a kitchen counter, morning light, natural still life photograph.
Seed: 260934
### Sample 10

Prompt: A snow-covered cabin among tall pine trees at blue hour, quiet winter landscape photograph.
Seed: 260935
### Sample 11

Prompt: A baker placing fresh bread on a cooling rack in a bright kitchen, documentary photograph.
Seed: 260936
### Sample 12

Prompt: A goldfinch perched on a thin branch among spring blossoms, close-up wildlife photograph.
Seed: 260937
### Sample 13

Prompt: A red bicycle leaning against a brick wall on a leafy neighborhood street, lifestyle photograph.
Seed: 260938
### Sample 14

Prompt: A lemon cake with a slice cut out on a ceramic plate, bright tabletop food photograph.
Seed: 260939
### Sample 15

Prompt: A small observatory beneath a clear star-filled sky, distant mountains, night landscape photograph.
Seed: 260940
## Files
- `step_00210000.pt`: EMA and raw weights, optimizer state, configuration, and training arguments.
- `solpix/`: the transformer, decoder adapter, configuration, data, and training components.
- `generate.py`: prompt-to-image generation. `train.py` starts training; `sample_latents.py` samples latents.
- `samples/`: the 15 PNGs and their metadata. `config.json` records the architecture and external-model manifest.
## License
The code, checkpoint weights, configuration, model card, and supplied banner use [Apache 2.0](LICENSE). Attribution is in [NOTICE](NOTICE). Dataset, Flan-T5, and SANA DC-AE licenses apply separately.
|