Encoders
Collection
My encoders from scratch • 2 items • Updated
Vision encoder for the Escarda family (~39.6M, 448px / patch16 / 784 tokens). Byrne-VE plus a JEPA head next to the HRM refine block - that's the Escarda trait.
| Byrne-VE | Escarda-VE | |
|---|---|---|
| Params | 39.34M | 39.60M (+JEPA head) |
| CLS cosine | 0.776 | 0.771 |
| PATCH cosine | 0.600 | 0.584 |
| JEPA self-consistency | - | 0.040 |
Escarda trades ~1-3% teacher-alignment (JEPA pulls a bit of capacity off pure mimicry) for a self-supervised spatial neighbour-prediction signal that costs nothing at inference. Same size class as Byrne-VE.