File size: 7,276 Bytes
245a427
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
---
license: other
license_name: dinov3-license
license_link: https://ai.meta.com/resources/models-and-libraries/dinov3-license
base_model: facebook/dinov3-vitb16-pretrain-lvd1689m
pipeline_tag: image-feature-extraction
library_name: codon
tags:
- dino
- dinov3
- vision
- feature-extraction
- fp16
---

# DINOv3 ViT-B/16 — converted weights (codon layout, float16)

This repository hosts the weights of
[`facebook/dinov3-vitb16-pretrain-lvd1689m`](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m)
converted for the `codon` implementation of DINOv3:

- **renamed** to the codon naming convention (`codon.impl.DINOv3ViT`),
- **cast to float16** (2× smaller than the original fp32 checkpoint),
- **numerically verified** against the original weights and against the
  `transformers` reference implementation.

The model itself is unchanged: same architecture, same parameter values, same
numerics up to fp16 quantization.

## Why this exists

The upstream checkpoints use the `transformers`/Meta naming layout
(`layer.{i}.attention.*`, `layer_scale{1,2}.lambda1`, `embeddings.*`) and ship in
fp32. `codon.impl.DINOv3ViT` uses its own convention (all linear layers are
`*_proj`, the block list is `blocks`, Layer Scale is `gamma{1,2}`, register tokens
are `storage_tokens`), so the weights need a one-time remap. This repository is
that remap, pre-applied and stored in fp16 so it can be loaded directly with a
single call.

## Files

| File | Description |
| --- | --- |
| `model.safetensors` | 163.4 MB, 211 tensors, all `F16`, codon key layout |
| `config.json` | Architecture summary plus `"dtype": "float16"` and `"key_layout": "codon"` |
| `preprocessor_config.json` | Image preprocessing parameters, carried over from the original model |
| `README.md` | This file |

## Key mapping

| Original (transformers / Meta) | codon |
| --- | --- |
| `embeddings.patch_embeddings.*` | `patch_embed.*` |
| `embeddings.cls_token` | `cls_token` |
| `embeddings.register_tokens` | `storage_tokens` |
| `layer.{i}.attention.q_proj.*` | `blocks.{i}.q_proj.*` |
| `layer.{i}.attention.k_proj.*` | `blocks.{i}.k_proj.*` |
| `layer.{i}.attention.v_proj.*` | `blocks.{i}.v_proj.*` |
| `layer.{i}.attention.o_proj.*` | `blocks.{i}.o_proj.*` |
| `layer.{i}.norm1.*` / `norm2.*` | `blocks.{i}.norm1.*` / `norm2.*` |
| `layer.{i}.mlp.up_proj.*` / `down_proj.*` | `blocks.{i}.up_proj.*` / `down_proj.*` |
| `layer.{i}.layer_scale1.lambda1` | `blocks.{i}.gamma1` |
| `layer.{i}.layer_scale2.lambda1` | `blocks.{i}.gamma2` |
| `norm.*` | `norm.*` |
| `rope_embeddings.inv_freq` | dropped (rebuilt dynamically from input shape) |

Two notes on the conversion:

- The `inv_freq` rotary table is **not** stored. In this implementation the 2D axial
  RoPE frequencies are recomputed from `rope_theta` and the input resolution at
  runtime (a non-persistent buffer), which is what lets the same weights run at any
  image size.
- A zero-initialized `mask_token` is included as a placeholder. It is part of this
  implementation's MAE architecture but is absent from the original checkpoint;
  `load_pretrained(strict=True)` does not require it.

## Model overview

| Property | Value |
| --- | --- |
| Architecture | DINOv3 ViT (pre-norm Transformer, bidirectional attention, 2D axial RoPE) |
| Parameters | 85.7 M |
| Hidden size | 768 |
| Layers / attention heads | 12 / 12 |
| MLP hidden size | 3072 (ratio 4, GELU) |
| Patch size / default resolution | 16 / 224×224 |
| Register tokens | 4 |
| RoPE theta | 100.0 |
| Layer-norm eps | 1e-5 |
| Layer Scale init | 1.0 |
| Attention biases | q/v/o yes, k no |
| Drop path | 0.0 (inference) |
| Weight dtype | float16 |

## Usage

```python
import torch
from codon.impl import DINOv3ViT_Base

# Pulls config.json + model.safetensors from this repository.
# config.json says dtype=float16, so the model is built in float16.
model = DINOv3ViT_Base.from_remote().eval()

# Prefer fp32 compute? The weights are upcast on load.
model32 = DINOv3ViT_Base.from_remote(dtype=torch.float32).eval()
```

`from_remote()` reads the architecture from `config.json` — `num_register_tokens`,
`rope_theta`, `pos_embed_rescale`, `layer_norm_eps`, `key_bias` and `dtype` are all
applied automatically, so no manual configuration is needed.

To load the file yourself:

```python
model = DINOv3ViT_Base().half()
model.load_pretrained('model.safetensors', strict=True, dtype=torch.float16)
```

### Feature extraction

```python
x = torch.randn(1, 3, 224, 224).half()

with torch.no_grad():
    feats = model.forward_features(x)

feats['x_norm_clstoken']        # [1, 768]        CLS token, post-norm
feats['x_storage_tokens']       # [1, 4, 768]     4 register tokens, post-norm
feats['x_norm_patchtokens']     # [1, 196, 768]   14x14 patch tokens, post-norm
feats['x_norm_alltokens']       # [1, 201, 768]   full post-norm sequence
feats['x_prenorm']              # [1, 201, 768]   full pre-norm sequence
```

`forward(x)` (the default, `is_training=False`) returns just the CLS token of shape
`[N, 768]`.

Dense features at arbitrary resolutions work without interpolation, because the RoPE
frequencies are derived from the actual patch grid:

```python
model.forward_features(torch.randn(1, 3, 256, 192).half())['x_norm_patchtokens'].shape
# torch.Size([1, 192, 768])   -> 16x12 patch grid
```

Intermediate layers, optionally reshaped to feature maps:

```python
with torch.no_grad():
    layers = model.get_intermediate_layers(x, n=[0, 5, 11])          # tuple of [N, HW, C]
    maps = model.get_intermediate_layers(x, n=1, reshape=True)       # [N, C, 14, 14]
```

## Verification

The conversion was checked at every step, and these checks are reproducible via
`test/test_dinov3_vit.py` in the codon repository:

| Check | Result |
| --- | --- |
| Key remap + `strict=True` load from `model.safetensors` | passes |
| fp32 remapped weights vs `transformers` `DINOv3ViTModel`, 224×224 and 196×252 | `max|diff| = 0.000e+00` (bit-exact) |
| Per-block outputs vs `transformers` `output_hidden_states` (layers 0, 5, 11) | `max|diff| < 1e-5` |
| fp16 forward vs original fp32 model (CLS token) | `max|diff| = 6.9e-03` |
| fp16 weight quantization error (per tensor) | `≤ 2.0e-03` |
| Save → reload round trip | `max|diff| = 0.000e+00` |

The fp16 error is well within the expected range for a 12-layer model: activations
reach an order of magnitude of ~10-30, and fp16 carries roughly three significant
decimal digits.
## Precision notes

- Inference in float16 works on CPU and CUDA. The rotary `cos`/`sin` tables are
  computed in fp32 and then **cast to the activation dtype**, so attention inputs
  stay in float16 throughout instead of being silently promoted to fp32.
- If you need maximum fidelity, use `dtype=torch.float32`; the weights are exact
  float16 representations of the original fp32 values, so this only recovers the
  rounding that happened at export time.

## License

The weights are derived from Meta's DINOv3 and remain subject to the
[DINOv3 License](https://ai.meta.com/resources/models-and-libraries/dinov3-license).
Please read and comply with that license before use. In particular, the license
governs acceptable use, redistribution and attribution; this conversion adds no
additional permissions and no warranty.