File size: 7,276 Bytes
245a427 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 | ---
license: other
license_name: dinov3-license
license_link: https://ai.meta.com/resources/models-and-libraries/dinov3-license
base_model: facebook/dinov3-vitb16-pretrain-lvd1689m
pipeline_tag: image-feature-extraction
library_name: codon
tags:
- dino
- dinov3
- vision
- feature-extraction
- fp16
---
# DINOv3 ViT-B/16 — converted weights (codon layout, float16)
This repository hosts the weights of
[`facebook/dinov3-vitb16-pretrain-lvd1689m`](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m)
converted for the `codon` implementation of DINOv3:
- **renamed** to the codon naming convention (`codon.impl.DINOv3ViT`),
- **cast to float16** (2× smaller than the original fp32 checkpoint),
- **numerically verified** against the original weights and against the
`transformers` reference implementation.
The model itself is unchanged: same architecture, same parameter values, same
numerics up to fp16 quantization.
## Why this exists
The upstream checkpoints use the `transformers`/Meta naming layout
(`layer.{i}.attention.*`, `layer_scale{1,2}.lambda1`, `embeddings.*`) and ship in
fp32. `codon.impl.DINOv3ViT` uses its own convention (all linear layers are
`*_proj`, the block list is `blocks`, Layer Scale is `gamma{1,2}`, register tokens
are `storage_tokens`), so the weights need a one-time remap. This repository is
that remap, pre-applied and stored in fp16 so it can be loaded directly with a
single call.
## Files
| File | Description |
| --- | --- |
| `model.safetensors` | 163.4 MB, 211 tensors, all `F16`, codon key layout |
| `config.json` | Architecture summary plus `"dtype": "float16"` and `"key_layout": "codon"` |
| `preprocessor_config.json` | Image preprocessing parameters, carried over from the original model |
| `README.md` | This file |
## Key mapping
| Original (transformers / Meta) | codon |
| --- | --- |
| `embeddings.patch_embeddings.*` | `patch_embed.*` |
| `embeddings.cls_token` | `cls_token` |
| `embeddings.register_tokens` | `storage_tokens` |
| `layer.{i}.attention.q_proj.*` | `blocks.{i}.q_proj.*` |
| `layer.{i}.attention.k_proj.*` | `blocks.{i}.k_proj.*` |
| `layer.{i}.attention.v_proj.*` | `blocks.{i}.v_proj.*` |
| `layer.{i}.attention.o_proj.*` | `blocks.{i}.o_proj.*` |
| `layer.{i}.norm1.*` / `norm2.*` | `blocks.{i}.norm1.*` / `norm2.*` |
| `layer.{i}.mlp.up_proj.*` / `down_proj.*` | `blocks.{i}.up_proj.*` / `down_proj.*` |
| `layer.{i}.layer_scale1.lambda1` | `blocks.{i}.gamma1` |
| `layer.{i}.layer_scale2.lambda1` | `blocks.{i}.gamma2` |
| `norm.*` | `norm.*` |
| `rope_embeddings.inv_freq` | dropped (rebuilt dynamically from input shape) |
Two notes on the conversion:
- The `inv_freq` rotary table is **not** stored. In this implementation the 2D axial
RoPE frequencies are recomputed from `rope_theta` and the input resolution at
runtime (a non-persistent buffer), which is what lets the same weights run at any
image size.
- A zero-initialized `mask_token` is included as a placeholder. It is part of this
implementation's MAE architecture but is absent from the original checkpoint;
`load_pretrained(strict=True)` does not require it.
## Model overview
| Property | Value |
| --- | --- |
| Architecture | DINOv3 ViT (pre-norm Transformer, bidirectional attention, 2D axial RoPE) |
| Parameters | 85.7 M |
| Hidden size | 768 |
| Layers / attention heads | 12 / 12 |
| MLP hidden size | 3072 (ratio 4, GELU) |
| Patch size / default resolution | 16 / 224×224 |
| Register tokens | 4 |
| RoPE theta | 100.0 |
| Layer-norm eps | 1e-5 |
| Layer Scale init | 1.0 |
| Attention biases | q/v/o yes, k no |
| Drop path | 0.0 (inference) |
| Weight dtype | float16 |
## Usage
```python
import torch
from codon.impl import DINOv3ViT_Base
# Pulls config.json + model.safetensors from this repository.
# config.json says dtype=float16, so the model is built in float16.
model = DINOv3ViT_Base.from_remote().eval()
# Prefer fp32 compute? The weights are upcast on load.
model32 = DINOv3ViT_Base.from_remote(dtype=torch.float32).eval()
```
`from_remote()` reads the architecture from `config.json` — `num_register_tokens`,
`rope_theta`, `pos_embed_rescale`, `layer_norm_eps`, `key_bias` and `dtype` are all
applied automatically, so no manual configuration is needed.
To load the file yourself:
```python
model = DINOv3ViT_Base().half()
model.load_pretrained('model.safetensors', strict=True, dtype=torch.float16)
```
### Feature extraction
```python
x = torch.randn(1, 3, 224, 224).half()
with torch.no_grad():
feats = model.forward_features(x)
feats['x_norm_clstoken'] # [1, 768] CLS token, post-norm
feats['x_storage_tokens'] # [1, 4, 768] 4 register tokens, post-norm
feats['x_norm_patchtokens'] # [1, 196, 768] 14x14 patch tokens, post-norm
feats['x_norm_alltokens'] # [1, 201, 768] full post-norm sequence
feats['x_prenorm'] # [1, 201, 768] full pre-norm sequence
```
`forward(x)` (the default, `is_training=False`) returns just the CLS token of shape
`[N, 768]`.
Dense features at arbitrary resolutions work without interpolation, because the RoPE
frequencies are derived from the actual patch grid:
```python
model.forward_features(torch.randn(1, 3, 256, 192).half())['x_norm_patchtokens'].shape
# torch.Size([1, 192, 768]) -> 16x12 patch grid
```
Intermediate layers, optionally reshaped to feature maps:
```python
with torch.no_grad():
layers = model.get_intermediate_layers(x, n=[0, 5, 11]) # tuple of [N, HW, C]
maps = model.get_intermediate_layers(x, n=1, reshape=True) # [N, C, 14, 14]
```
## Verification
The conversion was checked at every step, and these checks are reproducible via
`test/test_dinov3_vit.py` in the codon repository:
| Check | Result |
| --- | --- |
| Key remap + `strict=True` load from `model.safetensors` | passes |
| fp32 remapped weights vs `transformers` `DINOv3ViTModel`, 224×224 and 196×252 | `max|diff| = 0.000e+00` (bit-exact) |
| Per-block outputs vs `transformers` `output_hidden_states` (layers 0, 5, 11) | `max|diff| < 1e-5` |
| fp16 forward vs original fp32 model (CLS token) | `max|diff| = 6.9e-03` |
| fp16 weight quantization error (per tensor) | `≤ 2.0e-03` |
| Save → reload round trip | `max|diff| = 0.000e+00` |
The fp16 error is well within the expected range for a 12-layer model: activations
reach an order of magnitude of ~10-30, and fp16 carries roughly three significant
decimal digits.
## Precision notes
- Inference in float16 works on CPU and CUDA. The rotary `cos`/`sin` tables are
computed in fp32 and then **cast to the activation dtype**, so attention inputs
stay in float16 throughout instead of being silently promoted to fp32.
- If you need maximum fidelity, use `dtype=torch.float32`; the weights are exact
float16 representations of the original fp32 values, so this only recovers the
rounding that happened at export time.
## License
The weights are derived from Meta's DINOv3 and remain subject to the
[DINOv3 License](https://ai.meta.com/resources/models-and-libraries/dinov3-license).
Please read and comply with that license before use. In particular, the license
governs acceptable use, redistribution and attribution; this conversion adds no
additional permissions and no warranty.
|