--- license: other license_name: dinov3-license license_link: https://ai.meta.com/resources/models-and-libraries/dinov3-license base_model: facebook/dinov3-vitb16-pretrain-lvd1689m pipeline_tag: image-feature-extraction library_name: codon tags: - dino - dinov3 - vision - feature-extraction - fp16 --- # DINOv3 ViT-B/16 — converted weights (codon layout, float16) This repository hosts the weights of [`facebook/dinov3-vitb16-pretrain-lvd1689m`](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m) converted for the `codon` implementation of DINOv3: - **renamed** to the codon naming convention (`codon.impl.DINOv3ViT`), - **cast to float16** (2× smaller than the original fp32 checkpoint), - **numerically verified** against the original weights and against the `transformers` reference implementation. The model itself is unchanged: same architecture, same parameter values, same numerics up to fp16 quantization. ## Why this exists The upstream checkpoints use the `transformers`/Meta naming layout (`layer.{i}.attention.*`, `layer_scale{1,2}.lambda1`, `embeddings.*`) and ship in fp32. `codon.impl.DINOv3ViT` uses its own convention (all linear layers are `*_proj`, the block list is `blocks`, Layer Scale is `gamma{1,2}`, register tokens are `storage_tokens`), so the weights need a one-time remap. This repository is that remap, pre-applied and stored in fp16 so it can be loaded directly with a single call. ## Files | File | Description | | --- | --- | | `model.safetensors` | 163.4 MB, 211 tensors, all `F16`, codon key layout | | `config.json` | Architecture summary plus `"dtype": "float16"` and `"key_layout": "codon"` | | `preprocessor_config.json` | Image preprocessing parameters, carried over from the original model | | `README.md` | This file | ## Key mapping | Original (transformers / Meta) | codon | | --- | --- | | `embeddings.patch_embeddings.*` | `patch_embed.*` | | `embeddings.cls_token` | `cls_token` | | `embeddings.register_tokens` | `storage_tokens` | | `layer.{i}.attention.q_proj.*` | `blocks.{i}.q_proj.*` | | `layer.{i}.attention.k_proj.*` | `blocks.{i}.k_proj.*` | | `layer.{i}.attention.v_proj.*` | `blocks.{i}.v_proj.*` | | `layer.{i}.attention.o_proj.*` | `blocks.{i}.o_proj.*` | | `layer.{i}.norm1.*` / `norm2.*` | `blocks.{i}.norm1.*` / `norm2.*` | | `layer.{i}.mlp.up_proj.*` / `down_proj.*` | `blocks.{i}.up_proj.*` / `down_proj.*` | | `layer.{i}.layer_scale1.lambda1` | `blocks.{i}.gamma1` | | `layer.{i}.layer_scale2.lambda1` | `blocks.{i}.gamma2` | | `norm.*` | `norm.*` | | `rope_embeddings.inv_freq` | dropped (rebuilt dynamically from input shape) | Two notes on the conversion: - The `inv_freq` rotary table is **not** stored. In this implementation the 2D axial RoPE frequencies are recomputed from `rope_theta` and the input resolution at runtime (a non-persistent buffer), which is what lets the same weights run at any image size. - A zero-initialized `mask_token` is included as a placeholder. It is part of this implementation's MAE architecture but is absent from the original checkpoint; `load_pretrained(strict=True)` does not require it. ## Model overview | Property | Value | | --- | --- | | Architecture | DINOv3 ViT (pre-norm Transformer, bidirectional attention, 2D axial RoPE) | | Parameters | 85.7 M | | Hidden size | 768 | | Layers / attention heads | 12 / 12 | | MLP hidden size | 3072 (ratio 4, GELU) | | Patch size / default resolution | 16 / 224×224 | | Register tokens | 4 | | RoPE theta | 100.0 | | Layer-norm eps | 1e-5 | | Layer Scale init | 1.0 | | Attention biases | q/v/o yes, k no | | Drop path | 0.0 (inference) | | Weight dtype | float16 | ## Usage ```python import torch from codon.impl import DINOv3ViT_Base # Pulls config.json + model.safetensors from this repository. # config.json says dtype=float16, so the model is built in float16. model = DINOv3ViT_Base.from_remote().eval() # Prefer fp32 compute? The weights are upcast on load. model32 = DINOv3ViT_Base.from_remote(dtype=torch.float32).eval() ``` `from_remote()` reads the architecture from `config.json` — `num_register_tokens`, `rope_theta`, `pos_embed_rescale`, `layer_norm_eps`, `key_bias` and `dtype` are all applied automatically, so no manual configuration is needed. To load the file yourself: ```python model = DINOv3ViT_Base().half() model.load_pretrained('model.safetensors', strict=True, dtype=torch.float16) ``` ### Feature extraction ```python x = torch.randn(1, 3, 224, 224).half() with torch.no_grad(): feats = model.forward_features(x) feats['x_norm_clstoken'] # [1, 768] CLS token, post-norm feats['x_storage_tokens'] # [1, 4, 768] 4 register tokens, post-norm feats['x_norm_patchtokens'] # [1, 196, 768] 14x14 patch tokens, post-norm feats['x_norm_alltokens'] # [1, 201, 768] full post-norm sequence feats['x_prenorm'] # [1, 201, 768] full pre-norm sequence ``` `forward(x)` (the default, `is_training=False`) returns just the CLS token of shape `[N, 768]`. Dense features at arbitrary resolutions work without interpolation, because the RoPE frequencies are derived from the actual patch grid: ```python model.forward_features(torch.randn(1, 3, 256, 192).half())['x_norm_patchtokens'].shape # torch.Size([1, 192, 768]) -> 16x12 patch grid ``` Intermediate layers, optionally reshaped to feature maps: ```python with torch.no_grad(): layers = model.get_intermediate_layers(x, n=[0, 5, 11]) # tuple of [N, HW, C] maps = model.get_intermediate_layers(x, n=1, reshape=True) # [N, C, 14, 14] ``` ## Verification The conversion was checked at every step, and these checks are reproducible via `test/test_dinov3_vit.py` in the codon repository: | Check | Result | | --- | --- | | Key remap + `strict=True` load from `model.safetensors` | passes | | fp32 remapped weights vs `transformers` `DINOv3ViTModel`, 224×224 and 196×252 | `max|diff| = 0.000e+00` (bit-exact) | | Per-block outputs vs `transformers` `output_hidden_states` (layers 0, 5, 11) | `max|diff| < 1e-5` | | fp16 forward vs original fp32 model (CLS token) | `max|diff| = 6.9e-03` | | fp16 weight quantization error (per tensor) | `≤ 2.0e-03` | | Save → reload round trip | `max|diff| = 0.000e+00` | The fp16 error is well within the expected range for a 12-layer model: activations reach an order of magnitude of ~10-30, and fp16 carries roughly three significant decimal digits. ## Precision notes - Inference in float16 works on CPU and CUDA. The rotary `cos`/`sin` tables are computed in fp32 and then **cast to the activation dtype**, so attention inputs stay in float16 throughout instead of being silently promoted to fp32. - If you need maximum fidelity, use `dtype=torch.float32`; the weights are exact float16 representations of the original fp32 values, so this only recovers the rounding that happened at export time. ## License The weights are derived from Meta's DINOv3 and remain subject to the [DINOv3 License](https://ai.meta.com/resources/models-and-libraries/dinov3-license). Please read and comply with that license before use. In particular, the license governs acceptable use, redistribution and attribution; this conversion adds no additional permissions and no warranty.