DINOv3-ViT-Base / README.md
A03HCY's picture
Upload 4 files
245a427 verified
|
Raw History Blame Contribute Delete
7.28 kB
---
license: other
license_name: dinov3-license
license_link: https://ai.meta.com/resources/models-and-libraries/dinov3-license
base_model: facebook/dinov3-vitb16-pretrain-lvd1689m
pipeline_tag: image-feature-extraction
library_name: codon
tags:
- dino
- dinov3
- vision
- feature-extraction
- fp16
---
# DINOv3 ViT-B/16 — converted weights (codon layout, float16)
This repository hosts the weights of
[`facebook/dinov3-vitb16-pretrain-lvd1689m`](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m)
converted for the `codon` implementation of DINOv3:
- **renamed** to the codon naming convention (`codon.impl.DINOv3ViT`),
- **cast to float16** (2× smaller than the original fp32 checkpoint),
- **numerically verified** against the original weights and against the
`transformers` reference implementation.
The model itself is unchanged: same architecture, same parameter values, same
numerics up to fp16 quantization.
## Why this exists
The upstream checkpoints use the `transformers`/Meta naming layout
(`layer.{i}.attention.*`, `layer_scale{1,2}.lambda1`, `embeddings.*`) and ship in
fp32. `codon.impl.DINOv3ViT` uses its own convention (all linear layers are
`*_proj`, the block list is `blocks`, Layer Scale is `gamma{1,2}`, register tokens
are `storage_tokens`), so the weights need a one-time remap. This repository is
that remap, pre-applied and stored in fp16 so it can be loaded directly with a
single call.
## Files
| File | Description |
| --- | --- |
| `model.safetensors` | 163.4 MB, 211 tensors, all `F16`, codon key layout |
| `config.json` | Architecture summary plus `"dtype": "float16"` and `"key_layout": "codon"` |
| `preprocessor_config.json` | Image preprocessing parameters, carried over from the original model |
| `README.md` | This file |
## Key mapping
| Original (transformers / Meta) | codon |
| --- | --- |
| `embeddings.patch_embeddings.*` | `patch_embed.*` |
| `embeddings.cls_token` | `cls_token` |
| `embeddings.register_tokens` | `storage_tokens` |
| `layer.{i}.attention.q_proj.*` | `blocks.{i}.q_proj.*` |
| `layer.{i}.attention.k_proj.*` | `blocks.{i}.k_proj.*` |
| `layer.{i}.attention.v_proj.*` | `blocks.{i}.v_proj.*` |
| `layer.{i}.attention.o_proj.*` | `blocks.{i}.o_proj.*` |
| `layer.{i}.norm1.*` / `norm2.*` | `blocks.{i}.norm1.*` / `norm2.*` |
| `layer.{i}.mlp.up_proj.*` / `down_proj.*` | `blocks.{i}.up_proj.*` / `down_proj.*` |
| `layer.{i}.layer_scale1.lambda1` | `blocks.{i}.gamma1` |
| `layer.{i}.layer_scale2.lambda1` | `blocks.{i}.gamma2` |
| `norm.*` | `norm.*` |
| `rope_embeddings.inv_freq` | dropped (rebuilt dynamically from input shape) |
Two notes on the conversion:
- The `inv_freq` rotary table is **not** stored. In this implementation the 2D axial
RoPE frequencies are recomputed from `rope_theta` and the input resolution at
runtime (a non-persistent buffer), which is what lets the same weights run at any
image size.
- A zero-initialized `mask_token` is included as a placeholder. It is part of this
implementation's MAE architecture but is absent from the original checkpoint;
`load_pretrained(strict=True)` does not require it.
## Model overview
| Property | Value |
| --- | --- |
| Architecture | DINOv3 ViT (pre-norm Transformer, bidirectional attention, 2D axial RoPE) |
| Parameters | 85.7 M |
| Hidden size | 768 |
| Layers / attention heads | 12 / 12 |
| MLP hidden size | 3072 (ratio 4, GELU) |
| Patch size / default resolution | 16 / 224×224 |
| Register tokens | 4 |
| RoPE theta | 100.0 |
| Layer-norm eps | 1e-5 |
| Layer Scale init | 1.0 |
| Attention biases | q/v/o yes, k no |
| Drop path | 0.0 (inference) |
| Weight dtype | float16 |
## Usage
```python
import torch
from codon.impl import DINOv3ViT_Base
# Pulls config.json + model.safetensors from this repository.
# config.json says dtype=float16, so the model is built in float16.
model = DINOv3ViT_Base.from_remote().eval()
# Prefer fp32 compute? The weights are upcast on load.
model32 = DINOv3ViT_Base.from_remote(dtype=torch.float32).eval()
```
`from_remote()` reads the architecture from `config.json` — `num_register_tokens`,
`rope_theta`, `pos_embed_rescale`, `layer_norm_eps`, `key_bias` and `dtype` are all
applied automatically, so no manual configuration is needed.
To load the file yourself:
```python
model = DINOv3ViT_Base().half()
model.load_pretrained('model.safetensors', strict=True, dtype=torch.float16)
```
### Feature extraction
```python
x = torch.randn(1, 3, 224, 224).half()
with torch.no_grad():
feats = model.forward_features(x)
feats['x_norm_clstoken'] # [1, 768] CLS token, post-norm
feats['x_storage_tokens'] # [1, 4, 768] 4 register tokens, post-norm
feats['x_norm_patchtokens'] # [1, 196, 768] 14x14 patch tokens, post-norm
feats['x_norm_alltokens'] # [1, 201, 768] full post-norm sequence
feats['x_prenorm'] # [1, 201, 768] full pre-norm sequence
```
`forward(x)` (the default, `is_training=False`) returns just the CLS token of shape
`[N, 768]`.
Dense features at arbitrary resolutions work without interpolation, because the RoPE
frequencies are derived from the actual patch grid:
```python
model.forward_features(torch.randn(1, 3, 256, 192).half())['x_norm_patchtokens'].shape
# torch.Size([1, 192, 768]) -> 16x12 patch grid
```
Intermediate layers, optionally reshaped to feature maps:
```python
with torch.no_grad():
layers = model.get_intermediate_layers(x, n=[0, 5, 11]) # tuple of [N, HW, C]
maps = model.get_intermediate_layers(x, n=1, reshape=True) # [N, C, 14, 14]
```
## Verification
The conversion was checked at every step, and these checks are reproducible via
`test/test_dinov3_vit.py` in the codon repository:
| Check | Result |
| --- | --- |
| Key remap + `strict=True` load from `model.safetensors` | passes |
| fp32 remapped weights vs `transformers` `DINOv3ViTModel`, 224×224 and 196×252 | `max|diff| = 0.000e+00` (bit-exact) |
| Per-block outputs vs `transformers` `output_hidden_states` (layers 0, 5, 11) | `max|diff| < 1e-5` |
| fp16 forward vs original fp32 model (CLS token) | `max|diff| = 6.9e-03` |
| fp16 weight quantization error (per tensor) | `≤ 2.0e-03` |
| Save → reload round trip | `max|diff| = 0.000e+00` |
The fp16 error is well within the expected range for a 12-layer model: activations
reach an order of magnitude of ~10-30, and fp16 carries roughly three significant
decimal digits.
## Precision notes
- Inference in float16 works on CPU and CUDA. The rotary `cos`/`sin` tables are
computed in fp32 and then **cast to the activation dtype**, so attention inputs
stay in float16 throughout instead of being silently promoted to fp32.
- If you need maximum fidelity, use `dtype=torch.float32`; the weights are exact
float16 representations of the original fp32 values, so this only recovers the
rounding that happened at export time.
## License
The weights are derived from Meta's DINOv3 and remain subject to the
[DINOv3 License](https://ai.meta.com/resources/models-and-libraries/dinov3-license).
Please read and comply with that license before use. In particular, the license
governs acceptable use, redistribution and attribution; this conversion adds no
additional permissions and no warranty.