Expose patch tokens alongside the CLS embedding
Browse filesThe backbone already computed all 789 tokens, so returning the 784 patch ones
costs no extra inference. Registers stay dropped; they are artifact sinks, not
features. Patches are unnormalized because dense heads want the magnitude.
Dropping CLS from the artifact name since it no longer describes the outputs.
The weights are byte-identical, so LFS reuses the object already pushed.
- README.md +18 -11
- SHA256SUMS +5 -5
- convert.py +31 -15
- examples/embed.py +7 -2
- models/{DINOv3ViTB16CLS-FP32-448.mlpackage → DINOv3ViTB16-FP32-448.mlpackage}/Data/com.apple.CoreML/model.mlmodel +2 -2
- models/{DINOv3ViTB16CLS-FP32-448.mlpackage → DINOv3ViTB16-FP32-448.mlpackage}/Data/com.apple.CoreML/weights/weight.bin +0 -0
- models/{DINOv3ViTB16CLS-FP32-448.mlpackage → DINOv3ViTB16-FP32-448.mlpackage}/Manifest.json +8 -8
- models/{DINOv3ViTB16CLS-FP32-448.mlpackage → DINOv3ViTB16-FP32-448.mlpackage}/executorch_debug_handle_mapping.json +0 -0
- models/{DINOv3ViTB16CLS-FP32-448.validation.json → DINOv3ViTB16-FP32-448.validation.json} +5 -0
README.md
CHANGED
|
@@ -14,8 +14,9 @@ tags:
|
|
| 14 |
|
| 15 |
# DINOv3 Core ML
|
| 16 |
|
| 17 |
-
A Core ML conversion of Meta's DINOv3 ViT-B/16 image encoder. It returns
|
| 18 |
-
L2-normalized, 768-value CLS embedding
|
|
|
|
| 19 |
This is a format conversion of pretrained weights; no additional training was performed.
|
| 20 |
|
| 21 |
## Model
|
|
@@ -24,18 +25,24 @@ This is a format conversion of pretrained weights; no additional training was pe
|
|
| 24 |
|---|---|
|
| 25 |
| Base model | `facebook/dinov3-vitb16-pretrain-lvd1689m` |
|
| 26 |
| timm implementation | `vit_base_patch16_dinov3.lvd1689m` |
|
| 27 |
-
| Artifact | `models/
|
| 28 |
| Format | Core ML ML Program, FP32 |
|
| 29 |
| Target | macOS 14 or newer |
|
| 30 |
| Input | `image`: 448×448 RGB, pixel values 0–255 |
|
| 31 |
| Output | `embedding`: float32, shape `[1, 768]`, unit L2 norm |
|
|
|
|
| 32 |
|
| 33 |
Correct EXIF orientation, convert to RGB, and resize to 448×448 before inference.
|
| 34 |
The examples use a square resize, which can distort non-square images. Use the same
|
| 35 |
preprocessing for all images being compared. Pixel scaling and ImageNet normalization
|
| 36 |
-
are inside the model; do not apply them again.
|
| 37 |
-
|
| 38 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
|
| 40 |
## Run
|
| 41 |
|
|
@@ -56,7 +63,7 @@ For native Swift, provide an already resized, orientation-corrected 448×448 ima
|
|
| 56 |
```sh
|
| 57 |
mkdir -p build
|
| 58 |
swiftc -O examples/encode.swift -o build/encode
|
| 59 |
-
build/encode models/
|
| 60 |
```
|
| 61 |
|
| 62 |
The example compiles the package at runtime. An app can instead add the package to
|
|
@@ -66,21 +73,21 @@ resource once for repeated predictions.
|
|
| 66 |
## Convert
|
| 67 |
|
| 68 |
```sh
|
| 69 |
-
uv run python convert.py --output build/
|
| 70 |
```
|
| 71 |
|
| 72 |
The converter downloads pretrained weights through timm/Hugging Face. If the upstream
|
| 73 |
checkpoint requires access, accept its terms and authenticate with Hugging Face first.
|
| 74 |
It exports through `torch.export`, converts to Core ML, and compares both runtimes
|
| 75 |
-
on a synthetic RGB gradient. A cosine below 0.999 fails validation.
|
| 76 |
|
| 77 |
`--size` changes the square input resolution; use a positive multiple of 16.
|
| 78 |
`--precision fp16` is experimental and must pass its own validation.
|
| 79 |
Existing output packages are protected unless `--force` is supplied.
|
| 80 |
|
| 81 |
The copied FP32 artifact's original validation result is in
|
| 82 |
-
`models/
|
| 83 |
-
and
|
| 84 |
above 1. This single-image check establishes limited conversion parity, not retrieval
|
| 85 |
accuracy across datasets. The supplied package retains its original conversion metadata.
|
| 86 |
|
|
|
|
| 14 |
|
| 15 |
# DINOv3 Core ML
|
| 16 |
|
| 17 |
+
A Core ML conversion of Meta's DINOv3 ViT-B/16 image encoder. It returns an
|
| 18 |
+
L2-normalized, 768-value CLS embedding for image similarity and retrieval, plus the
|
| 19 |
+
784 raw patch tokens for dense tasks such as segmentation, depth, and correspondence.
|
| 20 |
This is a format conversion of pretrained weights; no additional training was performed.
|
| 21 |
|
| 22 |
## Model
|
|
|
|
| 25 |
|---|---|
|
| 26 |
| Base model | `facebook/dinov3-vitb16-pretrain-lvd1689m` |
|
| 27 |
| timm implementation | `vit_base_patch16_dinov3.lvd1689m` |
|
| 28 |
+
| Artifact | `models/DINOv3ViTB16-FP32-448.mlpackage` |
|
| 29 |
| Format | Core ML ML Program, FP32 |
|
| 30 |
| Target | macOS 14 or newer |
|
| 31 |
| Input | `image`: 448×448 RGB, pixel values 0–255 |
|
| 32 |
| Output | `embedding`: float32, shape `[1, 768]`, unit L2 norm |
|
| 33 |
+
| Output | `patch_embeddings`: float32, shape `[1, 784, 768]`, unnormalized |
|
| 34 |
|
| 35 |
Correct EXIF orientation, convert to RGB, and resize to 448×448 before inference.
|
| 36 |
The examples use a square resize, which can distort non-square images. Use the same
|
| 37 |
preprocessing for all images being compared. Pixel scaling and ImageNet normalization
|
| 38 |
+
are inside the model; do not apply them again. Similarity is the dot product of two
|
| 39 |
+
`embedding` vectors. It does not provide text embeddings, captions, or coordinates.
|
| 40 |
+
|
| 41 |
+
`patch_embeddings` is a 28x28 grid of 768-value tokens flattened in row-major order,
|
| 42 |
+
one per 16x16 input patch. Unlike `embedding` these are not normalized, because dense
|
| 43 |
+
heads generally want the magnitude; normalize per token yourself for cosine. The four
|
| 44 |
+
register tokens are dropped: they exist to absorb high-norm artifacts that would
|
| 45 |
+
otherwise pollute the patch tokens, and are not useful as features.
|
| 46 |
|
| 47 |
## Run
|
| 48 |
|
|
|
|
| 63 |
```sh
|
| 64 |
mkdir -p build
|
| 65 |
swiftc -O examples/encode.swift -o build/encode
|
| 66 |
+
build/encode models/DINOv3ViTB16-FP32-448.mlpackage image-448.png
|
| 67 |
```
|
| 68 |
|
| 69 |
The example compiles the package at runtime. An app can instead add the package to
|
|
|
|
| 73 |
## Convert
|
| 74 |
|
| 75 |
```sh
|
| 76 |
+
uv run python convert.py --output build/DINOv3ViTB16-FP32-448.mlpackage
|
| 77 |
```
|
| 78 |
|
| 79 |
The converter downloads pretrained weights through timm/Hugging Face. If the upstream
|
| 80 |
checkpoint requires access, accept its terms and authenticate with Hugging Face first.
|
| 81 |
It exports through `torch.export`, converts to Core ML, and compares both runtimes
|
| 82 |
+
on a synthetic RGB gradient. A cosine below 0.999 on either output fails validation.
|
| 83 |
|
| 84 |
`--size` changes the square input resolution; use a positive multiple of 16.
|
| 85 |
`--precision fp16` is experimental and must pass its own validation.
|
| 86 |
Existing output packages are protected unless `--force` is supplied.
|
| 87 |
|
| 88 |
The copied FP32 artifact's original validation result is in
|
| 89 |
+
`models/DINOv3ViTB16-FP32-448.validation.json`: cosine approximately 1.0 for both
|
| 90 |
+
outputs, and CLS norm approximately 1.0. Floating-point rounding can put cosine slightly
|
| 91 |
above 1. This single-image check establishes limited conversion parity, not retrieval
|
| 92 |
accuracy across datasets. The supplied package retains its original conversion metadata.
|
| 93 |
|
SHA256SUMS
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
-
|
| 2 |
-
12153d208093c1bfb72eed482d43a695e4062a7b654802355615145b088d5203 models/
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
|
|
|
| 1 |
+
a88e2a595ae587c134fa664280be243d2b9246643fa28e455601464125db7ea5 models/DINOv3ViTB16-FP32-448.mlpackage/Data/com.apple.CoreML/model.mlmodel
|
| 2 |
+
12153d208093c1bfb72eed482d43a695e4062a7b654802355615145b088d5203 models/DINOv3ViTB16-FP32-448.mlpackage/Data/com.apple.CoreML/weights/weight.bin
|
| 3 |
+
b080024a442c0a0d133905b7d79d25619451c784b10f456e0572dc81ac2270c9 models/DINOv3ViTB16-FP32-448.mlpackage/executorch_debug_handle_mapping.json
|
| 4 |
+
a52ae1bc4b1203e88e0ff3355a347f8fbfbcad9150ea0df5073f16d1bfaca380 models/DINOv3ViTB16-FP32-448.mlpackage/Manifest.json
|
| 5 |
+
9b01f646e006c36f84c5a8bdd4e7d006c805c25f0346380d97ff01506611a500 models/DINOv3ViTB16-FP32-448.validation.json
|
convert.py
CHANGED
|
@@ -17,18 +17,26 @@ IMAGENET_MEAN = (0.485, 0.456, 0.406)
|
|
| 17 |
IMAGENET_STD = (0.229, 0.224, 0.225)
|
| 18 |
|
| 19 |
|
| 20 |
-
class
|
| 21 |
def __init__(self, backbone: torch.nn.Module) -> None:
|
| 22 |
super().__init__()
|
| 23 |
self.backbone = backbone
|
|
|
|
| 24 |
self.register_buffer("mean", torch.tensor(IMAGENET_MEAN).view(1, 3, 1, 1))
|
| 25 |
self.register_buffer("std", torch.tensor(IMAGENET_STD).view(1, 3, 1, 1))
|
| 26 |
|
| 27 |
-
def forward(self, image: torch.Tensor) -> torch.Tensor:
|
| 28 |
pixels = (image / 255.0 - self.mean) / self.std
|
| 29 |
-
|
|
|
|
| 30 |
norm = torch.sqrt(torch.sum(cls * cls, dim=-1, keepdim=True).clamp_min(1e-12))
|
| 31 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 32 |
|
| 33 |
|
| 34 |
def parse_args() -> argparse.Namespace:
|
|
@@ -57,7 +65,7 @@ def main() -> None:
|
|
| 57 |
raise SystemExit("--size must be a positive multiple of 16")
|
| 58 |
if args.output is None:
|
| 59 |
precision_name = args.precision.upper()
|
| 60 |
-
args.output = Path(f"models/
|
| 61 |
if args.output.exists():
|
| 62 |
if not args.force:
|
| 63 |
raise SystemExit(f"{args.output} already exists; pass --force to replace it")
|
|
@@ -70,7 +78,7 @@ def main() -> None:
|
|
| 70 |
img_size=args.size,
|
| 71 |
num_classes=0,
|
| 72 |
).eval()
|
| 73 |
-
model =
|
| 74 |
|
| 75 |
with torch.inference_mode():
|
| 76 |
exported = torch.export.export(model, (example,)).run_decompositions({})
|
|
@@ -86,16 +94,17 @@ def main() -> None:
|
|
| 86 |
color_layout=ct.colorlayout.RGB,
|
| 87 |
)
|
| 88 |
],
|
| 89 |
-
outputs=[ct.TensorType(name="embedding")],
|
| 90 |
minimum_deployment_target=ct.target.macOS14,
|
| 91 |
compute_precision=precision,
|
| 92 |
)
|
| 93 |
coreml_model.author = "dinov3-coreml; base model by Meta"
|
| 94 |
coreml_model.license = "DINOv3 License"
|
| 95 |
-
coreml_model.short_description = "DINOv3 ViT-B/16
|
| 96 |
coreml_model.user_defined_metadata["base_model"] = BASE_MODEL
|
| 97 |
coreml_model.input_description["image"] = f"RGB image resized to {args.size}x{args.size}"
|
| 98 |
coreml_model.output_description["embedding"] = "L2-normalized 768-value CLS embedding"
|
|
|
|
| 99 |
|
| 100 |
args.output.parent.mkdir(parents=True, exist_ok=True)
|
| 101 |
coreml_model.save(args.output)
|
|
@@ -103,25 +112,32 @@ def main() -> None:
|
|
| 103 |
image = synthetic_image(args.size)
|
| 104 |
array = np.asarray(image, dtype=np.float32).transpose(2, 0, 1)[None, ...]
|
| 105 |
with torch.inference_mode():
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 112 |
report = {
|
| 113 |
"model": DEFAULT_MODEL,
|
| 114 |
"base_model": BASE_MODEL,
|
| 115 |
"input_size": args.size,
|
| 116 |
"precision": args.precision,
|
| 117 |
"output_shape": list(coreml_output.shape),
|
|
|
|
| 118 |
"pytorch_coreml_cosine_similarity": cosine,
|
|
|
|
| 119 |
"coreml_output_l2_norm": float(np.linalg.norm(coreml_output)),
|
| 120 |
}
|
| 121 |
report_path = args.output.with_suffix(".validation.json")
|
| 122 |
report_path.write_text(json.dumps(report, indent=2) + "\n")
|
| 123 |
print(json.dumps({"output": str(args.output), **report}, indent=2))
|
| 124 |
-
if cosine < 0.999:
|
| 125 |
raise SystemExit("Core ML parity check failed: cosine similarity is below 0.999")
|
| 126 |
|
| 127 |
|
|
|
|
| 17 |
IMAGENET_STD = (0.229, 0.224, 0.225)
|
| 18 |
|
| 19 |
|
| 20 |
+
class DINOv3Encoder(torch.nn.Module):
|
| 21 |
def __init__(self, backbone: torch.nn.Module) -> None:
|
| 22 |
super().__init__()
|
| 23 |
self.backbone = backbone
|
| 24 |
+
self.num_prefix_tokens = backbone.num_prefix_tokens
|
| 25 |
self.register_buffer("mean", torch.tensor(IMAGENET_MEAN).view(1, 3, 1, 1))
|
| 26 |
self.register_buffer("std", torch.tensor(IMAGENET_STD).view(1, 3, 1, 1))
|
| 27 |
|
| 28 |
+
def forward(self, image: torch.Tensor) -> tuple[torch.Tensor, torch.Tensor]:
|
| 29 |
pixels = (image / 255.0 - self.mean) / self.std
|
| 30 |
+
features = self.backbone.forward_features(pixels)
|
| 31 |
+
cls = features[:, 0]
|
| 32 |
norm = torch.sqrt(torch.sum(cls * cls, dim=-1, keepdim=True).clamp_min(1e-12))
|
| 33 |
+
# The prefix is one CLS token plus four register tokens. Registers absorb
|
| 34 |
+
# high-norm artifacts that would otherwise pollute the patch tokens; they
|
| 35 |
+
# are not features and nothing downstream uses them.
|
| 36 |
+
patches = features[:, self.num_prefix_tokens:]
|
| 37 |
+
# Patches stay unnormalized because dense heads generally want the
|
| 38 |
+
# magnitude. Callers doing cosine can normalize per token themselves.
|
| 39 |
+
return cls / norm, patches
|
| 40 |
|
| 41 |
|
| 42 |
def parse_args() -> argparse.Namespace:
|
|
|
|
| 65 |
raise SystemExit("--size must be a positive multiple of 16")
|
| 66 |
if args.output is None:
|
| 67 |
precision_name = args.precision.upper()
|
| 68 |
+
args.output = Path(f"models/DINOv3ViTB16-{precision_name}-{args.size}.mlpackage")
|
| 69 |
if args.output.exists():
|
| 70 |
if not args.force:
|
| 71 |
raise SystemExit(f"{args.output} already exists; pass --force to replace it")
|
|
|
|
| 78 |
img_size=args.size,
|
| 79 |
num_classes=0,
|
| 80 |
).eval()
|
| 81 |
+
model = DINOv3Encoder(backbone).eval()
|
| 82 |
|
| 83 |
with torch.inference_mode():
|
| 84 |
exported = torch.export.export(model, (example,)).run_decompositions({})
|
|
|
|
| 94 |
color_layout=ct.colorlayout.RGB,
|
| 95 |
)
|
| 96 |
],
|
| 97 |
+
outputs=[ct.TensorType(name="embedding"), ct.TensorType(name="patch_embeddings")],
|
| 98 |
minimum_deployment_target=ct.target.macOS14,
|
| 99 |
compute_precision=precision,
|
| 100 |
)
|
| 101 |
coreml_model.author = "dinov3-coreml; base model by Meta"
|
| 102 |
coreml_model.license = "DINOv3 License"
|
| 103 |
+
coreml_model.short_description = "DINOv3 ViT-B/16 CLS embedding and patch tokens"
|
| 104 |
coreml_model.user_defined_metadata["base_model"] = BASE_MODEL
|
| 105 |
coreml_model.input_description["image"] = f"RGB image resized to {args.size}x{args.size}"
|
| 106 |
coreml_model.output_description["embedding"] = "L2-normalized 768-value CLS embedding"
|
| 107 |
+
coreml_model.output_description["patch_embeddings"] = "Unnormalized patch tokens, one per 16x16 patch"
|
| 108 |
|
| 109 |
args.output.parent.mkdir(parents=True, exist_ok=True)
|
| 110 |
coreml_model.save(args.output)
|
|
|
|
| 112 |
image = synthetic_image(args.size)
|
| 113 |
array = np.asarray(image, dtype=np.float32).transpose(2, 0, 1)[None, ...]
|
| 114 |
with torch.inference_mode():
|
| 115 |
+
torch_cls, torch_patches = (t.numpy()[0] for t in model(torch.from_numpy(array)))
|
| 116 |
+
prediction = coreml_model.predict({"image": image})
|
| 117 |
+
coreml_output = np.asarray(prediction["embedding"])[0]
|
| 118 |
+
coreml_patches = np.asarray(prediction["patch_embeddings"])[0]
|
| 119 |
+
|
| 120 |
+
def cosine_of(a: np.ndarray, b: np.ndarray) -> float:
|
| 121 |
+
a, b = a.reshape(-1), b.reshape(-1)
|
| 122 |
+
return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
|
| 123 |
+
|
| 124 |
+
cosine = cosine_of(torch_cls, coreml_output)
|
| 125 |
+
patch_cosine = cosine_of(torch_patches, coreml_patches)
|
| 126 |
report = {
|
| 127 |
"model": DEFAULT_MODEL,
|
| 128 |
"base_model": BASE_MODEL,
|
| 129 |
"input_size": args.size,
|
| 130 |
"precision": args.precision,
|
| 131 |
"output_shape": list(coreml_output.shape),
|
| 132 |
+
"patch_output_shape": list(coreml_patches.shape),
|
| 133 |
"pytorch_coreml_cosine_similarity": cosine,
|
| 134 |
+
"pytorch_coreml_patch_cosine_similarity": patch_cosine,
|
| 135 |
"coreml_output_l2_norm": float(np.linalg.norm(coreml_output)),
|
| 136 |
}
|
| 137 |
report_path = args.output.with_suffix(".validation.json")
|
| 138 |
report_path.write_text(json.dumps(report, indent=2) + "\n")
|
| 139 |
print(json.dumps({"output": str(args.output), **report}, indent=2))
|
| 140 |
+
if min(cosine, patch_cosine) < 0.999:
|
| 141 |
raise SystemExit("Core ML parity check failed: cosine similarity is below 0.999")
|
| 142 |
|
| 143 |
|
examples/embed.py
CHANGED
|
@@ -11,7 +11,7 @@ from PIL import Image, ImageOps
|
|
| 11 |
def main() -> None:
|
| 12 |
parser = argparse.ArgumentParser(description=__doc__)
|
| 13 |
parser.add_argument("image", type=Path, nargs="?")
|
| 14 |
-
parser.add_argument("--model", type=Path, default=Path("models/
|
| 15 |
parser.add_argument("--output", type=Path)
|
| 16 |
args = parser.parse_args()
|
| 17 |
model = ct.models.MLModel(str(args.model), compute_units=ct.ComputeUnit.ALL)
|
|
@@ -24,9 +24,13 @@ def main() -> None:
|
|
| 24 |
y, x = np.mgrid[0:size[1], 0:size[0]]
|
| 25 |
pixels = np.stack((x * 255 // size[0], y * 255 // size[1], (x + y) * 255 // sum(size)), axis=-1).astype(np.uint8)
|
| 26 |
image = Image.fromarray(pixels)
|
| 27 |
-
|
|
|
|
|
|
|
| 28 |
if vector.shape != (768,) or not np.isfinite(vector).all():
|
| 29 |
raise SystemExit("Invalid embedding")
|
|
|
|
|
|
|
| 30 |
norm = float(np.linalg.norm(vector))
|
| 31 |
if not np.isclose(norm, 1, atol=1e-4):
|
| 32 |
raise SystemExit(f"Embedding is not normalized: {norm}")
|
|
@@ -34,6 +38,7 @@ def main() -> None:
|
|
| 34 |
args.output.parent.mkdir(parents=True, exist_ok=True)
|
| 35 |
np.save(args.output, vector)
|
| 36 |
print(f"shape={vector.shape}, dtype={vector.dtype}, L2 norm={norm:.8f}")
|
|
|
|
| 37 |
|
| 38 |
|
| 39 |
if __name__ == "__main__":
|
|
|
|
| 11 |
def main() -> None:
|
| 12 |
parser = argparse.ArgumentParser(description=__doc__)
|
| 13 |
parser.add_argument("image", type=Path, nargs="?")
|
| 14 |
+
parser.add_argument("--model", type=Path, default=Path("models/DINOv3ViTB16-FP32-448.mlpackage"))
|
| 15 |
parser.add_argument("--output", type=Path)
|
| 16 |
args = parser.parse_args()
|
| 17 |
model = ct.models.MLModel(str(args.model), compute_units=ct.ComputeUnit.ALL)
|
|
|
|
| 24 |
y, x = np.mgrid[0:size[1], 0:size[0]]
|
| 25 |
pixels = np.stack((x * 255 // size[0], y * 255 // size[1], (x + y) * 255 // sum(size)), axis=-1).astype(np.uint8)
|
| 26 |
image = Image.fromarray(pixels)
|
| 27 |
+
prediction = model.predict({"image": image})
|
| 28 |
+
vector = np.asarray(prediction["embedding"], dtype=np.float32).reshape(-1)
|
| 29 |
+
patches = np.asarray(prediction["patch_embeddings"], dtype=np.float32).reshape(-1, 768)
|
| 30 |
if vector.shape != (768,) or not np.isfinite(vector).all():
|
| 31 |
raise SystemExit("Invalid embedding")
|
| 32 |
+
if not np.isfinite(patches).all():
|
| 33 |
+
raise SystemExit("Invalid patch embeddings")
|
| 34 |
norm = float(np.linalg.norm(vector))
|
| 35 |
if not np.isclose(norm, 1, atol=1e-4):
|
| 36 |
raise SystemExit(f"Embedding is not normalized: {norm}")
|
|
|
|
| 38 |
args.output.parent.mkdir(parents=True, exist_ok=True)
|
| 39 |
np.save(args.output, vector)
|
| 40 |
print(f"shape={vector.shape}, dtype={vector.dtype}, L2 norm={norm:.8f}")
|
| 41 |
+
print(f"patches={patches.shape} (unnormalized)")
|
| 42 |
|
| 43 |
|
| 44 |
if __name__ == "__main__":
|
models/{DINOv3ViTB16CLS-FP32-448.mlpackage → DINOv3ViTB16-FP32-448.mlpackage}/Data/com.apple.CoreML/model.mlmodel
RENAMED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:a88e2a595ae587c134fa664280be243d2b9246643fa28e455601464125db7ea5
|
| 3 |
+
size 182835
|
models/{DINOv3ViTB16CLS-FP32-448.mlpackage → DINOv3ViTB16-FP32-448.mlpackage}/Data/com.apple.CoreML/weights/weight.bin
RENAMED
|
File without changes
|
models/{DINOv3ViTB16CLS-FP32-448.mlpackage → DINOv3ViTB16-FP32-448.mlpackage}/Manifest.json
RENAMED
|
@@ -1,18 +1,18 @@
|
|
| 1 |
{
|
| 2 |
"fileFormatVersion": "1.0.0",
|
| 3 |
"itemInfoEntries": {
|
| 4 |
-
"
|
| 5 |
-
"author": "com.apple.CoreML",
|
| 6 |
-
"description": "CoreML Model Weights",
|
| 7 |
-
"name": "weights",
|
| 8 |
-
"path": "com.apple.CoreML/weights"
|
| 9 |
-
},
|
| 10 |
-
"4C46342C-1CCD-429A-AAC6-E4BE6AD8B1D0": {
|
| 11 |
"author": "com.apple.CoreML",
|
| 12 |
"description": "CoreML Model Specification",
|
| 13 |
"name": "model.mlmodel",
|
| 14 |
"path": "com.apple.CoreML/model.mlmodel"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 15 |
}
|
| 16 |
},
|
| 17 |
-
"rootModelIdentifier": "
|
| 18 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"fileFormatVersion": "1.0.0",
|
| 3 |
"itemInfoEntries": {
|
| 4 |
+
"5A08B9B3-DA0A-4DD2-9DBC-A68B509B51F6": {
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 5 |
"author": "com.apple.CoreML",
|
| 6 |
"description": "CoreML Model Specification",
|
| 7 |
"name": "model.mlmodel",
|
| 8 |
"path": "com.apple.CoreML/model.mlmodel"
|
| 9 |
+
},
|
| 10 |
+
"F9776A49-BC3F-4795-BFC2-9F9C12DEC863": {
|
| 11 |
+
"author": "com.apple.CoreML",
|
| 12 |
+
"description": "CoreML Model Weights",
|
| 13 |
+
"name": "weights",
|
| 14 |
+
"path": "com.apple.CoreML/weights"
|
| 15 |
}
|
| 16 |
},
|
| 17 |
+
"rootModelIdentifier": "5A08B9B3-DA0A-4DD2-9DBC-A68B509B51F6"
|
| 18 |
}
|
models/{DINOv3ViTB16CLS-FP32-448.mlpackage → DINOv3ViTB16-FP32-448.mlpackage}/executorch_debug_handle_mapping.json
RENAMED
|
The diff for this file is too large to render.
See raw diff
|
|
|
models/{DINOv3ViTB16CLS-FP32-448.validation.json → DINOv3ViTB16-FP32-448.validation.json}
RENAMED
|
@@ -6,6 +6,11 @@
|
|
| 6 |
"output_shape": [
|
| 7 |
768
|
| 8 |
],
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
"pytorch_coreml_cosine_similarity": 1.0000001192092896,
|
|
|
|
| 10 |
"coreml_output_l2_norm": 0.9999999403953552
|
| 11 |
}
|
|
|
|
| 6 |
"output_shape": [
|
| 7 |
768
|
| 8 |
],
|
| 9 |
+
"patch_output_shape": [
|
| 10 |
+
784,
|
| 11 |
+
768
|
| 12 |
+
],
|
| 13 |
"pytorch_coreml_cosine_similarity": 1.0000001192092896,
|
| 14 |
+
"pytorch_coreml_patch_cosine_similarity": 1.0000001192092896,
|
| 15 |
"coreml_output_l2_norm": 0.9999999403953552
|
| 16 |
}
|