yolos-small-coreml
Model Summary
This is an unofficial Core ML conversion of hustvl/yolos-small. All credit for the original model goes to its authors (Huazhong University of Science & Technology / HuggingFace team).
About the conversion
YOLOS's ViT embedding layer has the same trap as pattern 7 (Depth Anything)'s
DINOv2: position-embedding interpolation
(InterpolateInitialPositionEmbeddings and
InterpolateMidPositionEmbeddings) always forces bicubic interpolation
when traced. At this model's trained resolution, a fixed 512x864 input, this
interpolation was confirmed to be mathematically an identity map (max error
0.0 before/after), so both interpolations are skipped before conversion.
Usage (Python / coremltools)
import numpy as np
import coremltools as ct
from PIL import Image
mlmodel = ct.models.MLModel("yolos-small_fp16.mlpackage")
image = Image.open("photo.jpg").convert("RGB").resize((864, 512)) # (width, height)
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32) # (1, 3, 512, 864), NCHW
out = mlmodel.predict({"pixel_values": pixel_values})
logits = out["logits"][0] # (100, 92): 91 COCO classes + "no object"
pred_boxes = out["pred_boxes"][0] # (100, 4): normalized (center_x, center_y, width, height)
Input is a fixed 512x864 (height, width), NCHW.
Accuracy
Compared against the PyTorch fp32 reference on one COCO validation image (for the 5 objects actually detected):
| Item | Result |
|---|---|
| logits cosine similarity | 0.99975 |
| label match rate | 100% |
| boxes max abs diff | 0.0039 (normalized to 0-1) |
Specs
| Item | Value |
|---|---|
| Base model | hustvl/yolos-small (ViT-Small, 30.7M params) |
| Precision | float16 |
| Input | 512x864, NCHW, fixed size |
| Framework | Core ML (mlprogram, minimum_deployment_target=macOS14) |
Notes
- This is a community conversion, not an official release from the YOLOS authors.
- Security audit uses model-audit-lite
(see
SECURITY.mdfor details).
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
hustvl/yolos-smallの Core ML版です。元モデルの著作権はその作者 (Huazhong University of Science & Technology / HuggingFaceチーム)に帰属します。
変換について
YOLOSのViT埋め込み層には、パターン7(Depth Anything)のDINOv2と同じ罠があります:
位置埋め込みの補間(InterpolateInitialPositionEmbeddingsと
InterpolateMidPositionEmbeddings)が、トレース時には常にbicubic補間を強制する実装に
なっています。本モデルの学習解像度である512x864固定入力では、この補間が数学的に恒等写像に
なることを確認済み(補間前後の最大誤差0.0)のため、両方の補間をスキップしてから変換しています。
使い方(Python / coremltools)
import numpy as np
import coremltools as ct
from PIL import Image
mlmodel = ct.models.MLModel("yolos-small_fp16.mlpackage")
image = Image.open("photo.jpg").convert("RGB").resize((864, 512)) # (width, height)
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32) # (1, 3, 512, 864), NCHW
out = mlmodel.predict({"pixel_values": pixel_values})
logits = out["logits"][0] # (100, 92): COCO 91クラス + 「該当なし」
pred_boxes = out["pred_boxes"][0] # (100, 4): normalized (center_x, center_y, width, height)
入力は固定サイズ512x864(height, width)、NCHW。
精度検証
PyTorch fp32リファレンスと、COCO検証画像1枚で比較(実際に検出された5件に対して):
| 項目 | 結果 |
|---|---|
| logitsコサイン類似度 | 0.99975 |
| ラベル一致率 | 100% |
| boxes最大絶対誤差 | 0.0039(0〜1に正規化) |
Specs
| Item | Value |
|---|---|
| ベースモデル | hustvl/yolos-small(ViT-Small、30.7M params) |
| 精度 | float16 |
| 入力 | 512x864、NCHW、固定サイズ |
| フレームワーク | Core ML(mlprogram、minimum_deployment_target=macOS14) |
備考
- 本変換は非公式のコミュニティ版です。
- セキュリティー監査にはmodel-audit-liteを
使用しています(詳細は
SECURITY.md)。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
- Downloads last month
- 11
Model tree for masahiroid/yolos-small-coreml
Base model
hustvl/yolos-small