yolos-small-coreml

English | 日本語

Model Summary

This is an unofficial Core ML conversion of hustvl/yolos-small. All credit for the original model goes to its authors (Huazhong University of Science & Technology / HuggingFace team).

About the conversion

YOLOS's ViT embedding layer has the same trap as pattern 7 (Depth Anything)'s DINOv2: position-embedding interpolation (InterpolateInitialPositionEmbeddings and InterpolateMidPositionEmbeddings) always forces bicubic interpolation when traced. At this model's trained resolution, a fixed 512x864 input, this interpolation was confirmed to be mathematically an identity map (max error 0.0 before/after), so both interpolations are skipped before conversion.

Usage (Python / coremltools)

import numpy as np
import coremltools as ct
from PIL import Image

mlmodel = ct.models.MLModel("yolos-small_fp16.mlpackage")

image = Image.open("photo.jpg").convert("RGB").resize((864, 512))  # (width, height)
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32)  # (1, 3, 512, 864), NCHW

out = mlmodel.predict({"pixel_values": pixel_values})
logits = out["logits"][0]       # (100, 92): 91 COCO classes + "no object"
pred_boxes = out["pred_boxes"][0]  # (100, 4): normalized (center_x, center_y, width, height)

Input is a fixed 512x864 (height, width), NCHW.

Accuracy

Compared against the PyTorch fp32 reference on one COCO validation image (for the 5 objects actually detected):

Item Result
logits cosine similarity 0.99975
label match rate 100%
boxes max abs diff 0.0039 (normalized to 0-1)

Specs

Item Value
Base model hustvl/yolos-small (ViT-Small, 30.7M params)
Precision float16
Input 512x864, NCHW, fixed size
Framework Core ML (mlprogram, minimum_deployment_target=macOS14)

Notes

  • This is a community conversion, not an official release from the YOLOS authors.
  • Security audit uses model-audit-lite (see SECURITY.md for details).

Security

Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.


モデルの概要

hustvl/yolos-smallの Core ML版です。元モデルの著作権はその作者 (Huazhong University of Science & Technology / HuggingFaceチーム)に帰属します。

変換について

YOLOSのViT埋め込み層には、パターン7(Depth Anything)のDINOv2と同じ罠があります: 位置埋め込みの補間(InterpolateInitialPositionEmbeddingsと InterpolateMidPositionEmbeddings)が、トレース時には常にbicubic補間を強制する実装に なっています。本モデルの学習解像度である512x864固定入力では、この補間が数学的に恒等写像に なることを確認済み(補間前後の最大誤差0.0)のため、両方の補間をスキップしてから変換しています。

使い方(Python / coremltools)

import numpy as np
import coremltools as ct
from PIL import Image

mlmodel = ct.models.MLModel("yolos-small_fp16.mlpackage")

image = Image.open("photo.jpg").convert("RGB").resize((864, 512))  # (width, height)
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32)  # (1, 3, 512, 864), NCHW

out = mlmodel.predict({"pixel_values": pixel_values})
logits = out["logits"][0]       # (100, 92): COCO 91クラス + 「該当なし」
pred_boxes = out["pred_boxes"][0]  # (100, 4): normalized (center_x, center_y, width, height)

入力は固定サイズ512x864(height, width)、NCHW。

精度検証

PyTorch fp32リファレンスと、COCO検証画像1枚で比較(実際に検出された5件に対して):

項目 結果
logitsコサイン類似度 0.99975
ラベル一致率 100%
boxes最大絶対誤差 0.0039(0〜1に正規化)

Specs

Item Value
ベースモデル hustvl/yolos-small(ViT-Small、30.7M params)
精度 float16
入力 512x864、NCHW、固定サイズ
フレームワーク Core ML(mlprogram、minimum_deployment_target=macOS14)

備考

  • 本変換は非公式のコミュニティ版です。
  • セキュリティー監査にはmodel-audit-liteを 使用しています(詳細はSECURITY.md)。

セキュリティー

model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。

Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masahiroid/yolos-small-coreml

Quantized
(4)
this model