yolos-small-mlx

English | 日本語

Model Summary

This is an unofficial MLX conversion of hustvl/yolos-small (an object detector that's essentially a plain ViT with learned detection tokens, COCO 91 classes). All credit for the original model goes to its authors (Huazhong University of Science & Technology / HuggingFace team).

This cannot be loaded with mlx-vlm

CLS/detection tokens plus per-layer "mid" position embeddings aren't in mlx-vlm's list of supported architectures, so this was reimplemented from scratch for MLX and requires the bundled yolos_mlx.py. There's no DETR-style decoder -- the last 100 detection tokens from the ViT encoder's output are fed directly into classification/bbox-regression heads, a comparatively simple structure.

Usage

import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np

from yolos_mlx import YolosMLX, IMAGE_SIZE  # (512, 864) = (height, width)

model = YolosMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())

image = Image.open("photo.jpg").convert("RGB").resize((IMAGE_SIZE[1], IMAGE_SIZE[0]))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16))  # (1, 512, 864, 3), NHWC

logits, boxes = model(pixel_values)  # logits: (1, 100, 92), boxes: (1, 100, 4) normalized cxcywh

COCO's 91 classes plus "no object" (class 91). boxes are (center_x, center_y, width, height) normalized to [0, 1]. Only the fixed 512x864 input size is supported (other resolutions haven't been verified).

Accuracy

Compared against the PyTorch fp32 reference on a COCO validation image (restricted to the 5 actual detections):

Precision Logits cosine sim. Boxes cosine sim. Label agreement
MLX fp32 1.0 1.0 100%
MLX fp16 (this release) 1.0 1.0 100%

Specs

Item Value
Base model hustvl/yolos-small (ViT-Small, 30.7M params)
Precision float16
Input 512x864, NHWC, fixed size
Framework MLX (from-scratch yolos_mlx.py)

Notes

  • This is a community conversion, not an official release from the YOLOS authors.
  • Security audit uses model-audit-lite (see SECURITY.md for details).

Security

Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.


モデルの概要

hustvl/yolos-small(学習済み検出トークンを 使う、ViTそのままの構造による物体検出モデル。COCO 91クラス)の MLX版です。元モデルの著作権はその作者 (Huazhong University of Science & Technology / HuggingFaceチーム)に帰属します。

mlx-vlmでは読み込めません

CLS/検出トークン + 層ごとの"mid" position embeddingという構成はmlx-vlmの対応アーキテクチャ 一覧に含まれていないため、MLXでの実装をゼロから書き起こして変換しています。同梱の yolos_mlx.pyが必要です。DETRのようなデコーダーは無く、ViTエンコーダーの出力のうち末尾100個の 検出トークンをそのまま分類・bbox回帰ヘッドに通すだけの、比較的シンプルな構造です。

使い方

import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np

from yolos_mlx import YolosMLX, IMAGE_SIZE  # (512, 864) = (height, width)

model = YolosMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())

image = Image.open("photo.jpg").convert("RGB").resize((IMAGE_SIZE[1], IMAGE_SIZE[0]))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16))  # (1, 512, 864, 3), NHWC

logits, boxes = model(pixel_values)  # logits: (1, 100, 92), boxes: (1, 100, 4) normalized cxcywh

COCO 91クラス + 「該当なし」(クラス91)。boxesは(center_x, center_y, width, height)を 0〜1に正規化した値。固定サイズ512x864のみ対応(学習解像度と異なる入力は未検証)。

精度検証

PyTorch fp32リファレンスと、COCO検証画像1枚で比較(実際に検出された5件に対して):

精度 Logitsコサイン類似度 Boxesコサイン類似度 ラベル一致率
MLX fp32 1.0 1.0 100%
MLX fp16(本リリース) 1.0 1.0 100%

Specs

Item Value
ベースモデル hustvl/yolos-small(ViT-Small、30.7M params)
精度 float16
入力 512x864、NHWC、固定サイズ
フレームワーク MLX(ゼロから実装したyolos_mlx.py)

備考

  • 本変換は非公式のコミュニティ版です。
  • セキュリティー監査にはmodel-audit-liteを 使用しています(詳細はSECURITY.md)。

セキュリティー

model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
30.7M params
Tensor type
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masahiroid/yolos-small-mlx

Finetuned
(10)
this model