Instructions to use masahiroid/yolos-small-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use masahiroid/yolos-small-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download masahiroid/yolos-small-mlx --local-dir yolos-small-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
yolos-small-mlx
Model Summary
This is an unofficial MLX conversion of hustvl/yolos-small (an object detector that's essentially a plain ViT with learned detection tokens, COCO 91 classes). All credit for the original model goes to its authors (Huazhong University of Science & Technology / HuggingFace team).
This cannot be loaded with mlx-vlm
CLS/detection tokens plus per-layer "mid" position embeddings aren't in
mlx-vlm's list of supported architectures, so this was reimplemented
from scratch for MLX and requires the bundled yolos_mlx.py. There's no
DETR-style decoder -- the last 100 detection tokens from the ViT encoder's
output are fed directly into classification/bbox-regression heads, a
comparatively simple structure.
Usage
import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np
from yolos_mlx import YolosMLX, IMAGE_SIZE # (512, 864) = (height, width)
model = YolosMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())
image = Image.open("photo.jpg").convert("RGB").resize((IMAGE_SIZE[1], IMAGE_SIZE[0]))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16)) # (1, 512, 864, 3), NHWC
logits, boxes = model(pixel_values) # logits: (1, 100, 92), boxes: (1, 100, 4) normalized cxcywh
COCO's 91 classes plus "no object" (class 91). boxes are
(center_x, center_y, width, height) normalized to [0, 1]. Only the fixed
512x864 input size is supported (other resolutions haven't been verified).
Accuracy
Compared against the PyTorch fp32 reference on a COCO validation image (restricted to the 5 actual detections):
| Precision | Logits cosine sim. | Boxes cosine sim. | Label agreement |
|---|---|---|---|
| MLX fp32 | 1.0 | 1.0 | 100% |
| MLX fp16 (this release) | 1.0 | 1.0 | 100% |
Specs
| Item | Value |
|---|---|
| Base model | hustvl/yolos-small (ViT-Small, 30.7M params) |
| Precision | float16 |
| Input | 512x864, NHWC, fixed size |
| Framework | MLX (from-scratch yolos_mlx.py) |
Notes
- This is a community conversion, not an official release from the YOLOS authors.
- Security audit uses model-audit-lite
(see
SECURITY.mdfor details).
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
hustvl/yolos-small(学習済み検出トークンを 使う、ViTそのままの構造による物体検出モデル。COCO 91クラス)の MLX版です。元モデルの著作権はその作者 (Huazhong University of Science & Technology / HuggingFaceチーム)に帰属します。
mlx-vlmでは読み込めません
CLS/検出トークン + 層ごとの"mid" position embeddingという構成はmlx-vlmの対応アーキテクチャ
一覧に含まれていないため、MLXでの実装をゼロから書き起こして変換しています。同梱の
yolos_mlx.pyが必要です。DETRのようなデコーダーは無く、ViTエンコーダーの出力のうち末尾100個の
検出トークンをそのまま分類・bbox回帰ヘッドに通すだけの、比較的シンプルな構造です。
使い方
import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np
from yolos_mlx import YolosMLX, IMAGE_SIZE # (512, 864) = (height, width)
model = YolosMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())
image = Image.open("photo.jpg").convert("RGB").resize((IMAGE_SIZE[1], IMAGE_SIZE[0]))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16)) # (1, 512, 864, 3), NHWC
logits, boxes = model(pixel_values) # logits: (1, 100, 92), boxes: (1, 100, 4) normalized cxcywh
COCO 91クラス + 「該当なし」(クラス91)。boxesは(center_x, center_y, width, height)を
0〜1に正規化した値。固定サイズ512x864のみ対応(学習解像度と異なる入力は未検証)。
精度検証
PyTorch fp32リファレンスと、COCO検証画像1枚で比較(実際に検出された5件に対して):
| 精度 | Logitsコサイン類似度 | Boxesコサイン類似度 | ラベル一致率 |
|---|---|---|---|
| MLX fp32 | 1.0 | 1.0 | 100% |
| MLX fp16(本リリース) | 1.0 | 1.0 | 100% |
Specs
| Item | Value |
|---|---|
| ベースモデル | hustvl/yolos-small(ViT-Small、30.7M params) |
| 精度 | float16 |
| 入力 | 512x864、NHWC、固定サイズ |
| フレームワーク | MLX(ゼロから実装したyolos_mlx.py) |
備考
- 本変換は非公式のコミュニティ版です。
- セキュリティー監査にはmodel-audit-liteを
使用しています(詳細は
SECURITY.md)。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
Quantized
Model tree for masahiroid/yolos-small-mlx
Base model
hustvl/yolos-small