Popcorn 1.0

A 0.91M-parameter guinea pig detector, named for the jump a happy guinea pig does. Distilled from OWLv2 to run on a Raspberry Pi 5 with a Hailo-8L. YOLOX-nano (LeakyReLU), one class, Apache-2.0.

file format size
popcorn-1.0.onnx fp32, 1x3x640x640 in, (1, 8400, 6) out, undecoded 3.6 MB
popcorn-1.0-hailo8l.hef Hailo-8L int8, DFC 3.33.0; nine raw head tensors 7.4 MB

Results

Measured on the compiled HEF on a Hailo-8L, against 107 hand-labelled frames from the deployment camera (64 with an animal, 35 without), under the live tiling.

conf presence recall false positives box on animal
0.65 70.3% [58, 80] 0 / 35 59.8%
0.55 81.2% [70, 89] 1 / 35 70.1%

At 0/35 false positives, the previous YOLOv8n deployment (AGPL) scored 53.1% [41, 65]; paired McNemar p = 0.013. About 10 ms per 640px tile on the Hailo-8L.

Usage (ONNX)

Input is raw 0-255 BGR, letterboxed top-left with 114 padding โ€” no /255.

import cv2, numpy as np, onnxruntime as ort

def detect(path, onnx="popcorn-1.0.onnx", conf=0.65, iou=0.45):
    img = cv2.imread(path)
    h, w = img.shape[:2]; r = min(640 / h, 640 / w)
    pad = np.full((640, 640, 3), 114, np.uint8)
    pad[:int(h * r), :int(w * r)] = cv2.resize(img, (int(w * r), int(h * r)))
    out = ort.InferenceSession(onnx).run(None, {"images": pad.transpose(2, 0, 1)[None].astype(np.float32)})[0][0]
    grid, stride = [], []
    for s in (8, 16, 32):
        g = 640 // s; yv, xv = np.meshgrid(np.arange(g), np.arange(g), indexing="ij")
        grid.append(np.stack([xv, yv], -1).reshape(-1, 2)); stride.append(np.full((g * g, 1), s))
    grid, stride = np.concatenate(grid), np.concatenate(stride)
    xy, wh = (out[:, :2] + grid) * stride, np.exp(out[:, 2:4]) * stride
    score = out[:, 4] * out[:, 5]
    boxes = np.concatenate([xy - wh / 2, xy + wh / 2], 1) / r
    k = score >= conf
    idx = cv2.dnn.NMSBoxes(boxes[k].tolist(), score[k].tolist(), conf, iou)
    return [(boxes[k][i].tolist(), float(score[k][i])) for i in np.array(idx).flatten()]

The HEF takes the same pixels as uint8 NHWC, and decoding runs on the host.

Notes

  • Why LeakyReLU: with YOLOX's default SiLU, int8 on the Hailo collapsed (output correlation 0.12โ€“0.64, 91% false positives). LeakyReLU gave 0.89โ€“0.98 at the same fp32 accuracy. If you retrain for an NPU, score the compiled model, not the checkpoint.
  • Limits: one enclosure, one camera, three animals, daylight frames, a small test set. People are out of scope, and it has fired on a dark-haired head. It can't see through hides. It is not a health monitor.
  • Training: OWLv2 pseudo-labels on 640px tiles. Ambiguous tiles were dropped rather than kept as negatives; deleting suspect boxes instead cost 36 points of recall. 7,133 tiles, 30 epochs, about 42 minutes on one GPU. The test frames are excluded by an enforced test.
  • Licence: Apache-2.0 (YOLOX by Megvii, OWLv2 by Google). The training data is images of a private home and is not released.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support