Popcorn 1.0
A 0.91M-parameter guinea pig detector, named for the jump a happy guinea pig does. Distilled from OWLv2 to run on a Raspberry Pi 5 with a Hailo-8L. YOLOX-nano (LeakyReLU), one class, Apache-2.0.
| file | format | size |
|---|---|---|
popcorn-1.0.onnx |
fp32, 1x3x640x640 in, (1, 8400, 6) out, undecoded |
3.6 MB |
popcorn-1.0-hailo8l.hef |
Hailo-8L int8, DFC 3.33.0; nine raw head tensors | 7.4 MB |
Results
Measured on the compiled HEF on a Hailo-8L, against 107 hand-labelled frames from the deployment camera (64 with an animal, 35 without), under the live tiling.
| conf | presence recall | false positives | box on animal |
|---|---|---|---|
| 0.65 | 70.3% [58, 80] | 0 / 35 | 59.8% |
| 0.55 | 81.2% [70, 89] | 1 / 35 | 70.1% |
At 0/35 false positives, the previous YOLOv8n deployment (AGPL) scored 53.1% [41, 65]; paired McNemar p = 0.013. About 10 ms per 640px tile on the Hailo-8L.
Usage (ONNX)
Input is raw 0-255 BGR, letterboxed top-left with 114 padding โ no /255.
import cv2, numpy as np, onnxruntime as ort
def detect(path, onnx="popcorn-1.0.onnx", conf=0.65, iou=0.45):
img = cv2.imread(path)
h, w = img.shape[:2]; r = min(640 / h, 640 / w)
pad = np.full((640, 640, 3), 114, np.uint8)
pad[:int(h * r), :int(w * r)] = cv2.resize(img, (int(w * r), int(h * r)))
out = ort.InferenceSession(onnx).run(None, {"images": pad.transpose(2, 0, 1)[None].astype(np.float32)})[0][0]
grid, stride = [], []
for s in (8, 16, 32):
g = 640 // s; yv, xv = np.meshgrid(np.arange(g), np.arange(g), indexing="ij")
grid.append(np.stack([xv, yv], -1).reshape(-1, 2)); stride.append(np.full((g * g, 1), s))
grid, stride = np.concatenate(grid), np.concatenate(stride)
xy, wh = (out[:, :2] + grid) * stride, np.exp(out[:, 2:4]) * stride
score = out[:, 4] * out[:, 5]
boxes = np.concatenate([xy - wh / 2, xy + wh / 2], 1) / r
k = score >= conf
idx = cv2.dnn.NMSBoxes(boxes[k].tolist(), score[k].tolist(), conf, iou)
return [(boxes[k][i].tolist(), float(score[k][i])) for i in np.array(idx).flatten()]
The HEF takes the same pixels as uint8 NHWC, and decoding runs on the host.
Notes
- Why LeakyReLU: with YOLOX's default SiLU, int8 on the Hailo collapsed (output correlation 0.12โ0.64, 91% false positives). LeakyReLU gave 0.89โ0.98 at the same fp32 accuracy. If you retrain for an NPU, score the compiled model, not the checkpoint.
- Limits: one enclosure, one camera, three animals, daylight frames, a small test set. People are out of scope, and it has fired on a dark-haired head. It can't see through hides. It is not a health monitor.
- Training: OWLv2 pseudo-labels on 640px tiles. Ambiguous tiles were dropped rather than kept as negatives; deleting suspect boxes instead cost 36 points of recall. 7,133 tiles, 30 epochs, about 42 minutes on one GPU. The test frames are excluded by an enforced test.
- Licence: Apache-2.0 (YOLOX by Megvii, OWLv2 by Google). The training data is images of a private home and is not released.