DANIEL student depth: attention-free Depth Anything V2 for the Hailo-8L

A 24.5 M-parameter ResNet-34 U-Net distilled from Depth Anything V2-Large, built so that every operation quantises cleanly to int8 on the Raspberry Pi 5 + Hailo-8L (Raspberry Pi AI Kit). It predicts relative inverse depth (disparity) at 288×512.

From the paper Seeing to Drive on a Single Board (Itai M., 2026): project page · paper · dataset

Why it exists

The Hailo-8L runs attention in int8. At the 449 tokens of Depth Anything V2-Small at 224×392, 91 % of softmax weights fall below one int8 step and carry 17-33 % of the attention mass, so int8 transformer builds (ours and Hailo's Model Zoo build) erase obstacles. This student has no attention, no LayerNorm and no softmax.

Measured (400 held-out sim frames with metric GT, 100 real robot frames; on the real Hailo-8L)

Model on the Hailo-8L AbsRel all AbsRel obstacles δ1 r vs DAv2-L (real) FPS (Pi 5, batch 1)
student v5 (this repo, deployed) 0.269 0.197 0.669 0.996 27.1
student v3 (this repo) 0.269 0.204 0.674 0.995 27.1
Depth Anything V2-S, Hailo Model Zoo (224²) 0.354 0.334 0.593 0.979 21.6
Depth Anything V2-S, our int8 build (224×392) 0.477 0.330 0.423 0.982 16.6
SC-DepthV3, Hailo Model Zoo 0.414 0.352 0.506 0.950 53.6
FastDepth, Hailo Model Zoo 0.530 0.767 0.344 0.709 121.3
  • int8 = float on the chip: HEF vs the same weights in float, Pearson p50 0.99993 (v5), 0.99995 (v3), n = 500.
  • 35.1 ms on the NPU, 37.0 ms per frame end to end; 455 ms on the Pi 5's CPU alone; 4 ms on an RTX 5090.
  • Float references (RTX 5090): DAv2-L teacher 0.284 / 0.144 at 175 ms per frame; DAv2-S 0.283 / 0.150.

Files

File What
student_v5.pt, student_v3.pt PyTorch state dicts for student_depth.build()
student_v5_288x512.onnx, student_v3_288x512.onnx ONNX, input 1×3×288×512 RGB in [0, 1]
student_v5_288x512_hailo8l.hef, student_v3_288x512_hailo8l.hef Hailo-8L HEFs, input uint8 RGB 288×512 (normalisation inside)
student_depth.py, student_hailo.alls model definition, training/export code, the Hailo compile recipe

Versions: v1 distilled from DAv2-L teacher labels; v3 fine-tuned on rendered ground-truth disparity (simulated range bias −30 % → −7 %); v5 = v3 with the loss ×4 below the horizon (fewer near-floor phantoms). v5 and v3 tie in closed-loop navigation with simulator pose (7.9 vs 7.9 of 10, ten seeds); v5 leads on the robot's own pose (7.7 vs 7.1, not significant).

Use

import torch, cv2, numpy as np
from student_depth import build                          # in this repo
net = build().eval(); net.load_state_dict(torch.load("student_v5.pt", map_location="cpu"))
rgb = cv2.cvtColor(cv2.imread("frame.jpg"), cv2.COLOR_BGR2RGB)
x = torch.from_numpy(cv2.resize(rgb, (512, 288), interpolation=cv2.INTER_AREA)).permute(2, 0, 1)[None].float() / 255
with torch.no_grad():
    disparity = net(x)[0].numpy()                         # relative inverse depth, 288x512 (near = high)

On a Raspberry Pi 5 + Hailo-8L with HailoRT (the call sequence the project's onboard code uses):

import cv2, numpy as np
from hailo_platform import FormatType, HailoSchedulingAlgorithm, VDevice
p = VDevice.create_params(); p.scheduling_algorithm = HailoSchedulingAlgorithm.ROUND_ROBIN   # required, see below
vdev = VDevice(p)
im = vdev.create_infer_model("student_v5_288x512_hailo8l.hef"); im.set_batch_size(1)
im.output().set_format_type(FormatType.FLOAT32)
cim = im.configure(); out = np.empty(im.output().shape, np.float32)
rgb = cv2.cvtColor(cv2.imread("frame.jpg"), cv2.COLOR_BGR2RGB)
x = np.ascontiguousarray(cv2.resize(rgb, (512, 288), interpolation=cv2.INTER_AREA))   # uint8 RGB, normalisation is in the HEF
b = cim.create_bindings(); b.input().set_buffer(x); b.output().set_buffer(out)
cim.wait_for_async_ready(timeout_ms=2000); cim.run_async([b]).wait(2000)
disparity = out.reshape(288, 512)                                    # 35 ms on the NPU

Without the ROUND_ROBIN scheduler the async API returns instantly with an untouched (all-zero) buffer and no error. The output is relative (unknown scale and shift); the paper turns it into metric obstacle ranges with ground-contact geometry (Z = f_y·h / (v − c_y) at the row where an object meets the floor). The onboard code is available on request.

Limits

Trained and benchmarked on one Gaussian-splat reconstruction of a real yard plus real robot frames; the real frames have no ground truth (agreement with the teacher is not accuracy). Thin, low structure (cone plates, feet) below about 5 cm is not resolved. Licence: CC-BY-NC-4.0, because the Depth Anything V2-Large teacher is CC-BY-NC-4.0.

Citation

@misc{itaim2026seeing,
  title  = {Seeing to Drive on a Single Board: Attention-Free Distilled Depth on a Raspberry Pi 5 + Hailo-8L
            for Onboard Obstacle Avoidance, Planning and SLAM},
  author = {M., Itai},
  year   = {2026},
  url    = {https://itaim18.github.io/seeing-to-drive/}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train itaimizlish/daniel-student-depth