StreamPETR — nuCarla (Town04)

Camera-only 3D object detection for autonomous driving. Six surround-view cameras in, 3D bounding boxes out, with a temporal memory carried across frames. Trained on the nuCarla dataset, Town04 subset.

Ships as a PyTorch checkpoint plus ready-to-run ONNX graphs and TensorRT FP16 engines.

Architecture StreamPETR, ResNet-50 + CPFPN
Classes car, truck, bus, motorcycle, bicycle, pedestrian
Input 6 cameras, 1600×900 → 704×256
mAP @0.25 3D IoU 0.5120 (nuCarla Town04 val)
Latency 4.5 ms/frame, RTX 5070 Ti FP16 (~220 fps)
Init nuScenes-pretrained, class head sliced to 6 classes

Demo

demo/carla_live_detection.mp4 — the model run over a clip recorded live in CARLA 0.9.16: an ego vehicle driving on autopilot through 150 spawned vehicles and 16 pedestrians in Town04. Six camera views with projected 3D boxes plus a bird's-eye panel. Averages 6.7 detections per frame.

The clip is unlabelled, so it is a qualitative check, not a measured score.

Checkpoints

File Size mAP @0.25 What it is
checkpoints/best_mAP_0.25_iter_8512.pth 432M 0.5120 Best epoch (28). Everything below — ONNX, engines, the demo video, all reported numbers — comes from this one.
checkpoints/final_iter_13376_epoch44.pth 432M 0.4887 Last checkpoint of the run. mAP had plateaued; kept for continuing training.
checkpoints/init_nuscenes_6class_sliced.pth 148M Starting point: the nuScenes StreamPETR weights with the class head sliced from 10 to 6. Use this to reproduce training.

Intermediate epoch checkpoints are not published — they are rotation artifacts from max_keep_ckpts, not separately evaluated models.

Files

checkpoints/                               the three above
stream_petr_r50_nucarla_town04.py          training config
DEPLOYMENT.md                              full inference guide — read this
demo/carla_live_detection.mp4         12M  detection demo
deployment/
  stream_petr_6cam_encoder.onnx        95M
  stream_petr_6cam_temporal_head.onnx  47M
  *_fp16.engine                             TensorRT 10.9, RTX 5070 Ti
  streampetr_trt_runner.py                  runtime, handles temporal memory
  carla_stream_petr_postprocess.py          decoder
  build_streampetr_engines.py               rebuild engines for another GPU
  test_inputs/, reference_outputs/          one frame, to verify your setup

Quick start (TensorRT)

pip install tensorrt==10.9.0.34 numpy torch
import numpy as np
from streampetr_trt_runner import StreamPETRRunner
from carla_stream_petr_postprocess import decode_stream_petr

runner = StreamPETRRunner(
    "deployment/stream_petr_6cam_encoder_fp16.engine",
    "deployment/stream_petr_6cam_temporal_head_fp16.engine",
)

runner.reset()                       # call at the start of every sequence
for frame in sequence:
    out = runner(
        images=frame.images,         # [1, 6, 3, 256, 704] float32, normalized
        timestamp=frame.seconds,
        ego_pose=frame.lidar2global,       # [1, 4, 4]
        ego_pose_inv=frame.global2lidar,   # [1, 4, 4]
    )
    det = decode_stream_petr(
        out["class_logits"], out["bbox_predictions"], score_threshold=0.35
    )

The engines are tied to the GPU and TensorRT version they were built on. Rebuild from the ONNX for anything else — it takes about 15 seconds:

python deployment/build_streampetr_engines.py --deployment-dir deployment --cameras 6

Camera rig — must match

The 3D position embedding is baked into the exported graph. Images from a different camera layout produce meaningless geometry with no error raised. Offsets are ego-frame metres (x forward, y left, z up), yaw counter-clockwise:

Camera x y z yaw
CAM_FRONT 1.901 0.016 1.511 0.3°
CAM_FRONT_LEFT 1.524 0.495 1.509 55.2°
CAM_FRONT_RIGHT 1.551 −0.493 1.496 −56.4°
CAM_BACK 0.028 0.003 1.579 179.9°
CAM_BACK_LEFT 1.036 0.485 1.591 108.6°
CAM_BACK_RIGHT 1.015 −0.481 1.562 −110.8°

All six share 1600×900 at 65° horizontal FOV (fx = fy = 1255.7484, cx = 800, cy = 450) — what CARLA emits for image_size_x=1600, image_size_y=900, fov=65.

Preprocessing, tensor shapes and the temporal memory protocol are all in DEPLOYMENT.md.

Results

nuCarla Town04 validation, 29 scenes / 1160 frames:

Class AP @0.25 IoU
car 0.630
bus 0.623
truck 0.612
bicycle 0.358
motorcycle 0.317
pedestrian 0.018
mAP 0.5120

Pedestrian scores near zero under 3D IoU because the metric is brutal on small boxes: a 0.42 m wide pedestrian tolerates ~0.25 m of lateral error to reach IoU 0.25, against ~1.2 m for a car. Under the nuScenes center-distance metric the paper actually reports, pedestrian AP is 0.309 and overall mAP 0.4638 — the model does detect pedestrians.

Not comparable to the nuCarla paper's 0.745

The paper reports mAP 0.745 for PETR, but under different conditions:

Paper (PETR) This model
Training data 700 scenes / 7 maps 91 scenes / 1 map
Backbone VoVNet ResNet-50
Resolution 1600×640 704×256 (5.7× fewer pixels)
GPU-hours 150 ~5
Metric center-distance 3D IoU

Town04 is also CARLA's highway map, where objects sit far from the ego (cars average 50 m away), making it harder than the urban maps in the full set.

Training

  • Initialized from the official nuScenes StreamPETR checkpoint. All six target classes exist in nuScenes' ten, so the classification rows were sliced across rather than randomly re-initialized — nothing in the network starts from scratch. This lifted first-epoch mAP from 0.0497 to 0.2701.
  • 50 epochs planned, batch 12, AdamW at 1.2e-4 with cosine annealing, FP16. The run reached epoch 45; mAP had plateaued since epoch 23 and the LR was annealed to 3.6e-6.
  • Best checkpoint at iteration 8512 (epoch 28).

Verification

The ONNX wrapper reproduces the native PyTorch head exactly (max |Δ| = 0.0 on both outputs). TensorRT FP16 against PyTorch FP32 on the bundled frame:

tensor correlation rel. error
image_features 0.999996 2.9e-3
bbox_predictions 0.999992 2.8e-2
class_logits 0.999546 7.6e-2

Decoded detections agree: 17 vs 17 above score 0.2, every one placed within 0.02–0.10 m. Two pairs swap rank order, both between detections whose scores tie to within 0.003.

Known FP16 effect. 768 of the 1024 memory slots are bit-exact frame to frame; within the 256 slots selected fresh each frame, FP16 scoring picks a slightly different subset near the top-256 cut-off. Its accumulation over long sequences has not been measured — build the FP32 engine if you need to rule it out.

Limitations

  • Trained on one map (Town04). Expect degradation on urban maps or dense intersections.
  • Pedestrian localization is weak in absolute terms; suitable for perception research, not safety-critical use.
  • Simulation only — no real-world images were used, and no sim-to-real transfer has been evaluated.
  • The rig is fixed. Any change to camera placement, count or intrinsics requires re-exporting from the checkpoint.

Citation

@article{nucarla,
  title={nuCarla: A nuScenes-Style Bird's-Eye View Perception Dataset for CARLA Simulation},
  author={Qiao, Zhijie and Cao, Zhong and Liu, Henry X.},
  year={2025},
  url={https://arxiv.org/abs/2511.13744}
}

@article{streampetr,
  title={Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection},
  author={Wang, Shihao and Liu, Yingfei and Wang, Tiancai and Li, Ying and Zhang, Xiangyu},
  journal={arXiv preprint arXiv:2303.11926},
  year={2023}
}

License

Apache 2.0, following StreamPETR and the nuCarla dataset.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train shin0412/StreamPETR_nuCarla

Papers for shin0412/StreamPETR_nuCarla