lightd v4

A tiny traffic-light detector for comma cameras. Finds traffic lights and reads their color and arrow shape.

params compute format validated license

lightd v4 on comma's public San Diego demo drive: greens, a red, and red left-turn arrows, colors read correctly

comma's public demo drive, downtown San Diego. Thick boxes = lights the model would act on (score ≥ 0.54). Thin boxes = lights it sees but scores lower. Note the two red left arrows and the greens, all read correctly.

12-second clip of lightd v4 approaching a signalized intersection in San Diego

A 12-second clip from the same drive.

TL;DR

What it does Finds traffic lights in a forward comma camera image and outputs box, color (red / yellow / green / off) and shape (circle / left / right / straight arrow).
Who made it The TurnPilot team at Sculptor AI, as part of an openpilot fork with navigation and traffic-light handling.
Model EfficientNet-B0 CenterNet, 3.7M params, 384×1344 input, stride-4 output. Distilled from a ConvNeXt-T teacher.
Validated on Daytime frames from the comma four narrow camera. Night, precipitation and other cameras are not validated (Scope).
Color accuracy 100% on every detected light in our 240-frame comma test (445 lights), with zero red-read-as-green errors.
Finds 94% of lights within 40 m on a 9-minute daytime drive, at ~0.6 false alarms per minute and no phantom reds.
Doesn't do Decide which light applies to your lane, output facing or stop-line (both untrained), or track lights over time.
License CC BY-NC-SA 4.0. Non-commercial. See licenses.

Scope. Validated only on daytime frames from the comma four narrow camera. No safety validation has been performed. See Scope and Limits.

Quick start

pip install onnxruntime numpy
import numpy as np
import onnxruntime as ort

THRESH = 0.54                      # operating point, see below
COLORS = ["red", "yellow", "green", "off"]
SHAPES = ["circle", "left", "right", "straight"]

sess = ort.InferenceSession("lightd_v4.onnx")

def detect(img):
    """img: float32 array, shape 1x3x384x1344, RGB, values 0..255 (lightd view, see below)."""
    out = sess.run(None, {"img": img})[0][0]          # 16 x 96 x 336, stride 4
    heat, peak = out[0], out[1]
    dets = []
    for y, x in zip(*np.nonzero(heat * peak >= THRESH)):
        dets.append({
            "cx": (x + out[4, y, x]) * 4, "cy": (y + out[5, y, x]) * 4,
            "w": np.exp(out[2, y, x]) * 4, "h": np.exp(out[3, y, x]) * 4,
            "score": float(heat[y, x]),
            "color": COLORS[int(out[6:10, y, x].argmax())],
            "shape": SHAPES[int(out[10:14, y, x].argmax())],
        })
    return dets

class Hold:
    """Only act on a light once it has been over the threshold in 2 frames in a row."""
    def __init__(self, dist_px=24):
        self.prev, self.dist = [], dist_px
    def update(self, dets):
        kept = [d for d in dets if any(abs(d["cx"] - p["cx"]) < self.dist and
                                       abs(d["cy"] - p["cy"]) < self.dist for p in self.prev)]
        self.prev = dets
        return kept

Camera geometry matters. The input must be in the "lightd view": the comma four narrow camera, focal length 1141.5 px, the band above the horizon, resized to 384×1344. The warp from the camera's YUV frame into this view isn't included here. On a different camera, re-export from best.pt with the matching focal length and size.

Operating point (please use this)

Detection threshold 0.54 on heat at peaks, plus the 2-frames-in-a-row hold shown above.

We tuned the threshold on our hand-labeled comma-camera test for about 95% precision per frame, then checked it per approach on the Bosch test drive: 9.3 minutes of continuous, fully labeled daytime video.

measure result at 0.54 + 2-frame hold
Lights within 40 m that we caught 49 of 52 (94%)
All approached lights (red, yellow, green) 76%
Distance at first catch median 58 m; a quarter of lights caught at 38 m or closer
False alarms (2+ frames, not a labeled light) ~0.6 per minute, all green or unlit (unlit heads, distant greens). Zero phantom reds.
Color at first catch 79 of 81 right. The two slips were yellow→red and green→off. Red was never read as green.
Why not a different threshold?
  • At 0.50, that drive produced 4 phantom reds in 9 minutes. At 0.45, it produced 24.
  • The "99% precision" calibrated cutoff (0.706) catches only 3% of lights on the same drive. Don't use it.
  • Averaging scores over 4, 8 or 15 frames did not beat the plain 2-frame hold. The false alarms are steady (unlit heads, distant greens), not flickery.
  • heat is not a calibrated probability. Real lights usually score 0.5 to 0.7, because each light's score is split across two neighboring grid cells (median 0.58 on the peak + 0.40 on its neighbor). The color probabilities are temperature-calibrated (expected calibration error 0.006), so put any confidence rule on those, not on heat.

Output layout

out is 1×16×96×336 (stride 4). Channel order:

idx channel meaning
0, 1 heat, is_peak detection score and local-maximum mask
2, 3 log_w, log_h box size: exp(·) × 4 px
4, 5 off_x, off_y sub-cell center offset: (cell + off) × 4 px
6–9 p_red, p_yellow, p_green, p_off color probabilities
10–13 p_circle, p_left, p_right, p_straight shape probabilities
14, 15 p_facing, stopline_over_20m untrained: ignore

The same metadata is embedded in the ONNX and in lightd_v4.json.

Results

Lights found at 95% and 99% precision on the comma-camera test, across model versions

Comma-camera test: 240 hand-labeled comma10k validation frames, 445 lights, from cars the model never trained on.

metric v4
AP 0.898
Recall at 95% precision 62% (63% at 94.6% precision when the cutoff is tuned on half the frames and measured on the other half)
Recall at 99% precision 38%
Color accuracy on detected lights 100%, 0 red→green
Shape accuracy 95.6%

Official test splits (LISA sequences, Bosch test, S2TLD test): AP 0.856. LISA's test labels miss many real lights, so its precision is understated.

Version history

Each step was measured on the same comma-camera test set. "Found" is recall at 95% precision.

model what changed AP found
Teacher 1 (ConvNeXt-T) public data only 0.841 30%
Student v2 (EfficientNet-B0) distilled from teacher 1 (a bug kept comma frames out of the mix) 0.837 40%
Teacher 3 + hand-cleaned LISA, pseudo-labeled comma10k, synthetic lamp bloom 0.896 71%
Student v3 distilled from teacher 3, comma frames actually in the mix 0.908 52%
Student v4 (this model) + 71 hand-checked comma yellows, + 129 streetlight hard negatives 0.898 62%

The teacher is better than the student on this scorecard. We ship the student because it's small enough to run on-device.

Teacher vs. student on the official test sets

How it was trained

  • Architecture: EfficientNet-B0 (efficientnet_b0.ra_in1k) CenterNet, distilled from a ConvNeXt-T teacher.
  • Data: LISA, Bosch Small Traffic Lights, S2TLD, plus comma10k frames pseudo-labeled by the teacher.
  • Fixing the public labels: LISA leaves a lot of real lights unlabeled, so the teacher's confident detections that match no label became "ignore" regions (about 9 in 10 are real lights).
  • Hand labeling: 71 comma yellows, 129 streetlight hard negatives, and a night pass of 205 of the model's most confident night detections that added 160 lights and 48 hard negatives.
  • Augmentation: a synthetic lamp bloom that imitates how the comma camera blows close lights out into glowing discs, plus copy-paste of rare lights (yellows, arrows).
Augmented training crops: lights rescaled to comma-camera size, with pasted-in rare lights

Training crops: lights rescaled to how they'd look on a comma four camera. Box color is the labeled state; soft rectangles are pasted-in rare lights.

What we learned

1. Public datasets are full of unlabeled lights

The teacher's most confident "false positives" on LISA were almost all real, lit lights that LISA never labeled. Training them as background taught the model to be under-confident.

Teacher false positives on LISA are mostly real unlabeled lights

Turning those detections into "ignore" regions fixed most of it:

Auto-ignore regions on LISA
2. Auto-labeling real comma frames: red and green worked, yellow didn't

comma10k is the only openly licensed set of real comma-camera frames. The teacher's red and green labels were nearly all correct. Its yellows were mostly sodium streetlights and diamond road signs, so we threw out all 198 and hand-checked replacements instead.

Teacher pseudo-labels: red Teacher pseudo-labels: yellow failures

3. The comma camera blows out close lights

The public training sets rarely show the big white-hot discs that a close light becomes on the comma camera, so the model missed lights it should have caught. Here is what it finds and misses on real comma frames at the 95%-precision cutoff (green border = found, red = missed, with score):

Found vs missed lights on comma frames
4. Night: tail lights and label gaps

At night the model scores the backs of cars and trucks (paired red tail lights) about as high as real signals. We hand-checked 205 of its most confident night detections: most were real lights LISA never labeled, and the rest (tail lights, neon and business signs, glare) went back into training as negatives. The comma test set is almost all daytime, so we can't yet measure night on comma cameras.

Night false positives Hand-checked night detections

5. Approaching a red Approaching a red light on the San Diego drive

The same drive, approaching a red. Close reds clear the bar; farther ones are seen but score below it until the car gets closer. (The close-up panels in this image come from a second-stage checker we tried and did not ship.)

More figures are in media/.

Scope

What was measured, and what wasn't.

dimension validated not validated
Camera comma four narrow camera, in the lightd view (focal 1141.5 px, horizon-up band, 384×1344). The sensor is assumed to be the OS04C10. Other cameras or focal lengths (re-export and re-validate), the wide camera.
Lighting Daytime. The 240-frame comma test and the 9.3-minute Bosch test drive are both daytime. Night: LISA night recall at 95% precision is 0.44, and there is no comma-camera night test. Direct sun in the lens, tunnels, dusk and dawn were not evaluated.
Weather Dry conditions in the test sets. Rain, snow, fog, wet-lens glare.
Geography Whatever LISA, Bosch, S2TLD and comma10k cover. No per-region evaluation, so we make no claim for any specific country or signal design.
Outputs Box, color and arrow shape. p_facing and stopline_over_20m (never trained). Lane relevance is not modeled, so the model does not say whether a detected light applies to your lane.
Temporal A 2-frame hold (see the operating point). Tracking, flicker handling (LED PWM) and longer-term state estimation.
Safety None. No safety analysis, failure-mode analysis or on-vehicle closed-loop validation has been done.

Limits

  • Night is weaker. LISA night recall at 95% precision is 0.44. The model scores paired red tail lights about as high as real signals.
  • p_facing and stopline_over_20m are untrained. Don't use them. Training them needs simulator data we didn't have in this run.
  • Distance. Far lights are seen but often score below 0.54 until the car closes in. Expect dependable detection at roughly 40 to 60 m, not 100 m.
  • Small, sparse training scenes. LISA's training set is 13 video clips and Bosch's is 5 recording sessions, so unfamiliar scenes will be harder.
  • Arrows. Shape accuracy is 95.6% on the comma test; the per-approach drive check covers color only.
  • Sensor assumption. view_focal 1141.5 assumes the comma four narrow camera is the OS04C10 sensor. If it's the OX03C10, re-export with the matching focal length.
  • Not yet verified: the tinygrad ONNX runner's output (eval/export_checks.json has onnxruntime only) and latency on comma hardware. onnxruntime and PyTorch agree to 7.8e-5 on a real comma frame.

Intended use

Intended for: research on traffic-light handling for driver-assistance systems, supervised test drives in daytime, benchmarks, fine-tuning on your own comma-camera data, and reading our write-up.

Out of scope: use without a human driver responsible for the vehicle, night or bad-weather driving, any commercial product (non-commercial license), and camera setups other than the comma narrow camera without re-export and re-validation.

Files

file what
lightd_v4.onnx Deployment graph (opset 17), operating-point metadata embedded.
lightd_v4.json The operating-point metadata as a sidecar.
best.pt PyTorch checkpoint (EMA weights) for re-export or fine-tuning.
config.yaml The training config.
eval/ Test reports, calibration and export checks. These tables are at the 99%-precision cutoff (0.706), not the recommended operating point.
media/ Figures and the preview clip used in this card.

Training data and licenses

dataset license used for
LISA Traffic Light Dataset CC BY-NC-SA 4.0 training and test
Bosch Small Traffic Lights non-commercial use only training and test
S2TLD MIT training and test
comma10k MIT frames pseudo-labeled by the teacher for training; 240 validation frames hand-labeled as our test set

Initial weights come from timm's efficientnet_b0.ra_in1k (ImageNet-pretrained).

These weights are released under CC BY-NC-SA 4.0: non-commercial use only, share alike, with attribution. LISA is CC BY-NC-SA 4.0 and Bosch's license is non-commercial and not for production, so weights trained on them inherit those terms. Check the terms again before any commercial use. The figures and preview clip here use only LISA, comma10k and comma's public demo drive; we don't redistribute Bosch images.

Citations

Please cite the datasets this model was trained on:

  • Jensen, M. B., Philipsen, M. P., Møgelmose, A., Moeslund, T. B., Trivedi, M. M. Vision for Looking at Traffic Lights: Issues, Survey, and Perspectives. IEEE T-ITS, 2016. (LISA)
  • Philipsen, M. P., Jensen, M. B., Møgelmose, A., Moeslund, T. B., Trivedi, M. M. Traffic Light Detection: A Learning Algorithm and Evaluations on Challenging Dataset. ITSC, 2015. (LISA)
  • Behrendt, K., Novak, L. A Deep Learning Approach to Traffic Lights: Detection, Tracking, and Classification. ICRA, 2017. (Bosch)
  • Yang, X., et al. SCRDet++: Detecting Small, Cluttered and Rotated Objects via Instance-Level Feature Denoising and Rotation Loss Smoothing. IEEE TPAMI, 2022. (S2TLD)
  • comma.ai. comma10k. https://github.com/commaai/comma10k

Acknowledgements

Built by the TurnPilot team at Sculptor AI. openpilot and comma are trademarks of comma.ai; this project is not affiliated with or endorsed by comma.ai.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Sculptor-AI/lightd-v4

Quantized
(2)
this model

Evaluation results

  • AP on comma10k hand-labeled test (240 frames, 445 lights)
    self-reported
    0.898
  • Recall at 95% precision on comma10k hand-labeled test (240 frames, 445 lights)
    self-reported
    0.620
  • AP on LISA sequences + Bosch test + S2TLD test
    self-reported
    0.856