RampNet Curb Ramp Detection Model

Stage-2 model from RampNet: A Two-Stage Pipeline for Bootstrapping Curb Ramp Detection in Streetscape Images from Open Government Metadata (O'Meara et al., ICCV'25 CV4A11y workshop, arXiv:2508.09415).

Takes a 2048x4096 equirectangular street-view panorama (ImageNet-normalized) and predicts a 512x1024 heatmap of curb ramp locations. Extract detections with skimage.feature.peak_local_max.

Provenance

Field Value
Training code https://github.com/ProjectSidewalk/RampNet @ 842ec33
Source checkpoint epoch_1_step_9378.pth (sha256 prefix b0c3ff7a10fc)
Training dataset projectsidewalk/rampnet-dataset revision main
Exported 2026-07-24 by scripts/export_hf_model.py

The model weights are unchanged from the originally published artifact — this revision only fixes the packaging (a transformers-compatible wrapper; see Requirements) and corrects the evaluation numbers (see Erratum). The prior artifact remains addressable at its tagged Hub revision.

Evaluation (1,000-panorama manually labeled gold set)

Metric Value
Average Precision (interpolated, full sweep) 0.9205
Precision @ threshold 0.55 0.9492
Recall @ threshold 0.55 0.8727
Ground-truth points 3919
Matching radius (normalized) 0.022
Flip TTA on

Important: these numbers were measured with horizontal-flip test-time augmentation (evaluate the original and mirrored panorama, combine heatmaps with elementwise max). Single-pass inference will land somewhat below them; derive your own threshold curve without TTA before choosing an operating point. Average Precision is the interpolated area under the full precision–recall sweep; precision and recall are reported at the operating threshold.

Erratum: Evaluation Metrics

After publication we found that the evaluation protocol described in §3.3 of the paper (and implemented at tag v1.0-iccv2025) differs from standard detection evaluation in two ways that bias precision and recall upward: predictions matching an already-claimed ground-truth point were counted as additional true positives rather than false positives, and one prediction could satisfy multiple ground-truth points. The numbers above use corrected one-to-one matching. For transparency, the originally published figures are kept here alongside the corrected ones:

Metric Originally published Corrected (this revision)
Precision @ 0.55 0.938 0.949
Recall @ 0.55 0.935 0.873
Average Precision 0.9236 0.9205

Swapping only the matching rule on identical model outputs moves precision by −1.0 points and recall by −4.4 points; the remainder is reproduction drift (environment, JPEG re-encoding). The comparison with prior work is unaffected (both systems were scored under the same protocol, and the gap is far larger than the correction). Full write-up: docs/eval_protocol_verification.html and the repository README's erratum.

Choosing a detection threshold

  • 0.55 is the recommended default operating point (see evaluation above).
  • Sweeping thresholds: run stage_two/evaluate.py in the training repo, which emits full precision/recall-vs-confidence curves as CSV.
  • Per-city deployments should calibrate on ~100 locally labeled panoramas; see the repo README's "Choosing a Detection Threshold" section.

Requirements

Loads via trust_remote_code, so a compatible transformers is required:

  • transformers >= 5.13 — supported (the custom-model auto-class contract this version introduced is satisfied; see issue #19).
  • transformers 5.12.x — supported.

Both are covered by a round-trip load smoke test (tests/test_hf_load.py in the training repo). Earlier revisions of this model fail to load on transformers >= 5.13 with AttributeError: ... 'KeypointModel' has no attribute 'register_for_auto_class' (or ... has no attribute 'generate') — pull the current revision.

Usage

For a step-by-step walkthrough, see the Google Colab notebook, which adds a visualization to the code below.

import torch
from transformers import AutoModel
from PIL import Image
import numpy as np
from torchvision import transforms
from skimage.feature import peak_local_max

DEVICE = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = AutoModel.from_pretrained("projectsidewalk/rampnet-model", trust_remote_code=True).to(DEVICE).eval()

preprocess = transforms.Compose([
    transforms.Resize((2048, 4096), interpolation=transforms.InterpolationMode.BILINEAR),
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])

img = Image.open("panorama.jpg").convert("RGB")
with torch.no_grad():
    heatmap = model(preprocess(img).unsqueeze(0).to(DEVICE)).squeeze().cpu().numpy()

peaks = peak_local_max(np.clip(heatmap, 0, 1), min_distance=10, threshold_abs=0.55)
scale_w, scale_h = img.width / heatmap.shape[1], img.height / heatmap.shape[0]
print([(int(c * scale_w), int(r * scale_h)) for r, c in peaks])

Citation

@inproceedings{omeara2025rampnet,
  author    = {John S. O'Meara and Jared Hwang and Zeyu Wang and Michael Saugstad and Jon E. Froehlich},
  title     = {{RampNet: A Two-Stage Pipeline for Bootstrapping Curb Ramp Detection in Streetscape Images from Open Government Metadata}},
  booktitle = {{ICCV'25 Workshop on Vision Foundation Models and Generative AI for Accessibility: Challenges and Opportunities (ICCV 2025 Workshop)}},
  year      = {2025},
  doi       = {https://doi.org/10.48550/arXiv.2508.09415},
}
Downloads last month
105
Safetensors
Model size
90.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for projectsidewalk/rampnet-model

Finetuned
(2)
this model

Dataset used to train projectsidewalk/rampnet-model

Collection including projectsidewalk/rampnet-model

Paper for projectsidewalk/rampnet-model