File size: 3,642 Bytes
22e1f7f
 
 
 
 
 
 
 
 
 
 
 
 
 
48a216d
22e1f7f
48a216d
 
22e1f7f
48a216d
 
 
22e1f7f
 
 
48a216d
22e1f7f
48a216d
22e1f7f
 
 
 
 
48a216d
22e1f7f
48a216d
22e1f7f
 
 
48a216d
22e1f7f
48a216d
 
 
22e1f7f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
48a216d
22e1f7f
 
 
 
 
 
 
 
 
48a216d
22e1f7f
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
---
library_name: pytorch
pipeline_tag: object-detection
tags:
  - object-detection
  - faster-rcnn
  - torchvision
  - license-plate-detection
  - house-number-detection
  - anonymization
  - privacy
  - street-view
---

<p align="center">
  <img src="assets/banner.png" width="1200" alt="Vision-AN18 banner" />
</p>

# Vision-AN18 — Faster R-CNN

Privacy-preserving computer vision for detecting and anonymizing sensitive visual information in street images.

Vision-AN18 detects **license plates** and **house numbers** in street-level images (Google Street View style), then blurs the detected regions to preserve privacy while keeping the rest of the image useful.

## Examples

Original image on the left, output of the anonymization pipeline on the right (Gaussian blur applied to each detected box).

| Original | Anonymized |
| :---: | :---: |
| ![Plate original](examples/inputs/plate.jpg) | ![Plate anonymized](examples/outputs/plate_anonymized.jpg) |
| ![House and plate original](examples/inputs/house_and_plate.png) | ![House and plate anonymized](examples/outputs/house_and_plate_anonymized.png) |
| ![Street original](examples/inputs/street.png) | ![Street anonymized](examples/outputs/street_anonymized.png) |

### Raw detection

<p align="center">
  <img src="assets/detection_example.png" width="700" alt="License plate detection example" />
</p>

> ⚠️ **Known limitations** — license plate detection is reliable. House number detection is still imperfect (scale bias inherited from the SVHN dataset used for this class), and the model can produce false positives on large flat regions such as windows or walls (see the third example).

## Model

| | |
| --- | --- |
| Architecture | Faster R-CNN |
| Backbone | ResNet-50 + FPN |
| Framework | PyTorch / torchvision |
| Classes | `0: background`, `1: license_plate`, `2: house_number` |
| Input size | min 800 / max 1333 px (resized internally) |
| Training data | OpenLPR (plates) + SVHN (house numbers) |
| Training hardware | NVIDIA RTX 5050 |

## Usage

```python
import torch
from PIL import Image
import torchvision.transforms.functional as TF
from huggingface_hub import snapshot_download
import sys

repo_dir = snapshot_download("ApyHTML19/Faster-RCNN-Vision-ANN18")
sys.path.insert(0, repo_dir)
from models.faster_rcnn import create_model

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

model = create_model(num_classes=3, pretrained_backbone=False)
checkpoint = torch.load(f"{repo_dir}/best_model.pth", map_location=device)
model.load_state_dict(checkpoint.get("model_state_dict", checkpoint))
model.to(device).eval()

CLASS_NAMES = {1: "license_plate", 2: "house_number"}

image = Image.open(f"{repo_dir}/examples/inputs/plate.jpg").convert("RGB")
with torch.inference_mode():
    output = model([TF.to_tensor(image).to(device)])[0]

for box, label, score in zip(output["boxes"], output["labels"], output["scores"]):
    if score >= 0.5:
        print(CLASS_NAMES[int(label)], f"{score:.3f}", [round(v) for v in box.tolist()])
```

### Blurring the detections

```python
import cv2

img = cv2.imread(f"{repo_dir}/examples/inputs/plate.jpg")
for box, score in zip(output["boxes"], output["scores"]):
    if score < 0.5:
        continue
    x1, y1, x2, y2 = map(int, box.tolist())
    img[y1:y2, x1:x2] = cv2.GaussianBlur(img[y1:y2, x1:x2], (51, 51), 0)
cv2.imwrite("anonymized.jpg", img)
```
## Intended use

Research and privacy-oriented computer vision experiments: anonymizing license plates and house numbers in street-level imagery before storage or publication. Not intended for surveillance or for identifying individuals.