File size: 3,091 Bytes
583065e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
---
license: mit
tags:
  - point-cloud-registration
  - 3d-vision
  - 3dmatch
  - 3dlomatch
library_name: pytorch
---

# OCFNet -- Overlap-guided Coarse-to-fine Correspondence Prediction

Weights for the spconv port of OCFNet, trained on 3DMatch for 150 epochs.
Code: https://github.com/gfmei/OCFNet

## Model

Coarse-to-fine registration over stride-8 super-points and their disjoint Voronoi patches,
with log-domain Sinkhorn at both levels. This checkpoint adds, over the published model:

- three rounds of interleaved self/cross attention at the coarse level (the published port
  used one), with a 3D rotary position embedding on the super-point voxel indices;
- an overlap-aware circle loss on the super-point features, alongside the transport loss.

Uniform transport marginals with the dustbin on at both levels; the overlap head is trained
with the coarse and fine overlap losses but does not drive the marginals.

10.11 M parameters.

## Results

3DMatch and 3DLoMatch, 1000 sampled correspondences, correspondence-based RANSAC. RR is
registration recall, IR the inlier ratio, FMR the feature match recall (percentages); RRE is
the mean median rotation error in degrees, RTE the mean median translation error in metres.

| benchmark | RR | IR | FMR | RRE | RTE |
| --- | --- | --- | --- | --- | --- |
| 3DMatch | 89.1 | 69.8 | 96.4 | 2.25 | 0.071 |
| 3DLoMatch | 58.6 | 35.1 | 78.0 | 3.31 | 0.098 |

For reference, the published numbers are 90.2 / 58.7 / 98.5 on 3DMatch and 66.7 / 29.5 /
84.0 on 3DLoMatch. Inlier ratio here is well above the published model on both benchmarks;
3DLoMatch registration recall is below it. Two reasons, neither hidden: the backbone is
spconv rather than MinkowskiEngine and the two engines build sparse-convolution kernel maps
and handle submanifold layers differently, so this is not the same function even at
identical weights; and 21% of 3DLoMatch pairs fail at coarse patch selection, producing
almost no correct correspondences (coarse inlier ratio 0.008 against 0.477 for the rest).
Substituting ground-truth patch pairs takes that fraction to 0 and registration recall to
78.9%, so the limit is coarse selection rather than the fine features. The repository
documents this and the interventions that did not move it.

## Usage

```python
import torch
state = torch.load('ocfnet_3dmatch.pth', map_location='cpu', weights_only=False)
model.load_state_dict(state['state_dict'])   # 166 tensors, epoch 149
```

Or point a test config at it and run the repository's evaluation:

```shell
# configs/test/ocfnet_geo.yaml, field  misc.pretrain
BENCH=3DLoMatch sbatch scripts/slurm_eval_sweep.sh configs/test/ocfnet_geo.yaml
```

`config.yaml` is the training configuration this checkpoint was produced with.

## Citation

```bibtex
@inproceedings{mei2022overlap,
  title     = {Overlap-guided Coarse-to-fine Correspondence Prediction for Point Cloud Registration},
  author    = {Mei, Guofeng and Huang, Xiaoshui and Zhang, Juan and Wu, Qiang},
  booktitle = {IEEE International Conference on Multimedia and Expo (ICME)},
  year      = {2022}
}
```