File size: 6,893 Bytes
0adf671
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
# Datasets

Every figure below was **measured from the data on disk**, not copied from a paper or a dataset
schema. Where a corpus was only partially acquired, or where the local subset differs from the
official release, that is stated β€” not smoothed over.

**Status tags:** `ACQUIRED` Β· `MEASURED` Β· `PARTIAL` Β· `NOT DOWNLOADED` Β· `REJECTED`.

---

## 1. Summary

| Task | Dataset | Role | Local state |
|---|---|---|---|
| `change` | **LEVIR-CD-256** | train / val / test | ACQUIRED β€” split 7120 / 1024 / 2048 |
| `change_vqa` | **CDVQA** (+ **SECOND** imagery) | train / val / test | annotations + imagery ACQUIRED; name-verified 2,968/2,968 |
| `grounding` | **VRSBench** | eval (16,159 records) | ACQUIRED β€” boxes normalised 0–100 |
| `optical_sar` | **BigEarthNet** (CLC-19) | train / held-out test | PARTIAL β€” 28,000-patch local subset |
| `vlm` | **BigEarthNet** (instruction pairs) | LoRA adaptation | PARTIAL β€” same 28k subset |
| β€” | SECOND (SCD) | CDVQA imagery source | ACQUIRED, extracted to `data/cdvqa/{im1,im2,label1,label2}/` |

## 2. LEVIR-CD-256 β€” change detection

The standard change-detection benchmark, 256Γ—256 tiles. Split used (from `configs/base.yaml`):

| Split | Count |
|---|---|
| train | 7,120 |
| val | 1,024 |
| test | **2,048** |

**Measured test-split detail** (`artifacts/change/eval_test/eval_result.json`):

| Property | Value |
|---|---|
| n | 2,048 |
| images containing change | **935** |
| mean change fraction | **0.0509** (β‰ˆ 5 % of pixels) |
| threshold | 0.50 |
| tile size | 256 |
| device (eval) | cuda |
| total pixels scored | 134,217,728 |

The 5 % change fraction is why **pooled** and **macro** metrics are both reported: with a strong
class imbalance, pooled IoU (0.8122) and macro IoU (0.8457) answer different questions. The
per-pixel confusion counts (tp 5,978,997 / fp 523,658 / fn 858,407 / tn 126,856,666) are stored so
any metric can be recomputed rather than trusted.

## 3. CDVQA (+ SECOND) β€” change-VQA

**The CDVQA repository publishes annotations only β€” no imagery.** The imagery is publicly available
as **SECOND (SCD)**. This project acquired SECOND, **name-verified it (2,968 / 2,968 MATCH)**,
extracted it to `data/cdvqa/{im1,im2,label1,label2}/`, and verified a CDVQA example loads against it
end-to-end.

**Measured corpus figures:**

| Property | Value |
|---|---|
| Val images | 400 |
| Val questions | **16,441** |
| Test questions | **39,686** |
| Test2 questions | second held-out set (accuracy 0.651469) |

**Temporal and label semantics β€” established from evidence, with honest uncertainty:**

| Mapping | State |
|---|---|
| `label1` = pre, `label2` = post | **established** |
| `im1` = pre, `im2` = post | **supported, not proven** |

**Corpus root:** `data/cdvqa` β€” the directory that *contains* `annotations/`. Passing
`data/cdvqa/annotations` is **rejected** by `_resolve_root` (verified by execution: it raises
`no 'annotations' directory under data\cdvqa\annotations`). The correct call is
`load_cdvqa('data/cdvqa')`.

> **Not established: trainability of the raw loader.** The adapter is pair-aware and
> `require_images=True` succeeds, but the raw corpus has no training loop of its own β€” the shipped
> `change_vqa` head trains on **cached features**, not on the raw loader. See
> [`TRAINING.md`](TRAINING.md).

## 4. VRSBench β€” grounding evaluation

The grounding head is evaluated on VRSBench. **Measured: 16,159 eval records.**

VRSBench stores boxes **normalised to 0–100**; this project stores boxes **normalised to 0–1**. The
conversion is declared explicitly as `grounding.benchmark_box_scale: 100.0` so it cannot be applied
twice or forgotten.

Because the box convention is a common source of silent error, grounding is reported under **two
protocols** (canonical and matched6) and **two decode variants** β€” see
[`BENCHMARKS.md`](BENCHMARKS.md) Β§1.

## 5. BigEarthNet β€” optical-SAR fusion and VLM adaptation

BigEarthNet is used in two places: as the **label space and evaluation benchmark** for the
optical-SAR fusion head, and as the **instruction-pair source** for the SmolVLM LoRA adaptation.

**Measured local subset:**

| Property | Value |
|---|---|
| S2 patches (local subset) | **28,000** |
| tiles | 98 |
| bands per patch | 12 |
| full official corpus | 480,038 patches |
| label space | **19 CLC classes** |

### 5.1 Label-semantics caveat (important)

> **The local 28k subset is 100 % single-label, against the official 1–11 multi-label scheme.**
> Metrics computed on this subset are therefore **not comparable** to published multi-label
> BigEarthNet numbers. Any statement of the form "BigEarthNet mAP = X" is **false** for this subset.

### 5.2 Fusion evaluation detail

**Measured** (`artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`):

| Property | Value |
|---|---|
| split | test |
| n scored | **4,000** |
| classes in label space | 19 |
| classes present in the scored split | 14 |
| classes absent | 5 |
| macro-F1 denominator | **all 19 classes (absent classes contribute 0.0)** |
| accuracy | 0.931 |
| macro F1 | 0.434161 |
| deciding statistic | **False** |

The macro-F1 denominator is recorded explicitly so a reader can see that absent classes drag the
macro score down by construction. This is why accuracy (0.931) and macro-F1 (0.434161) must be read
together.

### 5.3 The BigEarthNet data-format contradiction

> **PARTIALLY β€” reported, not silently resolved.** The BigEarthNet data format contradicts the
> original plan. This was reported rather than quietly patched, because silently changing the
> preprocessing would move the frozen config hash. See
> [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) for the full record.

The BigEarthNet documentation β€” its uses, mentions, or endorsements β€” does **not** specify a
percentile stretch. This project nevertheless applies percentile normalisation for optical inputs
(2/98) to match the CROMA contract. That is a deliberate, documented choice, not an upstream fact.

## 6. Data hygiene and leakage controls

- **Splits are by group, never by example**, for the router: template / hard-negative families are
  kept whole, and hard-negative families are placed in the **test** split so their accuracy measures
  generalisation rather than memorisation.
- **Leakage split key: `scene_id`** (declared in `configs/base.yaml`).
- **Immutable public test: true.** `hidden_data_access: false`. The evaluation config forbids
  touching hidden data.

## 7. What is NOT available

| Corpus | State |
|---|---|
| BigEarthNet S1+S2 full corpus | **NOT DOWNLOADED** (only the 28k S2 subset is local) |
| BigEarthNet multi-label (reBEN) results | **not produced** β€” the local subset is single-label |
| Cross-dataset generalisation sets | **not used** |
| Any private / hidden evaluation data | **not accessed** (`hidden_data_access: false`) |