thundercode commited on
Commit
0adf671
Β·
verified Β·
1 Parent(s): dfe79b9

release: add docs/DATASETS.md

Browse files
Files changed (1) hide show
  1. docs/DATASETS.md +162 -0
docs/DATASETS.md ADDED
@@ -0,0 +1,162 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Datasets
2
+
3
+ Every figure below was **measured from the data on disk**, not copied from a paper or a dataset
4
+ schema. Where a corpus was only partially acquired, or where the local subset differs from the
5
+ official release, that is stated β€” not smoothed over.
6
+
7
+ **Status tags:** `ACQUIRED` Β· `MEASURED` Β· `PARTIAL` Β· `NOT DOWNLOADED` Β· `REJECTED`.
8
+
9
+ ---
10
+
11
+ ## 1. Summary
12
+
13
+ | Task | Dataset | Role | Local state |
14
+ |---|---|---|---|
15
+ | `change` | **LEVIR-CD-256** | train / val / test | ACQUIRED β€” split 7120 / 1024 / 2048 |
16
+ | `change_vqa` | **CDVQA** (+ **SECOND** imagery) | train / val / test | annotations + imagery ACQUIRED; name-verified 2,968/2,968 |
17
+ | `grounding` | **VRSBench** | eval (16,159 records) | ACQUIRED β€” boxes normalised 0–100 |
18
+ | `optical_sar` | **BigEarthNet** (CLC-19) | train / held-out test | PARTIAL β€” 28,000-patch local subset |
19
+ | `vlm` | **BigEarthNet** (instruction pairs) | LoRA adaptation | PARTIAL β€” same 28k subset |
20
+ | β€” | SECOND (SCD) | CDVQA imagery source | ACQUIRED, extracted to `data/cdvqa/{im1,im2,label1,label2}/` |
21
+
22
+ ## 2. LEVIR-CD-256 β€” change detection
23
+
24
+ The standard change-detection benchmark, 256Γ—256 tiles. Split used (from `configs/base.yaml`):
25
+
26
+ | Split | Count |
27
+ |---|---|
28
+ | train | 7,120 |
29
+ | val | 1,024 |
30
+ | test | **2,048** |
31
+
32
+ **Measured test-split detail** (`artifacts/change/eval_test/eval_result.json`):
33
+
34
+ | Property | Value |
35
+ |---|---|
36
+ | n | 2,048 |
37
+ | images containing change | **935** |
38
+ | mean change fraction | **0.0509** (β‰ˆ 5 % of pixels) |
39
+ | threshold | 0.50 |
40
+ | tile size | 256 |
41
+ | device (eval) | cuda |
42
+ | total pixels scored | 134,217,728 |
43
+
44
+ The 5 % change fraction is why **pooled** and **macro** metrics are both reported: with a strong
45
+ class imbalance, pooled IoU (0.8122) and macro IoU (0.8457) answer different questions. The
46
+ per-pixel confusion counts (tp 5,978,997 / fp 523,658 / fn 858,407 / tn 126,856,666) are stored so
47
+ any metric can be recomputed rather than trusted.
48
+
49
+ ## 3. CDVQA (+ SECOND) β€” change-VQA
50
+
51
+ **The CDVQA repository publishes annotations only β€” no imagery.** The imagery is publicly available
52
+ as **SECOND (SCD)**. This project acquired SECOND, **name-verified it (2,968 / 2,968 MATCH)**,
53
+ extracted it to `data/cdvqa/{im1,im2,label1,label2}/`, and verified a CDVQA example loads against it
54
+ end-to-end.
55
+
56
+ **Measured corpus figures:**
57
+
58
+ | Property | Value |
59
+ |---|---|
60
+ | Val images | 400 |
61
+ | Val questions | **16,441** |
62
+ | Test questions | **39,686** |
63
+ | Test2 questions | second held-out set (accuracy 0.651469) |
64
+
65
+ **Temporal and label semantics β€” established from evidence, with honest uncertainty:**
66
+
67
+ | Mapping | State |
68
+ |---|---|
69
+ | `label1` = pre, `label2` = post | **established** |
70
+ | `im1` = pre, `im2` = post | **supported, not proven** |
71
+
72
+ **Corpus root:** `data/cdvqa` β€” the directory that *contains* `annotations/`. Passing
73
+ `data/cdvqa/annotations` is **rejected** by `_resolve_root` (verified by execution: it raises
74
+ `no 'annotations' directory under data\cdvqa\annotations`). The correct call is
75
+ `load_cdvqa('data/cdvqa')`.
76
+
77
+ > **Not established: trainability of the raw loader.** The adapter is pair-aware and
78
+ > `require_images=True` succeeds, but the raw corpus has no training loop of its own β€” the shipped
79
+ > `change_vqa` head trains on **cached features**, not on the raw loader. See
80
+ > [`TRAINING.md`](TRAINING.md).
81
+
82
+ ## 4. VRSBench β€” grounding evaluation
83
+
84
+ The grounding head is evaluated on VRSBench. **Measured: 16,159 eval records.**
85
+
86
+ VRSBench stores boxes **normalised to 0–100**; this project stores boxes **normalised to 0–1**. The
87
+ conversion is declared explicitly as `grounding.benchmark_box_scale: 100.0` so it cannot be applied
88
+ twice or forgotten.
89
+
90
+ Because the box convention is a common source of silent error, grounding is reported under **two
91
+ protocols** (canonical and matched6) and **two decode variants** β€” see
92
+ [`BENCHMARKS.md`](BENCHMARKS.md) Β§1.
93
+
94
+ ## 5. BigEarthNet β€” optical-SAR fusion and VLM adaptation
95
+
96
+ BigEarthNet is used in two places: as the **label space and evaluation benchmark** for the
97
+ optical-SAR fusion head, and as the **instruction-pair source** for the SmolVLM LoRA adaptation.
98
+
99
+ **Measured local subset:**
100
+
101
+ | Property | Value |
102
+ |---|---|
103
+ | S2 patches (local subset) | **28,000** |
104
+ | tiles | 98 |
105
+ | bands per patch | 12 |
106
+ | full official corpus | 480,038 patches |
107
+ | label space | **19 CLC classes** |
108
+
109
+ ### 5.1 Label-semantics caveat (important)
110
+
111
+ > **The local 28k subset is 100 % single-label, against the official 1–11 multi-label scheme.**
112
+ > Metrics computed on this subset are therefore **not comparable** to published multi-label
113
+ > BigEarthNet numbers. Any statement of the form "BigEarthNet mAP = X" is **false** for this subset.
114
+
115
+ ### 5.2 Fusion evaluation detail
116
+
117
+ **Measured** (`artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`):
118
+
119
+ | Property | Value |
120
+ |---|---|
121
+ | split | test |
122
+ | n scored | **4,000** |
123
+ | classes in label space | 19 |
124
+ | classes present in the scored split | 14 |
125
+ | classes absent | 5 |
126
+ | macro-F1 denominator | **all 19 classes (absent classes contribute 0.0)** |
127
+ | accuracy | 0.931 |
128
+ | macro F1 | 0.434161 |
129
+ | deciding statistic | **False** |
130
+
131
+ The macro-F1 denominator is recorded explicitly so a reader can see that absent classes drag the
132
+ macro score down by construction. This is why accuracy (0.931) and macro-F1 (0.434161) must be read
133
+ together.
134
+
135
+ ### 5.3 The BigEarthNet data-format contradiction
136
+
137
+ > **PARTIALLY β€” reported, not silently resolved.** The BigEarthNet data format contradicts the
138
+ > original plan. This was reported rather than quietly patched, because silently changing the
139
+ > preprocessing would move the frozen config hash. See
140
+ > [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) for the full record.
141
+
142
+ The BigEarthNet documentation β€” its uses, mentions, or endorsements β€” does **not** specify a
143
+ percentile stretch. This project nevertheless applies percentile normalisation for optical inputs
144
+ (2/98) to match the CROMA contract. That is a deliberate, documented choice, not an upstream fact.
145
+
146
+ ## 6. Data hygiene and leakage controls
147
+
148
+ - **Splits are by group, never by example**, for the router: template / hard-negative families are
149
+ kept whole, and hard-negative families are placed in the **test** split so their accuracy measures
150
+ generalisation rather than memorisation.
151
+ - **Leakage split key: `scene_id`** (declared in `configs/base.yaml`).
152
+ - **Immutable public test: true.** `hidden_data_access: false`. The evaluation config forbids
153
+ touching hidden data.
154
+
155
+ ## 7. What is NOT available
156
+
157
+ | Corpus | State |
158
+ |---|---|
159
+ | BigEarthNet S1+S2 full corpus | **NOT DOWNLOADED** (only the 28k S2 subset is local) |
160
+ | BigEarthNet multi-label (reBEN) results | **not produced** β€” the local subset is single-label |
161
+ | Cross-dataset generalisation sets | **not used** |
162
+ | Any private / hidden evaluation data | **not accessed** (`hidden_data_access: false`) |