task1 / HANDOVER.md
siddhant20's picture
Add files using upload-large-folder tool
d667566 verified
|
Raw History Blame Contribute Delete
21.2 kB
# Handover β€” Packaging Material Difference Mining
Everything needed to continue this work: what the task is, what the data actually looks like,
what was built, what the numbers are, what failed and why, and what to do next.
**Current result: fold-0 out-of-fold global F1 = 0.905** (precision 0.922, recall 0.888, threshold 0.31).
Classical baseline floor is 0.03. No submission has been produced yet.
---
## 1. The task
Given 100 (template, photo) image pairs, predict bounding boxes of every real content difference
between a clean design template and a photo of the printed packaging. Differences are semantic
changes to text, graphics or layout β€” not printing or scanning artefacts.
Scoring is **global F1**: TP/FP/FN are accumulated across all images *before* precision and recall
are computed, so one image with many false positives damages the whole score. A prediction is a TP
if its IoU with an unmatched ground-truth box is β‰₯ 0.5. Submission is a CSV with the same header as
`train.csv`, one row per predicted box.
Source: `Task1/Task1 Description.docx`. Data: `Task1/PackagingMaterialDifferenceMiningDataset/`
(200 annotated train pairs, 100 unlabelled test pairs, 2.5 GB).
---
## 2. What the data actually looks like
Every claim below is reproducible with `scripts/analyze_dataset.py` (~2 min, CPU).
These measurements drove every architectural decision, so re-run them before changing direction.
| Property | Measurement | Why it matters |
|---|---|---|
| Annotations | 1407 boxes over 200 pairs; 3–11 per image, median 7 | Expect ~7 predictions per test image |
| Image size | Template and photo are **identical in size** for all 200 pairs; 55 distinct sizes, mostly 1700Γ—2200 and 1654Γ—2339 | Full-resolution tiling, no resizing |
| Alignment | Block phase-correlation residual p95 ≀ 0.26 px. In the real Modal run, 284/300 pairs needed no warp, 16 a sub-pixel translation, 0 needed ECC | **Registration is not a problem in this dataset.** Do not spend effort on SIFT/LoFTR matching |
| Box size | Median 22Γ—22 px (~1.5% of image width); 419/1407 have both sides under 16 px; 83 are exactly 8Γ—8 | Small-object regime; output stride 2, not 4 |
| Change polarity | 842 additions, 558 modifications, **0 pure deletions**; 96.5% of boxes have more ink in the photo | Signed difference channel; the polarity filter in `decode.py` |
| Degradation | Photo is a blurred, noisy, tone-shifted render; Laplacian variance 2–6Γ— lower than the template. ~18% of pairs carry heavy shadow | Blur matching and local background normalization in `preprocess.py` |
| Blur match found | Οƒ=0.4 for 189 pairs, Οƒ=0.6 for 68, Οƒ=0 for 41 (from the real preprocessing run) | Confirms the blur gap is real but small and consistent |
**The two hard parts, quantified.**
1. *Precision.* A naive normalized-difference baseline gets recall 0.57 at precision **0.015**
(F1 0.03) β€” every glyph edge fires because the photo is blurred. Blur matching before
differencing is what makes the problem tractable.
2. *Localization.* An 8Γ—8 box at IoU β‰₯ 0.5 tolerates only ~2.7 px of shift. Thresholded difference
blobs match ground truth at median IoU 0.58, with only 34% within Β±2 px on all sides. This is why
the model regresses box centers and sizes (CenterNet-style) instead of thresholding a mask.
---
## 3. Architecture
Three stages. Stage 1 is trained and working; stage 2 is built and measured but currently
ineffective (see Β§5); stage 3 is deterministic post-processing.
### Stage 0 β€” preprocessing (`src/pmdm/preprocess.py`)
1. Verify alignment by phase correlation; sub-pixel translate if needed; ECC fallback (never
triggered so far).
2. **Blur matching**: search a Οƒ grid, blur the template until its Laplacian variance matches the
photo's. This removes the dominant false-positive source.
3. **Local background normalization** (`img βˆ’ GaussianBlur(img, Οƒ=31)`) producing an "ink map" that
is invariant to shadow and tone shift.
4. Cache four PNGs per pair to the volume: blur-matched template, aligned photo, and both ink maps.
### Stage 1 β€” siamese change detector (`src/pmdm/model.py`)
- Shared-weight `convnext_tiny` encoder over two 4-channel streams (BGR + ink map).
- `concat(a, b, |aβˆ’b|)` fusion at every scale β€” the U-Net SiamDiff/SiamConc shape that a
reality-check study found still competitive with heavier change-detection transformers.
- U-Net decoder to **output stride 2**, with a stride-2 stem taken straight from the stacked input
so fine detail is not invented by upsampling.
- CenterNet heads: center heatmap, width/height, sub-pixel offset, plus an auxiliary change mask.
- Loss: Gaussian focal + masked L1 on wh/offset + Dice/BCE on the mask.
- Trained on 768 px tiles, 60% sampled around an annotated difference.
### Stage 2 β€” patch verifier (`src/pmdm/verifier.py`)
96Γ—96 crops around each candidate through a ResNet-18 with an 8-channel stem, producing a
real-difference probability and a 4-coordinate box delta. **Currently buys almost nothing** β€” see Β§5.
### Stage 3 β€” post-processing (`src/pmdm/decode.py`)
- Weighted Boxes Fusion across overlapping tiles (averages coordinates; better than NMS for tiny boxes).
- **Polarity filter**: drop candidates where the template has ink and the photo does not β€” 0 of 1404
ground-truth boxes look like that.
- **Ink snapping**: refit the box to the local ink-difference blob, then apply the measured constant
margin (+0, +0, +1, +1), with a Β±6 px sanity bound.
- **Global threshold sweep**: one threshold for all images, chosen to maximize global F1 on
out-of-fold predictions. Never guess this value.
---
## 4. Results
### Stage 1, fold 0 (160 train pairs / 40 validation pairs, no synthetic data)
Two full runs were done. Run 1 hit a NaN bug from epoch 26 (fixed, see Β§6); run 2 is the clean one.
| Epoch | Run 1 F1 (NaN bug) | Run 2 F1 (fixed) |
|---|---|---|
| 4 | 0.762 | 0.789 |
| 9 | 0.819 | 0.840 |
| 14 | 0.849 | 0.850 |
| 19 | 0.881 | **0.905** |
| 24 | 0.900 | 0.896 |
| 29 | **0.911** | 0.889 |
| 34 | 0.896 | 0.889 |
| 39 | 0.903 | not run (cancelled at ~36) |
**Read this carefully before drawing conclusions**: the two runs agree within Β±0.02 at every
checkpoint, which is the run-to-run noise on a 40-pair, 268-box validation set where a single box is
worth ~0.004 F1. The NaN fix was a genuine bug fix but produced **no measurable F1 gain**. Both runs
plateau from epoch ~19–24 onward while training loss keeps falling β€” the model has saturated on 160
real pairs.
Run 1's 0.911 checkpoint was **deleted during cleanup** (see Β§6, mistake 3). The surviving artifact
is run 2's `best.pt` at 0.905.
### Out-of-fold verification (`predict_oof`, run 2 `best.pt`)
```
TP 238 FP 20 FN 30 precision 0.9225 recall 0.8881 F1 0.9049 threshold 0.31
```
Identical to the training-time evaluation, confirming the checkpoint and the inference path agree.
### Error breakdown (`scripts/error_analysis.py`, all 40 validation pairs, 268 boxes)
```
matched at IoU >= 0.5: 252
proposed but IoU < 0.5: 5 <- box tightness
never proposed (IoU == 0): 11 <- detection
recall ceiling with perfect boxes: 0.959
```
By box size:
| Longest side | n | matched | never proposed |
|---|---|---|---|
| <12 px | 64 | 0.828 | 10 |
| 12–24 px | 47 | 0.979 | 1 |
| >24 px | 157 | 0.975 | 0 |
**This is the single most important result in the handover.** Boxes above 12 px are essentially
solved at 97–98%. Ten of the eleven total misses are tiny marks under 12 px. Box regression is
nearly exhausted as a lever (only 5 boxes lost to loose coordinates), so effort spent on better
coordinate refinement will return almost nothing.
### Stage 2 verifier, honest A/B (`verify_ab`)
Validation pairs split by index parity: verifier trained on 16 pairs' candidates, measured on the
other 24, with the pre-verifier sweep on those same 24 as the control.
| | TP | FP | FN | precision | recall | F1 |
|---|---|---|---|---|---|---|
| before | 155 | 18 | 17 | 0.896 | 0.901 | 0.8986 |
| after | 152 | 14 | 20 | 0.916 | 0.884 | 0.8994 |
**Delta +0.0009 β€” no effect.** It traded 4 false positives for 3 true positives.
The cause is a pipeline design error, not a modelling one: `predict_oof` applies the full
post-processing before saving candidates, so the verifier received 137 training records of which 97
were already positive. A verifier trained on 40 negatives cannot learn what artefact noise looks
like. **Fix before retrying**: dump raw candidates at a much lower score threshold (~0.01) with
`use_snap=False, use_polarity=False`, so the verifier sees hundreds of negatives per image and has
something to discriminate.
---
## 5. What to do next, in priority order
1. **Synthetic data targeting tiny marks.** The generator (`src/pmdm/synth.py`) is written and
smoke-tested but **has never been run at scale**. It reproduces the dataset's own construction:
paste edits into a clean template, then degrade into a "photo" (blur, noise, JPEG, tone curve,
shadow field, sub-pixel warp). Its three edit types are `swap` (0.40), `insert` (0.30) and
`mark` (0.30); `mark` at 6–16 px is exactly the failing category. Raise that weight, generate
~6400 pairs (`modal run modal_app.py::synth --shards 16 --per-shard 400`, CPU, ~20 min), then
retrain fold 0 with `--n-synth 4000` and compare against the 0.905 control.
*Expected effect, extrapolated from the error breakdown and not yet measured*: closing the tiny-box
gap moves recall toward the 0.959 ceiling, putting F1 around 0.93–0.94.
2. **Retry the verifier with proper negatives** (see Β§4). Only worth doing after step 1, since more
candidates make the verifier's job meaningful.
3. **Train folds 1–4** (`modal run modal_app.py::train_all`) once the configuration is settled.
Do not do this before, it multiplies cost by five for no information.
4. **Ensemble and submit**: `predict --folds 0,1,2,3,4 --threshold <from the OOF sweep>`.
The threshold must come from `/ckpt/oof/fold*.json`, never a guess.
Ideas deliberately **not** pursued, with reasons: image registration (data is already aligned);
higher-capacity backbones (the failure is tiny-object recall, not representational capacity); more
epochs (curve flat since epoch 19); test-time flips (text is chiral).
---
## 6. Mistakes made, so they are not repeated
1. **bf16 NaN in the focal loss.** Under bfloat16 autocast, `1 βˆ’ 1e-4` rounds to exactly 1.0, so
`log(1 βˆ’ pred)` became `-inf` and the zero-weighted term became NaN from epoch 26 of run 1.
GradScaler skipped those steps, so weights were never corrupted, but heatmap updates were lost.
Fixed by forcing float32 inside `gaussian_focal` and `dice_bce`. Verified: zero NaN in run 2.
2. **`train_verifier` path bug.** It appended `"stage2"` to the caller's `out_dir`, so checkpoints
landed in `/ckpt/stage2_ab/stage2/` while the loader looked in `/ckpt/stage2_ab/`. Fixed, plus a
guard that raises if the DataLoader ends up empty instead of silently training zero steps.
(The stale `/ckpt/stage2_ab/stage2/` directory can be deleted.)
3. **A checkpoint was destroyed.** `modal volume cp` refuses directories but still **exits 0**, so a
`cp && rm` chain deleted run 1's 0.911 weights without a backup. Always verify a copy landed
before deleting; prefer `modal volume get` to a local path for backups.
4. **Launch pattern.** `nohup modal run --detach ... &` is unreliable β€” if the client is killed
during app startup, before the function call is enqueued, nothing runs at all and the log looks
like a successful start. One resume attempt silently trained zero epochs this way. Launch through
a supervised background process and confirm training lines appear.
5. **Client disconnects kill non-detached runs.** A local DNS blip killed run 1 at epoch ~1. Use
`--detach`, and note that `train_fold` now commits the volume after every epoch so an
interruption costs at most one epoch.
---
## 7. Code map
```
TASK1_CYBER/
β”œβ”€β”€ HANDOVER.md <- this file
β”œβ”€β”€ IMPLEMENTATION_PLAN.md <- original architecture rationale
β”œβ”€β”€ RUNBOOK.md <- exact commands
β”œβ”€β”€ modal_app.py <- all Modal functions and entrypoints (398 lines)
β”œβ”€β”€ pyproject.toml
β”œβ”€β”€ checkpoints/ <- trained weights, mirrored off Modal; see checkpoints/README.md
β”‚ β”œβ”€β”€ stage1_fold0_convnext_tiny/best.pt the 0.905 model, 128 MB
β”‚ β”œβ”€β”€ stage1_fold0_convnext_tiny/last.pt resumable state at epoch 35, 384 MB
β”‚ β”œβ”€β”€ oof/ out-of-fold candidates and sweep
β”‚ └── stage2_ab/ verifier from the A/B; reference only
β”œβ”€β”€ runs/ <- raw evidence from every run; see runs/README.md
β”‚ β”œβ”€β”€ fold0_run2_fixed.log the primary run, full step-level losses
β”‚ β”œβ”€β”€ fold0_run2_history.json per-epoch structured metrics
β”‚ β”œβ”€β”€ fold0_run1_nanbug.log first run, epochs 19-31 only, shows NaN onset
β”‚ β”œβ”€β”€ oof_fold0_candidates.npz out-of-fold boxes and scores, 40 pairs
β”‚ β”œβ”€β”€ oof_fold0_sweep.json the 0.9049 headline result
β”‚ β”œβ”€β”€ error_analysis_output.txt the size breakdown
β”‚ β”œβ”€β”€ analyze_dataset_output.txt full dataset analysis output
β”‚ β”œβ”€β”€ verifier_ab_raw.txt verifier A/B, including record counts
β”‚ └── prep_info.json per-pair alignment mode and blur sigma
β”œβ”€β”€ scripts/
β”‚ β”œβ”€β”€ local_smoke.py <- CPU pre-flight; run before any Modal spend
β”‚ β”œβ”€β”€ train_local.py <- run training without Modal, on any GPU box
β”‚ β”œβ”€β”€ analyze_dataset.py <- reproduces every claim in Β§2
β”‚ └── error_analysis.py <- reproduces the breakdown in Β§4
└── src/pmdm/ <- 1500 lines total
β”œβ”€β”€ config.py <- paths (env-overridable) and hyperparameters
β”œβ”€β”€ preprocess.py <- alignment, blur match, ink maps
β”œβ”€β”€ dataset.py <- tiling dataset, CenterNet target encoding, GroupedSampler
β”œβ”€β”€ model.py <- SiamCenterNet
β”œβ”€β”€ losses.py <- focal / masked L1 / Dice
β”œβ”€β”€ decode.py <- heatmap decode, WBF, polarity filter, ink snapping
β”œβ”€β”€ train.py <- stage-1 loop with per-epoch resume
β”œβ”€β”€ infer.py <- tiled full-image inference
β”œβ”€β”€ verifier.py <- stage-2 model, training, application
β”œβ”€β”€ synth.py <- synthetic pair generator (never run at scale)
β”œβ”€β”€ metric.py <- global F1 scorer and threshold sweep
β”œβ”€β”€ folds.py <- deterministic 5-fold split
β”œβ”€β”€ tiles.py <- tiling helpers
└── baseline.py <- classical floor
```
### How training is actually invoked
There is one training loop, `src/pmdm/train.py::train_fold`, reached two ways:
| Path | Command | When to use |
|---|---|---|
| Modal (used for all results here) | `modal run --detach modal_app.py::train --fold 0 --epochs 40` | Normal path. `modal_app.py::train_stage1` is a thin wrapper that pins A100-40GB, mounts both volumes, and passes `on_epoch_end=ckpt_vol.commit` so each epoch is durable |
| Local / any GPU box | `python scripts/train_local.py --fold 0 --epochs 40` | No Modal account needed. Same loop, same checkpoints, same resume behaviour. Requires stage 0 first: `python scripts/train_local.py --preprocess-only` |
Both write to `$PMDM_CKPT/stage1_fold<N>_<backbone>/` and resume from `last.pt` automatically.
The exact command that produced the current 0.905 model was
`modal run --detach modal_app.py::train --fold 0 --epochs 40` with no synthetic data.
**Conventions worth knowing before editing.**
- All paths come from `config.py` and are environment-overridable (`PMDM_DATA`, `PMDM_WORK`,
`PMDM_CKPT`), which is how the same code runs locally and on Modal volumes.
- Ground-truth boxes are keyed by `(split, idx)` tuples inside the dataset, but `load_gt()` returns
them keyed by integer index β€” `train.py::build_gt` bridges the two.
- `GroupedSampler` keeps a pair's tile samples adjacent so the dataset's one-entry cache serves them;
without it every sample decodes four full-page PNGs.
- The metric module is the authority on scoring. If you change decoding, re-check against
`metric.global_f1`, not against intuition.
---
## 8. Infrastructure state (Modal account `bigbalak`)
**Volume `pmdm-data`**
```
/raw/train, /raw/test uploaded dataset (2.5 GB)
/work/prep/{train,test}/NNN/ preprocessed pairs, all 300 done
/work/prep_info.json per-pair alignment mode and blur sigma
```
**Volume `pmdm-ckpt`**
```
/stage1_fold0_convnext_tiny/best.pt <- THE model, F1 0.905, epoch 19
/stage1_fold0_convnext_tiny/last.pt <- epoch 35 state (optimizer + scheduler, resumable)
/stage1_fold0_convnext_tiny/history.json <- per-epoch losses and evaluations
/oof/fold0_convnext_tiny.{npz,json} <- out-of-fold candidates and sweep result
/stage2_ab/ <- verifier A/B artifacts (weak result, see Β§4)
/hf/ <- cached timm weights
```
Nothing is running. Fold 0 stopped at epoch 35 of 40; the missing 4 epochs are not worth running
given the flat curve.
**Hugging Face**: the whole thing β€” weights, code, docs, `runs/`, and the dataset β€” is pushed to the
private repo <https://huggingface.co/siddhant20/task1>. That is the single link to hand a teammate;
they need a collaborator invite on the `siddhant20` account, nothing else. Clone with
`git clone https://huggingface.co/siddhant20/task1` (needs git-lfs) or
`hf download siddhant20/task1 --local-dir task1`.
Everything on `pmdm-ckpt` except the timm cache and one duplicate directory has been mirrored into
`checkpoints/` in this repo, so the weights survive without access to the Modal account. The
dataset on `pmdm-data` is not mirrored β€” it is the 2.5 GB competition download plus its
preprocessed derivatives, both re-creatable from `Task1/` with `prep`.
**Hardware and rough timings**: preprocessing is a CPU fan-out over 300 pairs (a few minutes);
stage-1 training is A100-40GB at ~52 s/epoch plus ~60 s per evaluation, so a 40-epoch fold is
roughly 45 minutes; inference and verifier work run on L4.
---
## 9. Getting started as the new owner
```bash
# 1. Local environment
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -e . torch torchvision timm opencv-python-headless "numpy<2" pandas
# 2. Confirm the pipeline works end to end on CPU, before spending anything
PMDM_DATA=Task1/PackagingMaterialDifferenceMiningDataset \
PMDM_WORK=/tmp/pmdm_work PMDM_CKPT=/tmp/pmdm_ckpt \
.venv/bin/python scripts/local_smoke.py
# 3. Confirm the data claims for yourself
PMDM_DATA=Task1/PackagingMaterialDifferenceMiningDataset \
.venv/bin/python scripts/analyze_dataset.py
# 4. Reproduce the current headline number (needs Modal access to the volumes)
modal run modal_app.py::oof --fold 0 # expect F1 0.9049, threshold 0.31
# 5. Then start on Β§5 step 1
modal run modal_app.py::synth --shards 16 --per-shard 400
modal run modal_app.py::train --fold 0 --epochs 40 --n-synth 4000
```
Modal access requires the `bigbalak` account's token (`~/.modal.toml`) or re-uploading the dataset
to your own volumes with the commands in `RUNBOOK.md`. Without either, step 4 still works offline
against the local mirror:
```bash
PMDM_DATA=Task1/PackagingMaterialDifferenceMiningDataset \
.venv/bin/python scripts/error_analysis.py checkpoints/oof/fold0_convnext_tiny.npz
```
and `checkpoints/stage1_fold0_convnext_tiny/best.pt` loads directly into `SiamCenterNet` for
inference on any machine (see `checkpoints/README.md`).
---
## 10. Reference reading
The design draws on these; the first two are the most directly relevant.
- [Spot the Difference by Object Detection (arXiv:1801.01051)](https://arxiv.org/pdf/1801.01051) β€” printed book cover vs digital design, the closest published analogue
- [A Change Detection Reality Check (arXiv:2402.06994)](https://arxiv.org/abs/2402.06994) β€” simple siamese U-Nets remain competitive with elaborate architectures
- [Changer: Feature Interaction is What You Need (arXiv:2209.08290)](https://arxiv.org/abs/2209.08290) β€” stream interaction design
- [Self-Pair: Synthesizing Changes from Single Source (arXiv:2212.10236)](https://arxiv.org/abs/2212.10236) β€” the synthetic pair strategy
- [SAHI: Slicing Aided Hyper Inference (arXiv:2202.06934)](https://arxiv.org/abs/2202.06934) β€” tiled training and inference for small objects
- [Weighted Boxes Fusion (arXiv:1910.13302)](https://arxiv.org/abs/1910.13302) β€” coordinate-averaging box merge
- [Open-CD toolbox](https://github.com/likyoo/open-cd) β€” reference implementations of change-detection baselines
- [OmniDocBench (arXiv:2412.07626)](https://arxiv.org/abs/2412.07626) β€” the corpus the dataset was derived from, and a source of extra templates