task1 / IMPLEMENTATION_PLAN.md
siddhant20's picture
Add files using upload-large-folder tool
d667566 verified
|
Raw History Blame Contribute Delete
7.92 kB
# Packaging Material Difference Mining β€” Implementation Plan
> **Status note.** This document is the original design rationale, written before any training run.
> It is kept as-is so the reasoning behind each choice stays readable. For current results, what was
> actually built, what failed, and what to do next, read **`HANDOVER.md`** β€” that file supersedes the
> phase plan in Β§4 below.
Target metric: global F1 (TP/FP/FN accumulated over all 100 test pairs, IoU β‰₯ 0.5 matching).
## 0. Measured dataset facts driving the design
| Fact | Value | Design consequence |
|---|---|---|
| Pairs | 200 train (1407 boxes, 3–11 per image, median 7) / 100 test | Data synthesis is the main lever |
| Image size | template and photo identical per pair; mostly 1654Γ—2339, 1700Γ—2200 | Work at native resolution with tiling |
| Alignment | block phase-correlation residual \|Ξ”\| p95 ≀ 0.26 px | No registration stage required |
| Box size | median 22Γ—22 px (~1% of width); 419/1407 under 16 px; 83 exactly 8Γ—8 | Small-object regime; sub-pixel box regression |
| Polarity | 842 additions, 558 modifications, **0 deletions**; 96.5% more ink in photo | Signed difference channel + polarity filter |
| Degradation | photo is blur + noise + tone shifted (Laplacian variance 2–6Γ— lower); ~18% of pairs carry heavy global shadow | Blur matching + local background normalization |
| Naive diff baseline | recall 0.57, precision 0.015, F1 0.03 | Precision is problem #1 |
| Threshold-blob box vs GT | median IoU 0.58, 34% within Β±2 px | Box tightness is problem #2 |
An 8Γ—8 box at IoU β‰₯ 0.5 tolerates only ~2.7 px of shift, so localization accuracy is worth as much as detection accuracy.
## 1. Architecture
Three stages plus metric-aware post-processing.
**Stage 0 β€” preprocessing (CPU, cached to volume)**
1. Per-pair alignment check with phase correlation; ECC homography fallback if the response is low.
2. Blur matching: estimate the photo's blur from the Laplacian variance ratio and apply the matched Gaussian to the template. Removes the largest false-positive class (sharp-vs-soft glyph edges).
3. Local background normalization (`img βˆ’ GaussianBlur(img, Οƒ=31)`) on both images to remove shadows and tone shift.
4. Stack an 8-channel tensor: template RGB (3), photo RGB (3), signed difference (1), gradient-magnitude difference (1).
**Stage 1 β€” siamese change detector**
- Shared-weight `convnext_tiny` encoder (timm, ImageNet initialization) over both streams, with concat + absolute-difference fusion at every scale and Changer-style stream exchange.
- U-Net decoder to output stride 2, since the objects are tiny.
- CenterNet-style heads: center heatmap, width/height, sub-pixel offset, plus an auxiliary change mask.
- Loss: Gaussian focal on the heatmap + L1 on width/height and offset + Dice/BCE on the auxiliary mask.
- Tiles of 768 px with 25% overlap, sampled around ground truth plus hard-mined negatives.
**Stage 2 β€” patch verifier and refiner**
- 96Γ—96 crops from both images around each stage-1 candidate, ResNet-18 with a 6-channel stem.
- Two outputs: real-difference probability and a 4-coordinate box delta.
- Trained on stage-1 out-of-fold predictions (matched candidates positive, unmatched negative).
**Stage 3 β€” post-processing**
- Weighted Boxes Fusion to merge tile predictions (better tiny-box coordinates than NMS).
- Polarity filter: drop candidates where the template has ink and the photo does not (0 of 1404 ground-truth boxes look like that).
- Ink snapping: fit the box to the photo-side ink connected component, then apply the measured constant margin (+0, +0, +1, +1).
- Global threshold sweep over out-of-fold predictions, optimizing global F1 with a single threshold shared by all images.
## 2. Data strategy
1. **Synthetic pair generator** replicating the dataset's own construction: clean template β†’ paste a synthetic difference (glyph swap, character insertion, 8Γ—8 mark, region content change, sized to match the observed distribution) β†’ degradation chain (Gaussian blur, sensor noise, JPEG, tone curve, shadow map, Β±0.3 px warp). Target 10k+ pairs from the 200 templates plus additional OmniDocBench pages.
2. **Copy-paste augmentation** of real annotated difference patches into other pairs.
3. **5-fold cross-validation** on the 200 real pairs; synthetic data is added to every training fold, never to validation.
## 3. Modal execution plan
| Modal function | Hardware | Purpose |
|---|---|---|
| `preprocess` | CPU Γ—8 | Alignment check, blur match, cache preprocessed pairs and tile index to the volume |
| `synth` | CPU Γ—16, parallel map | Generate synthetic pairs into the volume |
| `train_stage1` | A100-40GB | Train the siamese detector for one fold |
| `predict_oof` | A10G | Out-of-fold tiled inference, produces verifier training data |
| `train_verifier` | A10G | Train the stage-2 patch verifier |
| `predict_test` | A10G | Tiled test inference, WBF, verifier rescoring |
| `sweep` | CPU | Global-F1 threshold and post-processing calibration |
Two Modal volumes: `pmdm-data` (raw + preprocessed + synthetic) and `pmdm-ckpt` (checkpoints, predictions, submissions). Code ships via `add_local_python_source`, so edits take effect without rebuilding the image.
## 4. Phases
**Phase 1 β€” harness (no GPU)**
- Official-equivalent scorer, 5-fold split, preprocessing, tiling, and the naive diff baseline as a floor.
- Exit criterion: the scorer reproduces the F1 definition in the task description and the baseline reports F1 β‰ˆ 0.03.
**Phase 2 β€” stage-1 model**
- Train one fold on Modal, decode, sweep the threshold, measure out-of-fold global F1.
- Exit criterion: out-of-fold F1 substantially above the baseline, with the precision/recall split logged.
**Phase 3 β€” synthetic data**
- Build the generator, pretrain on synthetic pairs, fine-tune on real folds.
- Exit criterion: measurable out-of-fold F1 gain over Phase 2 at equal decode settings.
**Phase 4 β€” verifier and calibration**
- Train the stage-2 verifier on out-of-fold candidates, add ink snapping and the margin calibration, re-sweep.
- Exit criterion: precision gain without a matching recall loss.
**Phase 5 β€” ensemble and submission**
- 5 folds Γ— 2 backbones (ConvNeXt-UNet, SegFormer), WBF merge, tile-offset and multi-scale test-time augmentation (no flips β€” text is chiral), final threshold sweep, write `submission.csv`.
## 5. Repository layout
```
TASK1_CYBER/
β”œβ”€β”€ IMPLEMENTATION_PLAN.md
β”œβ”€β”€ modal_app.py # Modal app: image, volumes, all remote entrypoints
β”œβ”€β”€ pyproject.toml
└── src/pmdm/
β”œβ”€β”€ config.py # paths, hyperparameters
β”œβ”€β”€ metric.py # global F1 scorer, greedy IoU matching
β”œβ”€β”€ folds.py # 5-fold split
β”œβ”€β”€ preprocess.py # alignment, blur match, local normalization, channel stack
β”œβ”€β”€ tiles.py # tiling and box bookkeeping
β”œβ”€β”€ dataset.py # torch datasets, augmentation, target encoding
β”œβ”€β”€ model.py # siamese ConvNeXt-UNet with CenterNet heads
β”œβ”€β”€ losses.py # focal, L1, Dice
β”œβ”€β”€ decode.py # heatmap decode, WBF, ink snapping, polarity filter
β”œβ”€β”€ train.py # stage-1 training loop
β”œβ”€β”€ verifier.py # stage-2 model, training, application
β”œβ”€β”€ synth.py # synthetic pair generator
β”œβ”€β”€ infer.py # tiled inference, out-of-fold and test
└── baseline.py # classical diff baseline (floor)
```
## 6. Cost control
- Development runs on a single fold with a reduced tile budget before any 5-fold sweep.
- Every Modal function has an explicit timeout and writes checkpoints to the volume after each epoch, so a preempted run resumes rather than restarts.
- GPU functions are invoked explicitly; nothing trains on import.