|
Download IMPLEMENTATION_PLAN.md from siddhant20/task1: direct link, hf CLI and curl.
- Browser
- Download file 7.92 kB
-
https://huggingface.co/siddhant20/task1/resolve/main/IMPLEMENTATION_PLAN.md
- Command line
-
hf download hf://siddhant20/task1/IMPLEMENTATION_PLAN.md
-
curl -L -o IMPLEMENTATION_PLAN.md https://huggingface.co/siddhant20/task1/resolve/main/IMPLEMENTATION_PLAN.md
7.92 kB
| # Packaging Material Difference Mining β Implementation Plan | |
| > **Status note.** This document is the original design rationale, written before any training run. | |
| > It is kept as-is so the reasoning behind each choice stays readable. For current results, what was | |
| > actually built, what failed, and what to do next, read **`HANDOVER.md`** β that file supersedes the | |
| > phase plan in Β§4 below. | |
| Target metric: global F1 (TP/FP/FN accumulated over all 100 test pairs, IoU β₯ 0.5 matching). | |
| ## 0. Measured dataset facts driving the design | |
| | Fact | Value | Design consequence | | |
| |---|---|---| | |
| | Pairs | 200 train (1407 boxes, 3β11 per image, median 7) / 100 test | Data synthesis is the main lever | | |
| | Image size | template and photo identical per pair; mostly 1654Γ2339, 1700Γ2200 | Work at native resolution with tiling | | |
| | Alignment | block phase-correlation residual \|Ξ\| p95 β€ 0.26 px | No registration stage required | | |
| | Box size | median 22Γ22 px (~1% of width); 419/1407 under 16 px; 83 exactly 8Γ8 | Small-object regime; sub-pixel box regression | | |
| | Polarity | 842 additions, 558 modifications, **0 deletions**; 96.5% more ink in photo | Signed difference channel + polarity filter | | |
| | Degradation | photo is blur + noise + tone shifted (Laplacian variance 2β6Γ lower); ~18% of pairs carry heavy global shadow | Blur matching + local background normalization | | |
| | Naive diff baseline | recall 0.57, precision 0.015, F1 0.03 | Precision is problem #1 | | |
| | Threshold-blob box vs GT | median IoU 0.58, 34% within Β±2 px | Box tightness is problem #2 | | |
| An 8Γ8 box at IoU β₯ 0.5 tolerates only ~2.7 px of shift, so localization accuracy is worth as much as detection accuracy. | |
| ## 1. Architecture | |
| Three stages plus metric-aware post-processing. | |
| **Stage 0 β preprocessing (CPU, cached to volume)** | |
| 1. Per-pair alignment check with phase correlation; ECC homography fallback if the response is low. | |
| 2. Blur matching: estimate the photo's blur from the Laplacian variance ratio and apply the matched Gaussian to the template. Removes the largest false-positive class (sharp-vs-soft glyph edges). | |
| 3. Local background normalization (`img β GaussianBlur(img, Ο=31)`) on both images to remove shadows and tone shift. | |
| 4. Stack an 8-channel tensor: template RGB (3), photo RGB (3), signed difference (1), gradient-magnitude difference (1). | |
| **Stage 1 β siamese change detector** | |
| - Shared-weight `convnext_tiny` encoder (timm, ImageNet initialization) over both streams, with concat + absolute-difference fusion at every scale and Changer-style stream exchange. | |
| - U-Net decoder to output stride 2, since the objects are tiny. | |
| - CenterNet-style heads: center heatmap, width/height, sub-pixel offset, plus an auxiliary change mask. | |
| - Loss: Gaussian focal on the heatmap + L1 on width/height and offset + Dice/BCE on the auxiliary mask. | |
| - Tiles of 768 px with 25% overlap, sampled around ground truth plus hard-mined negatives. | |
| **Stage 2 β patch verifier and refiner** | |
| - 96Γ96 crops from both images around each stage-1 candidate, ResNet-18 with a 6-channel stem. | |
| - Two outputs: real-difference probability and a 4-coordinate box delta. | |
| - Trained on stage-1 out-of-fold predictions (matched candidates positive, unmatched negative). | |
| **Stage 3 β post-processing** | |
| - Weighted Boxes Fusion to merge tile predictions (better tiny-box coordinates than NMS). | |
| - Polarity filter: drop candidates where the template has ink and the photo does not (0 of 1404 ground-truth boxes look like that). | |
| - Ink snapping: fit the box to the photo-side ink connected component, then apply the measured constant margin (+0, +0, +1, +1). | |
| - Global threshold sweep over out-of-fold predictions, optimizing global F1 with a single threshold shared by all images. | |
| ## 2. Data strategy | |
| 1. **Synthetic pair generator** replicating the dataset's own construction: clean template β paste a synthetic difference (glyph swap, character insertion, 8Γ8 mark, region content change, sized to match the observed distribution) β degradation chain (Gaussian blur, sensor noise, JPEG, tone curve, shadow map, Β±0.3 px warp). Target 10k+ pairs from the 200 templates plus additional OmniDocBench pages. | |
| 2. **Copy-paste augmentation** of real annotated difference patches into other pairs. | |
| 3. **5-fold cross-validation** on the 200 real pairs; synthetic data is added to every training fold, never to validation. | |
| ## 3. Modal execution plan | |
| | Modal function | Hardware | Purpose | | |
| |---|---|---| | |
| | `preprocess` | CPU Γ8 | Alignment check, blur match, cache preprocessed pairs and tile index to the volume | | |
| | `synth` | CPU Γ16, parallel map | Generate synthetic pairs into the volume | | |
| | `train_stage1` | A100-40GB | Train the siamese detector for one fold | | |
| | `predict_oof` | A10G | Out-of-fold tiled inference, produces verifier training data | | |
| | `train_verifier` | A10G | Train the stage-2 patch verifier | | |
| | `predict_test` | A10G | Tiled test inference, WBF, verifier rescoring | | |
| | `sweep` | CPU | Global-F1 threshold and post-processing calibration | | |
| Two Modal volumes: `pmdm-data` (raw + preprocessed + synthetic) and `pmdm-ckpt` (checkpoints, predictions, submissions). Code ships via `add_local_python_source`, so edits take effect without rebuilding the image. | |
| ## 4. Phases | |
| **Phase 1 β harness (no GPU)** | |
| - Official-equivalent scorer, 5-fold split, preprocessing, tiling, and the naive diff baseline as a floor. | |
| - Exit criterion: the scorer reproduces the F1 definition in the task description and the baseline reports F1 β 0.03. | |
| **Phase 2 β stage-1 model** | |
| - Train one fold on Modal, decode, sweep the threshold, measure out-of-fold global F1. | |
| - Exit criterion: out-of-fold F1 substantially above the baseline, with the precision/recall split logged. | |
| **Phase 3 β synthetic data** | |
| - Build the generator, pretrain on synthetic pairs, fine-tune on real folds. | |
| - Exit criterion: measurable out-of-fold F1 gain over Phase 2 at equal decode settings. | |
| **Phase 4 β verifier and calibration** | |
| - Train the stage-2 verifier on out-of-fold candidates, add ink snapping and the margin calibration, re-sweep. | |
| - Exit criterion: precision gain without a matching recall loss. | |
| **Phase 5 β ensemble and submission** | |
| - 5 folds Γ 2 backbones (ConvNeXt-UNet, SegFormer), WBF merge, tile-offset and multi-scale test-time augmentation (no flips β text is chiral), final threshold sweep, write `submission.csv`. | |
| ## 5. Repository layout | |
| ``` | |
| TASK1_CYBER/ | |
| βββ IMPLEMENTATION_PLAN.md | |
| βββ modal_app.py # Modal app: image, volumes, all remote entrypoints | |
| βββ pyproject.toml | |
| βββ src/pmdm/ | |
| βββ config.py # paths, hyperparameters | |
| βββ metric.py # global F1 scorer, greedy IoU matching | |
| βββ folds.py # 5-fold split | |
| βββ preprocess.py # alignment, blur match, local normalization, channel stack | |
| βββ tiles.py # tiling and box bookkeeping | |
| βββ dataset.py # torch datasets, augmentation, target encoding | |
| βββ model.py # siamese ConvNeXt-UNet with CenterNet heads | |
| βββ losses.py # focal, L1, Dice | |
| βββ decode.py # heatmap decode, WBF, ink snapping, polarity filter | |
| βββ train.py # stage-1 training loop | |
| βββ verifier.py # stage-2 model, training, application | |
| βββ synth.py # synthetic pair generator | |
| βββ infer.py # tiled inference, out-of-fold and test | |
| βββ baseline.py # classical diff baseline (floor) | |
| ``` | |
| ## 6. Cost control | |
| - Development runs on a single fold with a reduced tile budget before any 5-fold sweep. | |
| - Every Modal function has an explicit timeout and writes checkpoints to the volume after each epoch, so a preempted run resumes rather than restarts. | |
| - GPU functions are invoked explicitly; nothing trains on import. | |