task1 / HANDOVER.md
siddhant20's picture
Add files using upload-large-folder tool
d667566 verified
|
Raw History Blame Contribute Delete
21.2 kB

Handover β€” Packaging Material Difference Mining

Everything needed to continue this work: what the task is, what the data actually looks like, what was built, what the numbers are, what failed and why, and what to do next.

Current result: fold-0 out-of-fold global F1 = 0.905 (precision 0.922, recall 0.888, threshold 0.31). Classical baseline floor is 0.03. No submission has been produced yet.


1. The task

Given 100 (template, photo) image pairs, predict bounding boxes of every real content difference between a clean design template and a photo of the printed packaging. Differences are semantic changes to text, graphics or layout β€” not printing or scanning artefacts.

Scoring is global F1: TP/FP/FN are accumulated across all images before precision and recall are computed, so one image with many false positives damages the whole score. A prediction is a TP if its IoU with an unmatched ground-truth box is β‰₯ 0.5. Submission is a CSV with the same header as train.csv, one row per predicted box.

Source: Task1/Task1 Description.docx. Data: Task1/PackagingMaterialDifferenceMiningDataset/ (200 annotated train pairs, 100 unlabelled test pairs, 2.5 GB).


2. What the data actually looks like

Every claim below is reproducible with scripts/analyze_dataset.py (~2 min, CPU). These measurements drove every architectural decision, so re-run them before changing direction.

Property Measurement Why it matters
Annotations 1407 boxes over 200 pairs; 3–11 per image, median 7 Expect ~7 predictions per test image
Image size Template and photo are identical in size for all 200 pairs; 55 distinct sizes, mostly 1700Γ—2200 and 1654Γ—2339 Full-resolution tiling, no resizing
Alignment Block phase-correlation residual p95 ≀ 0.26 px. In the real Modal run, 284/300 pairs needed no warp, 16 a sub-pixel translation, 0 needed ECC Registration is not a problem in this dataset. Do not spend effort on SIFT/LoFTR matching
Box size Median 22Γ—22 px (~1.5% of image width); 419/1407 have both sides under 16 px; 83 are exactly 8Γ—8 Small-object regime; output stride 2, not 4
Change polarity 842 additions, 558 modifications, 0 pure deletions; 96.5% of boxes have more ink in the photo Signed difference channel; the polarity filter in decode.py
Degradation Photo is a blurred, noisy, tone-shifted render; Laplacian variance 2–6Γ— lower than the template. ~18% of pairs carry heavy shadow Blur matching and local background normalization in preprocess.py
Blur match found Οƒ=0.4 for 189 pairs, Οƒ=0.6 for 68, Οƒ=0 for 41 (from the real preprocessing run) Confirms the blur gap is real but small and consistent

The two hard parts, quantified.

  1. Precision. A naive normalized-difference baseline gets recall 0.57 at precision 0.015 (F1 0.03) β€” every glyph edge fires because the photo is blurred. Blur matching before differencing is what makes the problem tractable.
  2. Localization. An 8Γ—8 box at IoU β‰₯ 0.5 tolerates only ~2.7 px of shift. Thresholded difference blobs match ground truth at median IoU 0.58, with only 34% within Β±2 px on all sides. This is why the model regresses box centers and sizes (CenterNet-style) instead of thresholding a mask.

3. Architecture

Three stages. Stage 1 is trained and working; stage 2 is built and measured but currently ineffective (see Β§5); stage 3 is deterministic post-processing.

Stage 0 β€” preprocessing (src/pmdm/preprocess.py)

  1. Verify alignment by phase correlation; sub-pixel translate if needed; ECC fallback (never triggered so far).
  2. Blur matching: search a Οƒ grid, blur the template until its Laplacian variance matches the photo's. This removes the dominant false-positive source.
  3. Local background normalization (img βˆ’ GaussianBlur(img, Οƒ=31)) producing an "ink map" that is invariant to shadow and tone shift.
  4. Cache four PNGs per pair to the volume: blur-matched template, aligned photo, and both ink maps.

Stage 1 β€” siamese change detector (src/pmdm/model.py)

  • Shared-weight convnext_tiny encoder over two 4-channel streams (BGR + ink map).
  • concat(a, b, |aβˆ’b|) fusion at every scale β€” the U-Net SiamDiff/SiamConc shape that a reality-check study found still competitive with heavier change-detection transformers.
  • U-Net decoder to output stride 2, with a stride-2 stem taken straight from the stacked input so fine detail is not invented by upsampling.
  • CenterNet heads: center heatmap, width/height, sub-pixel offset, plus an auxiliary change mask.
  • Loss: Gaussian focal + masked L1 on wh/offset + Dice/BCE on the mask.
  • Trained on 768 px tiles, 60% sampled around an annotated difference.

Stage 2 β€” patch verifier (src/pmdm/verifier.py)

96Γ—96 crops around each candidate through a ResNet-18 with an 8-channel stem, producing a real-difference probability and a 4-coordinate box delta. Currently buys almost nothing β€” see Β§5.

Stage 3 β€” post-processing (src/pmdm/decode.py)

  • Weighted Boxes Fusion across overlapping tiles (averages coordinates; better than NMS for tiny boxes).
  • Polarity filter: drop candidates where the template has ink and the photo does not β€” 0 of 1404 ground-truth boxes look like that.
  • Ink snapping: refit the box to the local ink-difference blob, then apply the measured constant margin (+0, +0, +1, +1), with a Β±6 px sanity bound.
  • Global threshold sweep: one threshold for all images, chosen to maximize global F1 on out-of-fold predictions. Never guess this value.

4. Results

Stage 1, fold 0 (160 train pairs / 40 validation pairs, no synthetic data)

Two full runs were done. Run 1 hit a NaN bug from epoch 26 (fixed, see Β§6); run 2 is the clean one.

Epoch Run 1 F1 (NaN bug) Run 2 F1 (fixed)
4 0.762 0.789
9 0.819 0.840
14 0.849 0.850
19 0.881 0.905
24 0.900 0.896
29 0.911 0.889
34 0.896 0.889
39 0.903 not run (cancelled at ~36)

Read this carefully before drawing conclusions: the two runs agree within Β±0.02 at every checkpoint, which is the run-to-run noise on a 40-pair, 268-box validation set where a single box is worth ~0.004 F1. The NaN fix was a genuine bug fix but produced no measurable F1 gain. Both runs plateau from epoch ~19–24 onward while training loss keeps falling β€” the model has saturated on 160 real pairs.

Run 1's 0.911 checkpoint was deleted during cleanup (see Β§6, mistake 3). The surviving artifact is run 2's best.pt at 0.905.

Out-of-fold verification (predict_oof, run 2 best.pt)

TP 238  FP 20  FN 30   precision 0.9225  recall 0.8881  F1 0.9049  threshold 0.31

Identical to the training-time evaluation, confirming the checkpoint and the inference path agree.

Error breakdown (scripts/error_analysis.py, all 40 validation pairs, 268 boxes)

matched at IoU >= 0.5:        252
proposed but IoU < 0.5:         5      <- box tightness
never proposed (IoU == 0):     11      <- detection
recall ceiling with perfect boxes: 0.959

By box size:

Longest side n matched never proposed
<12 px 64 0.828 10
12–24 px 47 0.979 1
>24 px 157 0.975 0

This is the single most important result in the handover. Boxes above 12 px are essentially solved at 97–98%. Ten of the eleven total misses are tiny marks under 12 px. Box regression is nearly exhausted as a lever (only 5 boxes lost to loose coordinates), so effort spent on better coordinate refinement will return almost nothing.

Stage 2 verifier, honest A/B (verify_ab)

Validation pairs split by index parity: verifier trained on 16 pairs' candidates, measured on the other 24, with the pre-verifier sweep on those same 24 as the control.

TP FP FN precision recall F1
before 155 18 17 0.896 0.901 0.8986
after 152 14 20 0.916 0.884 0.8994

Delta +0.0009 β€” no effect. It traded 4 false positives for 3 true positives.

The cause is a pipeline design error, not a modelling one: predict_oof applies the full post-processing before saving candidates, so the verifier received 137 training records of which 97 were already positive. A verifier trained on 40 negatives cannot learn what artefact noise looks like. Fix before retrying: dump raw candidates at a much lower score threshold (~0.01) with use_snap=False, use_polarity=False, so the verifier sees hundreds of negatives per image and has something to discriminate.


5. What to do next, in priority order

  1. Synthetic data targeting tiny marks. The generator (src/pmdm/synth.py) is written and smoke-tested but has never been run at scale. It reproduces the dataset's own construction: paste edits into a clean template, then degrade into a "photo" (blur, noise, JPEG, tone curve, shadow field, sub-pixel warp). Its three edit types are swap (0.40), insert (0.30) and mark (0.30); mark at 6–16 px is exactly the failing category. Raise that weight, generate ~6400 pairs (modal run modal_app.py::synth --shards 16 --per-shard 400, CPU, ~20 min), then retrain fold 0 with --n-synth 4000 and compare against the 0.905 control. Expected effect, extrapolated from the error breakdown and not yet measured: closing the tiny-box gap moves recall toward the 0.959 ceiling, putting F1 around 0.93–0.94.
  2. Retry the verifier with proper negatives (see Β§4). Only worth doing after step 1, since more candidates make the verifier's job meaningful.
  3. Train folds 1–4 (modal run modal_app.py::train_all) once the configuration is settled. Do not do this before, it multiplies cost by five for no information.
  4. Ensemble and submit: predict --folds 0,1,2,3,4 --threshold <from the OOF sweep>. The threshold must come from /ckpt/oof/fold*.json, never a guess.

Ideas deliberately not pursued, with reasons: image registration (data is already aligned); higher-capacity backbones (the failure is tiny-object recall, not representational capacity); more epochs (curve flat since epoch 19); test-time flips (text is chiral).


6. Mistakes made, so they are not repeated

  1. bf16 NaN in the focal loss. Under bfloat16 autocast, 1 βˆ’ 1e-4 rounds to exactly 1.0, so log(1 βˆ’ pred) became -inf and the zero-weighted term became NaN from epoch 26 of run 1. GradScaler skipped those steps, so weights were never corrupted, but heatmap updates were lost. Fixed by forcing float32 inside gaussian_focal and dice_bce. Verified: zero NaN in run 2.
  2. train_verifier path bug. It appended "stage2" to the caller's out_dir, so checkpoints landed in /ckpt/stage2_ab/stage2/ while the loader looked in /ckpt/stage2_ab/. Fixed, plus a guard that raises if the DataLoader ends up empty instead of silently training zero steps. (The stale /ckpt/stage2_ab/stage2/ directory can be deleted.)
  3. A checkpoint was destroyed. modal volume cp refuses directories but still exits 0, so a cp && rm chain deleted run 1's 0.911 weights without a backup. Always verify a copy landed before deleting; prefer modal volume get to a local path for backups.
  4. Launch pattern. nohup modal run --detach ... & is unreliable β€” if the client is killed during app startup, before the function call is enqueued, nothing runs at all and the log looks like a successful start. One resume attempt silently trained zero epochs this way. Launch through a supervised background process and confirm training lines appear.
  5. Client disconnects kill non-detached runs. A local DNS blip killed run 1 at epoch ~1. Use --detach, and note that train_fold now commits the volume after every epoch so an interruption costs at most one epoch.

7. Code map

TASK1_CYBER/
β”œβ”€β”€ HANDOVER.md              <- this file
β”œβ”€β”€ IMPLEMENTATION_PLAN.md   <- original architecture rationale
β”œβ”€β”€ RUNBOOK.md               <- exact commands
β”œβ”€β”€ modal_app.py             <- all Modal functions and entrypoints (398 lines)
β”œβ”€β”€ pyproject.toml
β”œβ”€β”€ checkpoints/             <- trained weights, mirrored off Modal; see checkpoints/README.md
β”‚   β”œβ”€β”€ stage1_fold0_convnext_tiny/best.pt   the 0.905 model, 128 MB
β”‚   β”œβ”€β”€ stage1_fold0_convnext_tiny/last.pt   resumable state at epoch 35, 384 MB
β”‚   β”œβ”€β”€ oof/                                 out-of-fold candidates and sweep
β”‚   └── stage2_ab/                           verifier from the A/B; reference only
β”œβ”€β”€ runs/                    <- raw evidence from every run; see runs/README.md
β”‚   β”œβ”€β”€ fold0_run2_fixed.log        the primary run, full step-level losses
β”‚   β”œβ”€β”€ fold0_run2_history.json     per-epoch structured metrics
β”‚   β”œβ”€β”€ fold0_run1_nanbug.log       first run, epochs 19-31 only, shows NaN onset
β”‚   β”œβ”€β”€ oof_fold0_candidates.npz    out-of-fold boxes and scores, 40 pairs
β”‚   β”œβ”€β”€ oof_fold0_sweep.json        the 0.9049 headline result
β”‚   β”œβ”€β”€ error_analysis_output.txt   the size breakdown
β”‚   β”œβ”€β”€ analyze_dataset_output.txt  full dataset analysis output
β”‚   β”œβ”€β”€ verifier_ab_raw.txt         verifier A/B, including record counts
β”‚   └── prep_info.json              per-pair alignment mode and blur sigma
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ local_smoke.py       <- CPU pre-flight; run before any Modal spend
β”‚   β”œβ”€β”€ train_local.py       <- run training without Modal, on any GPU box
β”‚   β”œβ”€β”€ analyze_dataset.py   <- reproduces every claim in Β§2
β”‚   └── error_analysis.py    <- reproduces the breakdown in Β§4
└── src/pmdm/                <- 1500 lines total
    β”œβ”€β”€ config.py            <- paths (env-overridable) and hyperparameters
    β”œβ”€β”€ preprocess.py        <- alignment, blur match, ink maps
    β”œβ”€β”€ dataset.py           <- tiling dataset, CenterNet target encoding, GroupedSampler
    β”œβ”€β”€ model.py             <- SiamCenterNet
    β”œβ”€β”€ losses.py            <- focal / masked L1 / Dice
    β”œβ”€β”€ decode.py            <- heatmap decode, WBF, polarity filter, ink snapping
    β”œβ”€β”€ train.py             <- stage-1 loop with per-epoch resume
    β”œβ”€β”€ infer.py             <- tiled full-image inference
    β”œβ”€β”€ verifier.py          <- stage-2 model, training, application
    β”œβ”€β”€ synth.py             <- synthetic pair generator (never run at scale)
    β”œβ”€β”€ metric.py            <- global F1 scorer and threshold sweep
    β”œβ”€β”€ folds.py             <- deterministic 5-fold split
    β”œβ”€β”€ tiles.py             <- tiling helpers
    └── baseline.py          <- classical floor

How training is actually invoked

There is one training loop, src/pmdm/train.py::train_fold, reached two ways:

Path Command When to use
Modal (used for all results here) modal run --detach modal_app.py::train --fold 0 --epochs 40 Normal path. modal_app.py::train_stage1 is a thin wrapper that pins A100-40GB, mounts both volumes, and passes on_epoch_end=ckpt_vol.commit so each epoch is durable
Local / any GPU box python scripts/train_local.py --fold 0 --epochs 40 No Modal account needed. Same loop, same checkpoints, same resume behaviour. Requires stage 0 first: python scripts/train_local.py --preprocess-only

Both write to $PMDM_CKPT/stage1_fold<N>_<backbone>/ and resume from last.pt automatically. The exact command that produced the current 0.905 model was modal run --detach modal_app.py::train --fold 0 --epochs 40 with no synthetic data.

Conventions worth knowing before editing.

  • All paths come from config.py and are environment-overridable (PMDM_DATA, PMDM_WORK, PMDM_CKPT), which is how the same code runs locally and on Modal volumes.
  • Ground-truth boxes are keyed by (split, idx) tuples inside the dataset, but load_gt() returns them keyed by integer index β€” train.py::build_gt bridges the two.
  • GroupedSampler keeps a pair's tile samples adjacent so the dataset's one-entry cache serves them; without it every sample decodes four full-page PNGs.
  • The metric module is the authority on scoring. If you change decoding, re-check against metric.global_f1, not against intuition.

8. Infrastructure state (Modal account bigbalak)

Volume pmdm-data

/raw/train, /raw/test            uploaded dataset (2.5 GB)
/work/prep/{train,test}/NNN/     preprocessed pairs, all 300 done
/work/prep_info.json             per-pair alignment mode and blur sigma

Volume pmdm-ckpt

/stage1_fold0_convnext_tiny/best.pt       <- THE model, F1 0.905, epoch 19
/stage1_fold0_convnext_tiny/last.pt       <- epoch 35 state (optimizer + scheduler, resumable)
/stage1_fold0_convnext_tiny/history.json  <- per-epoch losses and evaluations
/oof/fold0_convnext_tiny.{npz,json}       <- out-of-fold candidates and sweep result
/stage2_ab/                               <- verifier A/B artifacts (weak result, see Β§4)
/hf/                                      <- cached timm weights

Nothing is running. Fold 0 stopped at epoch 35 of 40; the missing 4 epochs are not worth running given the flat curve.

Hugging Face: the whole thing β€” weights, code, docs, runs/, and the dataset β€” is pushed to the private repo https://huggingface.co/siddhant20/task1. That is the single link to hand a teammate; they need a collaborator invite on the siddhant20 account, nothing else. Clone with git clone https://huggingface.co/siddhant20/task1 (needs git-lfs) or hf download siddhant20/task1 --local-dir task1.

Everything on pmdm-ckpt except the timm cache and one duplicate directory has been mirrored into checkpoints/ in this repo, so the weights survive without access to the Modal account. The dataset on pmdm-data is not mirrored β€” it is the 2.5 GB competition download plus its preprocessed derivatives, both re-creatable from Task1/ with prep.

Hardware and rough timings: preprocessing is a CPU fan-out over 300 pairs (a few minutes); stage-1 training is A100-40GB at ~52 s/epoch plus ~60 s per evaluation, so a 40-epoch fold is roughly 45 minutes; inference and verifier work run on L4.


9. Getting started as the new owner

# 1. Local environment
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -e . torch torchvision timm opencv-python-headless "numpy<2" pandas

# 2. Confirm the pipeline works end to end on CPU, before spending anything
PMDM_DATA=Task1/PackagingMaterialDifferenceMiningDataset \
PMDM_WORK=/tmp/pmdm_work PMDM_CKPT=/tmp/pmdm_ckpt \
.venv/bin/python scripts/local_smoke.py

# 3. Confirm the data claims for yourself
PMDM_DATA=Task1/PackagingMaterialDifferenceMiningDataset \
.venv/bin/python scripts/analyze_dataset.py

# 4. Reproduce the current headline number (needs Modal access to the volumes)
modal run modal_app.py::oof --fold 0        # expect F1 0.9049, threshold 0.31

# 5. Then start on Β§5 step 1
modal run modal_app.py::synth --shards 16 --per-shard 400
modal run modal_app.py::train --fold 0 --epochs 40 --n-synth 4000

Modal access requires the bigbalak account's token (~/.modal.toml) or re-uploading the dataset to your own volumes with the commands in RUNBOOK.md. Without either, step 4 still works offline against the local mirror:

PMDM_DATA=Task1/PackagingMaterialDifferenceMiningDataset \
.venv/bin/python scripts/error_analysis.py checkpoints/oof/fold0_convnext_tiny.npz

and checkpoints/stage1_fold0_convnext_tiny/best.pt loads directly into SiamCenterNet for inference on any machine (see checkpoints/README.md).


10. Reference reading

The design draws on these; the first two are the most directly relevant.