task1 / checkpoints /README.md
siddhant20's picture
Add files using upload-large-folder tool
d667566 verified
|
Raw History Blame Contribute Delete
2.59 kB

Local checkpoint mirror

Downloaded from the Modal volume pmdm-ckpt (account bigbalak) on 2026-08-16 with modal volume get. These are the actual trained weights behind every number in HANDOVER.md, kept locally so the next owner does not need access to that Modal account.

Path Size What it is
stage1_fold0_convnext_tiny/best.pt 128 MB The model. Fold-0 detector, ConvNeXt-tiny backbone, saved at epoch 19 — the best out-of-fold evaluation of run 2. 318 tensors, 32.8 M parameters.
stage1_fold0_convnext_tiny/last.pt 384 MB Final training state (epoch 35): weights plus optimizer and scheduler state. Use this to resume training, not to run inference.
stage1_fold0_convnext_tiny/history.json 12 KB Per-epoch losses, timings, and the seven evaluations. Same file as runs/fold0_run2_history.json.
oof/fold0_convnext_tiny.npz 24 KB Out-of-fold candidate boxes and scores on the 40 validation pairs.
oof/fold0_convnext_tiny.json 4 KB The threshold sweep on those candidates.
stage2_ab/last.pt 43 MB Verifier from the leakage-free A/B. Kept for reference only — its measured delta was +0.0009, i.e. nothing.
stage2_ab/records.json 20 KB The crops the verifier trained on. Explains why it failed: too few negatives.

Not downloaded: pmdm-ckpt:/stage2_ab/stage2/, a duplicate written by the path bug described in HANDOVER.md §6 before it was fixed, and pmdm-ckpt:/hf/, which is just the timm pretrained-weight cache and re-downloads on demand.

Checkpoint contents

best.pt holds three keys:

ck = torch.load("checkpoints/stage1_fold0_convnext_tiny/best.pt", map_location="cpu")
ck["model"]    # state dict for pmdm.model.SiamCenterNet(backbone="convnext_tiny")
ck["epoch"]    # 19
ck["metric"]   # {'f1': 0.9049, 'precision': 0.9225, 'recall': 0.8881, 'threshold': 0.31, ...}

Load for inference:

import torch
from pmdm.model import SiamCenterNet

model = SiamCenterNet(backbone="convnext_tiny")
ck = torch.load("checkpoints/stage1_fold0_convnext_tiny/best.pt", map_location="cpu")
model.load_state_dict(ck["model"])
model.eval()

last.pt additionally carries optimizer, scaler, and scheduler state, which is why it is three times the size. train_fold picks it up automatically when the checkpoint directory is present, so copying this directory back onto a Modal volume resumes training from epoch 36.

Pushing back to Modal

modal volume put pmdm-ckpt checkpoints/stage1_fold0_convnext_tiny /stage1_fold0_convnext_tiny