pomidori73 commited on
Commit
d1b047a
·
verified ·
1 Parent(s): c90819c

Note that the union checkpoint is the last epoch, not the best-validation epoch

Browse files
Files changed (1) hide show
  1. README.md +2 -0
README.md CHANGED
@@ -28,6 +28,8 @@ All models share the same backbone (d_model 128, 6 layers, 4 heads, ~6.0M parame
28
  | `e1_v2_manifold_noaux/model.ckpt` | Ablation E1-V2 | on-shell (manifold) | no | 2 epochs, 48k steps |
29
  | `e1_v3_euclidean_noaux/model.ckpt` | Ablation E1-V3 | free (Euclidean) | no | 2 epochs, 48k steps |
30
 
 
 
31
  Checkpoint format: PyTorch Lightning (weights + EMA weights + hyperparameters; optimizer state stripped). The hyperparameters carry `architecture: legacy_type_film`, which selects the paper architecture in the ShellFlow code. Distribution matching (a training-time reweighting) is disabled in the stored hyperparameters so the checkpoints load without any dataset present; it has no effect on generation.
32
 
33
  ## Usage
 
28
  | `e1_v2_manifold_noaux/model.ckpt` | Ablation E1-V2 | on-shell (manifold) | no | 2 epochs, 48k steps |
29
  | `e1_v3_euclidean_noaux/model.ckpt` | Ablation E1-V3 | free (Euclidean) | no | 2 epochs, 48k steps |
30
 
31
+ Checkpoint selection: `union_1LMET30_2to4lep/model.ckpt` holds the weights at the end of the last training epoch (stored `epoch` = 29, `global_step` = 717210). It is not the checkpoint with the lowest validation loss, which occurred at epoch 26 (`val/loss/total` = 1.1685). Epochs are counted from 0, as in PyTorch Lightning.
32
+
33
  Checkpoint format: PyTorch Lightning (weights + EMA weights + hyperparameters; optimizer state stripped). The hyperparameters carry `architecture: legacy_type_film`, which selects the paper architecture in the ShellFlow code. Distribution matching (a training-time reweighting) is disabled in the stored hyperparameters so the checkpoints load without any dataset present; it has no effect on generation.
34
 
35
  ## Usage