Note that the union checkpoint is the last epoch, not the best-validation epoch
Browse files
README.md
CHANGED
|
@@ -28,6 +28,8 @@ All models share the same backbone (d_model 128, 6 layers, 4 heads, ~6.0M parame
|
|
| 28 |
| `e1_v2_manifold_noaux/model.ckpt` | Ablation E1-V2 | on-shell (manifold) | no | 2 epochs, 48k steps |
|
| 29 |
| `e1_v3_euclidean_noaux/model.ckpt` | Ablation E1-V3 | free (Euclidean) | no | 2 epochs, 48k steps |
|
| 30 |
|
|
|
|
|
|
|
| 31 |
Checkpoint format: PyTorch Lightning (weights + EMA weights + hyperparameters; optimizer state stripped). The hyperparameters carry `architecture: legacy_type_film`, which selects the paper architecture in the ShellFlow code. Distribution matching (a training-time reweighting) is disabled in the stored hyperparameters so the checkpoints load without any dataset present; it has no effect on generation.
|
| 32 |
|
| 33 |
## Usage
|
|
|
|
| 28 |
| `e1_v2_manifold_noaux/model.ckpt` | Ablation E1-V2 | on-shell (manifold) | no | 2 epochs, 48k steps |
|
| 29 |
| `e1_v3_euclidean_noaux/model.ckpt` | Ablation E1-V3 | free (Euclidean) | no | 2 epochs, 48k steps |
|
| 30 |
|
| 31 |
+
Checkpoint selection: `union_1LMET30_2to4lep/model.ckpt` holds the weights at the end of the last training epoch (stored `epoch` = 29, `global_step` = 717210). It is not the checkpoint with the lowest validation loss, which occurred at epoch 26 (`val/loss/total` = 1.1685). Epochs are counted from 0, as in PyTorch Lightning.
|
| 32 |
+
|
| 33 |
Checkpoint format: PyTorch Lightning (weights + EMA weights + hyperparameters; optimizer state stripped). The hyperparameters carry `architecture: legacy_type_film`, which selects the paper architecture in the ShellFlow code. Distribution matching (a training-time reweighting) is disabled in the stored hyperparameters so the checkpoints load without any dataset present; it has no effect on generation.
|
| 34 |
|
| 35 |
## Usage
|