PDE-OBS: trained baseline checkpoints
Final checkpoints of the PDE-OBS campaign: one model per (PDE family, method, training observation pattern). Each checkpoint is the final recorded checkpoint of its attempt, not one selected on validation or test error.
All 441 credited settings are published here. models_manifest.<cluster-label>.json lists the SHA-256
of every published file; the benchmark repository binds each published checkpoint digest to the checkpoint
identity its results index uses (results/public_deposits/release_map.json there).
This repository accompanies an anonymous submission under double-blind review.
Layout
models/<cluster-label>/<pde>/<method>/<train_view>/
checkpoints/last.pt final checkpoint: weights, scheduler state, history, resolved config
checkpoints/training_config.json resolved trainer configuration
identity.json provenance.json resolved.yaml split_manifest.json factor_coverage.json
completion.json health.json history.json (where the attempt wrote them)
budget-protocol.json | dynamic-protocol.json (cohort dependent)
models_manifest.<cluster-label>.json per-file SHA-256 of the published files (repository root)
scrub-manifest.json which record files the last de-identification pass rewrote
The cluster labels (cluster-A, cluster-B, cluster-C) stand for the three machines the campaign
ran on; they carry no institutional meaning. resolved.yaml beside each checkpoint records the
architecture (method.name, method.kwargs), the training view and the training settings.
Scoring a checkpoint with the benchmark
from pdeobs import api
data = api.load_dataset("./pdeobs-data", verify=True) # from PDE-OBS/pdeobs-data
package, targets = api.inference_input_from_dataset(data, api.make_observation("paper:R50"), task="recovery")
pred = api.load_legacy_checkpoint("models/cluster-A/poisson/fno/random_50pct/checkpoints/last.pt",
model={"name": "fno", "preset": "paper"}, task="recovery")
score = api.evaluate(api.predict(pred, package), targets)
score["status"], score["scoring_version"], score["summary"]["rel_l2_joint_mean"]
The loader takes the structure from the caller (the campaign checkpoints use the paper preset of
their method), cross-checks the checkpoint's own stored training configuration, and records the
file's SHA-256 and epoch in the predictor's provenance. Without the benchmark, the file is a plain
PyTorch payload: torch.load(path, map_location="cpu", weights_only=False) returns a dict with
model_state, config, history, epoch and scheduler_state.
Important properties
- Weights only. Optimizer, gradient-scaler and RNG states were removed (about two thirds of each
file);
dropped_keysin the manifest records this per checkpoint. The checkpoints support inference and rescoring, not exact continuation of training. - Anonymized metadata, unmodified parameters. Every string naming a filesystem path, account, node,
host, login node, scheduler job, partition, GPU hardware identifier, interpreter or environment
directory, run directory or source-code revision was rewritten to a placeholder such as
<ROOT_A>,<USER>,<NODE>,<JOB>,<PYTHON>,<ENV>,<RUN>or<COMMIT-1>, and raw scheduler records were reduced to their resource fields (<SCHEDULER_RECORD> ...); three passes, the files rewritten by the last one are listed inscrub-manifest.json. Model parameters were not modified.models_manifest.*.jsonrecords the SHA-256 of each released file; the correspondence to the campaign's checkpoint identities is kept in the benchmark repository. - Training cohorts are not interchangeable. Each entry carries its cohort label: the original 500-epoch protocol, a recovered continuation of it, a declared 200-epoch budget, a budgeted max-200 patience protocol, or a salvaged interrupted attempt. Any table that pools cohorts must say so per row.
- Scoring provenance. All 441 checkpoints were later rescored from retained prediction arrays with the
benchmark's strict scorer (
pdeobs-strict-v1, nine views x 200 held-out identities each); those results and their per-identity errors are in the benchmark repository underresults/prediction_verification_20260924/. The campaign's own evaluator metrics remain in the archived-results index and are not the paper's current numbers.
Citation
Anonymous submission under review. Please cite the paper once it is public.