CORD-PHM
CORD: Learning Reusable Degradation Representations Across Heterogeneous Physical Systems
- Paper: arXiv:2609.39784
- DOI: 10.48550/arXiv.2609.39784
- Full training and evaluation code: github.com/HelpLee/CORD-PHM
- Authors: Haibo Li and Zhiguo Zeng
- Affiliation: CentraleSupélec, Université Paris-Saclay
CORD learns reusable degradation representations across heterogeneous physical systems. The released checkpoint is the single selected multi-domain encoder used for the paper's CORD (Multi-domain) results.
Model description
CORD combines system-type-specific observation interfaces with a shared degradation backbone. The interfaces retain the different channel structures and measurement semantics of bearings, batteries, and cutting tools while mapping all three physical system types into a shared representation space.
The encoder is trained with two complementary self-supervised objectives:
- Intra-Observation Structure Modeling (ISM): masks valid local descriptors and reconstructs their observed components.
- Inter-Observation Dynamics Modeling (IDM): predicts the next observation embedding from a history of six observations.
The released checkpoint contains the reusable encoder. It does not include a downstream RUL prediction head because CORD uses target-aware temporal heads for different physical system types and datasets.
Released checkpoint
| Field | Value |
|---|---|
| File | cord_multidomain_encoder.pt |
| Training experiment | E37 |
| Training arm | joint_bearingfloor_cagrad |
| Selected epoch | 545 |
| Selection rule | Best macro source-validation score |
| File size | 942,828 bytes |
| SHA256 | 09422359745CF47C69D94A241D92B9018A940A9DCBCC7804A3D08C86A2485103 |
| Pretraining system types | Bearing, battery, cutting tool |
| Output size | 96 |
The single-domain checkpoints used as experimental controls are intentionally excluded from this model repository.
Architecture
| Component | Configuration |
|---|---|
| Structured descriptor size | 26 features |
| Local slots | 64 per physical channel |
| Hidden size | 96 |
| Transformer blocks | 2 |
| Attention heads | 4 |
| Feed-forward size | 192 |
| Residual adapter bottleneck | 24 |
| Dropout | 0.1 |
| ISM mask ratio | 0.3 |
| IDM history | 6 observations, followed by one target observation |
Each represented system type has separate local/global stems and residual adapters. Validity-aware channel attention handles variable channel counts. The Transformer blocks, final normalization, and observation projector are shared.
This implementation operates on structured health-state descriptors; it does not use a raw-signal patch size. The complete machine-readable configuration is in config.json.
Input and output contract
For a selected domain, the encoder expects:
| Argument | Shape | Description |
|---|---|---|
local_x |
[batch, channels, 64, 26] |
Local structured descriptors |
global_x |
[batch, channels, 26] |
Global structured descriptors |
channel_mask |
[batch, channels] |
Valid physical channels |
token_mask |
[batch, channels, 64] |
Valid local slots |
domain must be bearing, battery, or milling.
The encoder returns:
embedding:[batch, 96], the reusable observation representation.local_hidden:[batch, 64, 96], the final local-token representations after validity-aware channel pooling.
Installation
pip install torch huggingface_hub
Download the repository with:
huggingface-cli download haibo-lgi/CORD-PHM --local-dir CORD-PHM
Load the encoder
import torch
from joint_model import JointModel
encoder = JointModel().encoder
state = torch.load(
"cord_multidomain_encoder.pt",
map_location="cpu",
weights_only=True,
)
encoder.load_state_dict(state, strict=True)
encoder.eval()
with torch.inference_mode():
embedding, local_hidden = encoder(
"bearing",
local_x,
global_x,
channel_mask,
token_mask,
)
Run the included synthetic-input verification after downloading:
python example_load.py
It strictly loads the checkpoint and checks output shapes for bearing, battery, and milling inputs.
Pretraining and checkpoint selection
CORD jointly pretrains on 19 source datasets from three represented physical system types. Source and downstream datasets are separated as described in the paper and project manifests.
The selected checkpoint uses CAGrad to coordinate gradients from the three system types. A minimum Euclidean correction enforces the specified bearing-gradient directional floor before clipping and AdamW. ISM routing coefficients are 0.3, 0.3, and 1.0 for bearing, battery, and cutting-tool data; IDM gradients retain full strength. Source validation averages dataset loss ratios within each system type and then averages across system types. One joint checkpoint is selected for all represented types.
The relevant training configuration, bearing-floor implementation, and completion record are included under training/.
Evaluation summary
The paper evaluates the representation at two transfer boundaries:
- Pretraining-Included System Types: target datasets and units are unseen, while their broad physical system types are represented during source pretraining.
- Pretraining-Excluded System Types: the complete turbofan-engine system type is absent from source pretraining and introduced during adaptation.
Across bearings, batteries, and cutting tools, CORD (Multi-domain) improves over the matched CORD (Single-domain) initialization in all nine Frozen settings and eight of nine Full-FT settings. Source-pretrained initialization also improves low-label adaptation to the previously unseen turbofan-engine type. See the paper for complete protocols, baselines, uncertainty summaries, and per-dataset results.
Intended use
This checkpoint is intended for research on:
- degradation representation learning;
- prognostics and health management;
- remaining-useful-life prediction;
- predictive maintenance;
- transfer across devices, datasets, and physical system types.
For a represented system type, reuse the matching type-specific interface and attach a task-appropriate head. For a new system type, initialize a new input interface and downstream model while reusing the shared pretrained backbone, following the Level-II protocol in the paper.
Limitations
- Inputs must first be converted into CORD's structured 64-by-26 observation format using the project preprocessing code.
- The released file is an encoder checkpoint, not a ready-to-run end-to-end RUL predictor.
- Source pretraining uses one selected upstream seed; downstream experiments use multiple seeds, but the release does not measure pretraining-seed variability.
- Physical systems with substantially different degradation processes or observation construction may require a new interface and additional adaptation.
- The repository does not redistribute source or downstream datasets.
Repository contents
| Path | Purpose |
|---|---|
cord_multidomain_encoder.pt |
Selected multi-domain CORD encoder |
joint_model.py |
Joint encoder and self-supervised model definition |
compression_model.py |
Validity-aware channel pooling and base encoder |
global_local_model.py |
Shared Transformer backbone components |
masking.py |
ISM structured masking |
config.json |
Machine-readable model and input configuration |
example_load.py |
Strict checkpoint-loading smoke test |
training/ |
E37 selection status and gradient-routing implementation |
License
The arXiv paper is distributed under CC BY 4.0. No separate software or model-weight license has been specified in this repository. The paper license should not be interpreted as an additional license for the code or model weights.
Citation
@article{li2026cord,
title = {{CORD}: Learning Reusable Degradation Representations Across Heterogeneous Physical Systems},
author = {Li, Haibo and Zeng, Zhiguo},
journal = {arXiv preprint arXiv:2609.39784},
year = {2026},
doi = {10.48550/arXiv.2609.39784},
url = {https://arxiv.org/abs/2609.39784}
}
- Downloads last month
- 30