File size: 18,181 Bytes
278fee4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2965d3b
278fee4
73ef2a7
 
 
 
278fee4
73ef2a7
 
 
 
278fee4
ff6c805
 
 
278fee4
13dc2e4
 
 
73ef2a7
278fee4
 
 
 
 
 
 
73ef2a7
 
 
 
 
278fee4
ff6c805
8299b46
278fee4
 
 
 
73ef2a7
 
 
 
278fee4
 
 
 
 
 
 
73ef2a7
278fee4
73ef2a7
278fee4
8299b46
73ef2a7
 
 
 
 
278fee4
8299b46
 
 
 
278fee4
8299b46
73ef2a7
 
278fee4
73ef2a7
278fee4
8299b46
 
 
278fee4
8299b46
278fee4
 
 
 
 
 
 
73ef2a7
 
278fee4
 
73ef2a7
278fee4
8299b46
73ef2a7
 
278fee4
73ef2a7
 
278fee4
 
8299b46
 
 
 
278fee4
 
73ef2a7
 
278fee4
 
8299b46
 
 
278fee4
73ef2a7
278fee4
73ef2a7
 
 
 
 
278fee4
 
 
73ef2a7
 
 
 
 
 
 
 
 
 
 
 
278fee4
 
 
 
 
 
73ef2a7
 
 
 
 
 
 
 
 
 
20d2438
 
 
73ef2a7
 
 
 
278fee4
 
 
 
 
73ef2a7
278fee4
 
 
73ef2a7
 
 
 
 
278fee4
 
 
73ef2a7
278fee4
 
 
 
 
73ef2a7
278fee4
 
 
 
 
 
73ef2a7
 
 
278fee4
73ef2a7
278fee4
73ef2a7
 
8299b46
 
 
 
 
73ef2a7
 
 
8299b46
 
 
278fee4
 
 
 
73ef2a7
 
278fee4
 
5ddae3c
278fee4
 
 
 
 
 
73ef2a7
 
278fee4
 
5ddae3c
278fee4
 
 
 
73ef2a7
 
 
278fee4
20d2438
 
 
 
 
 
 
278fee4
 
 
 
 
 
73ef2a7
 
 
 
 
 
 
 
278fee4
 
 
73ef2a7
278fee4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
73ef2a7
 
278fee4
 
 
 
 
 
 
 
 
 
73ef2a7
278fee4
 
 
 
 
 
 
 
 
 
 
 
 
 
73ef2a7
 
 
 
 
 
 
8299b46
73ef2a7
 
 
 
 
278fee4
 
 
20d2438
 
278fee4
 
 
 
 
 
 
8299b46
20d2438
 
73ef2a7
 
278fee4
 
 
 
8299b46
 
278fee4
 
2965d3b
278fee4
2965d3b
278fee4
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
---
license: apache-2.0
library_name: pytorch
pipeline_tag: graph-ml
language:
  - en
tags:
  - graph-neural-network
  - gnn
  - meshgraphnet
  - film
  - surrogate-model
  - scientific-machine-learning
  - physics-ml
  - heat-transfer
  - casting
  - metal-casting
  - foundry
  - solidification
  - finite-element
  - fenicsx
  - pytorch-geometric
  - scikit-learn
---

# Casting solidification GNN (v5.1)

A surrogate model for the solidification of metal castings in a sand mold. It takes a
tetrahedral mesh of a casting together with the alloy and mold properties, and predicts the
time at which every mesh node solidifies. The nodes that solidify last are the hot spots, where
feeding problems and shrinkage defects usually start.

A transient finite-element solve of the same field takes minutes to hours. The model returns it
in seconds, which makes it useful for screening design variants before a full simulation. It
does not replace a calibrated process simulation or physical validation, and a hot spot is a
shrinkage-risk indicator, not a porosity label.

To look at predictions without installing anything, open the [demo Space](https://huggingface.co/spaces/eugenmik/castsolid-gnn-demo): it shows
four parts in two alloys in 3D, with the hot spots and a shrinkage screening.

<p align="center">
  <video src="https://huggingface.co/eugenmik/castsolid-gnn/resolve/main/assets/solidification.mp4"
         poster="https://huggingface.co/eugenmik/castsolid-gnn/resolve/main/assets/solidification_poster.jpg"
         width="720" autoplay loop muted playsinline controls></video><br>
  <em>A predicted field played back over time. Solidified material turns translucent while the
  remaining liquid contracts toward the hot spot.</em>
</p>

| | |
|---|---|
| Version | v5.1, frozen on 31 July 2026 |
| Task | Per-node regression on a tetrahedral mesh (graph ML) |
| Input | Tetrahedral mesh in mm (gmsh 2.2 format) with the cooling surface tagged; alloy `EN-GJS-400-15` or `G20Mn5`, or a custom property set; green-sand mold |
| Output | Solidification time in seconds for every mesh node |
| Architecture | 12-block MeshGraphNet-style GNN with FiLM conditioning, combined with two gradient-boosted heads |
| Parameters | about 1.85 M (GNN) plus two `HistGradientBoostingRegressor` heads |
| Training data | FEniCSx enthalpy-method conduction simulations: about 9.7k primitive cases, 4.7k multi-alloy cases, 500 real-part cases |
| License | Apache-2.0 for code and weights |
| Demo | [castsolid-gnn-demo](https://huggingface.co/spaces/eugenmik/castsolid-gnn-demo): precomputed predictions in 3D, no installation |
| Code | [github.com/eugenmik/castsolid-gnn](https://github.com/eugenmik/castsolid-gnn): loading, inference, and the STEP-to-result pipeline |
| Author | Eugen Miknevic, independent foundry engineer |

## Contents

- [Input and output](#input-and-output)
- [How to use](#how-to-use)
- [Model architecture](#model-architecture)
- [Files](#files)
- [Training data](#training-data)
- [Training procedure](#training-procedure)
- [Evaluation](#evaluation)
- [Intended use and limitations](#intended-use-and-limitations)
- [Reproducibility](#reproducibility)
- [License and citation](#license-and-citation)

## Input and output

The model reads one casting at a time.

The mesh is a tetrahedral volume mesh in millimetres, in gmsh 2.2 ASCII format (`.msh`). Surface
triangles that cool into the mold belong to a physical group named `Cooling`; in this release
that is the whole outer surface. The model was trained on meshes of roughly 7,500 to 10,000
nodes, and that is the density it should be given. It does not use node numbering or absolute
coordinates: node inputs are distances, flags and volume-type quantities in millimetres, and
edges carry only relative displacements between neighbouring nodes.

Eight numbers describe the alloy and the mold: thermal conductivity, density, specific heat,
latent heat, liquidus, solidus, pouring temperature and mold effusivity. The code ships two
tabulated alloys, `EN-GJS-400-15` (ductile iron, the default) and `G20Mn5` (cast steel), and one
mold, green sand. You can also pass your own eight values.

The output is one number per node: the time in seconds at which that node cools through the
solidus. The maximum over all nodes is the total solidification time, and the 10% of nodes with
the largest times are the hot spots.

## How to use

This repository holds the weights and this card. The code that loads and runs them is on
[GitHub](https://github.com/eugenmik/castsolid-gnn):

```bash
git clone https://github.com/eugenmik/castsolid-gnn
cd castsolid-gnn
python -m venv .venv && source .venv/bin/activate
pip install "torch==2.3.0" --index-url https://download.pytorch.org/whl/cpu
pip install torch-cluster -f https://data.pyg.org/whl/torch-2.3.0+cpu.html
pip install -r requirements.txt
```

Python 3.11 is recommended. For a GPU, replace `cpu` with your CUDA tag (for example `cu121`)
in both URLs. Then, from the repository root:

```python
from src.model_api import CastSolidModel

model = CastSolidModel.from_pretrained()         # downloads this repository's files once
t = model.predict("part.msh")                    # EN-GJS-400-15 in green sand
t_steel = model.predict("part.msh", alloy="G20Mn5")

t.max()                                          # total solidification time, s
hot = CastSolidModel.hotspots(t)                 # boolean mask, top 10% of nodes
```

`from_pretrained()` fetches `config.json` and the files it lists into the local Hugging Face
cache, pinned to the `v5.1` tag; later calls run offline. `predict` returns a NumPy array in the
node order of the `.msh` file. To use another alloy, pass `props_json="my_alloy.json"` with
exactly these keys (SI units, temperatures in °C; 1018 is the green-sand effusivity):

```json
{"k": 32.9, "rho": 6800.0, "cp": 515.0, "L_kJ_kg": 220.0,
 "TL": 1160.0, "TS": 1120.0, "Tpour": 1260.0, "mold_effusivity": 1018.0}
```

To start from a STEP file instead of a mesh, run `python -m src.predict_casting part.step` in
the GitHub repository. It heals the CAD, builds a mesh at the right density, runs the model and
writes the result for ParaView.

## Model architecture

The model is a hybrid of three trained parts. The GNN predicts the spatial shape of the field;
two gradient-boosted heads predict its overall level and its spread; a fixed recomposition step
combines them. The GNN alone is less reliable about absolute magnitude, which depends strongly
on alloy and part size. On the original benchmark of held-out geometry families, adding the
heads brought nodal MAE down from 5.77 s to 4.26 s.

```mermaid
flowchart LR
  M[Tetrahedral mesh] --> F[15 node features<br/>+ union graph]
  P[Alloy + mold<br/>properties] --> G[5-value conditioning vector]
  P --> D
  F --> N[12-block GNN<br/>field shape]
  G --> N
  F --> D[Part descriptors]
  D --> T[GBM total head]
  D --> S[GBM spread head]
  N --> R[Soft-anchor<br/>recomposition]
  T --> R
  S --> R
  R --> O[Time per node, s]
```

### Graph network

Nodes are mesh vertices. Edges are the union of the tetrahedral mesh edges and the 16 nearest
neighbours of each node. Each edge carries its displacement vector, its length and a flag that
says whether it is a mesh edge.

Each node gets 15 features computed from the mesh and its surface tags:

- geodesic distance through the mesh to the nearest cooling surface and to the nearest
  adiabatic surface, plus straight-line distance to the surface and to adiabatic faces;
- cooling, adiabatic and interior flags;
- part volume, surface area and the fraction of surface that cools;
- local modulus, the volume-to-surface ratio inside geodesic balls of 5, 15 and 40 mm, which is
  a local version of Chvorinov's modulus;
- two boundary-condition channels: the mold effusivity on surface nodes (zero inside the
  part), and a boundary heat-input channel, which is zero because a sand-mold surface adds no
  heat.

The alloy and mold enter as five dimensionless numbers: Stefan number, normalized freezing
range, thermal diffusivity ratio, normalized superheat and normalized mold effusivity. They
condition every processor block through FiLM (`h ← h·(1+γ(g)) + β(g)`).

The network follows the encode-process-decode design of
[MeshGraphNets](https://arxiv.org/abs/2010.03409). MLP encoders embed nodes and edges, 12
message-passing blocks (hidden width 128, residual connections, LayerNorm, mean aggregation)
process them, and a two-head decoder predicts the per-node deviation and a graph-level scale
separately. The training target is the standardized `log1p` of the solidification time.

### Gradient-boosted heads

Both heads are scikit-learn `HistGradientBoostingRegressor` models with one row per part: the
alloy properties, the part's volume-to-area ratio, the mold effusivity, and boundary-condition
descriptors that are fixed to the plain sand mold in this release. The total head (16 inputs)
predicts `log1p` of the total solidification time. The spread head (17 inputs, absolute-error
loss) predicts `log1p` of the standard deviation of the nodal field.

### Soft-anchor recomposition

The GNN field is scaled, shifted and clipped at zero,

```
pred(a, m) = max(0, a · (gnn − mean(gnn)) + m)
```

with `(a, m)` chosen to minimize

```
J = w_total · ((q99.5(pred) − T_total) / T_total)²  +  w_std · ((std(pred) − S) / S)²
```

where `T_total` and `S` come from the two heads, `w_total = 1` and `w_std = 4`. When both
targets can be met exactly, this equals the closed-form affine solve. When clipping makes that
impossible, the total anchor may move by a few percent so the spread stays close. The post-clip
spread is not monotone in `a`, so the solver brackets the smallest root explicitly.

## Files

| File | Size | Content |
|---|---:|---|
| `config.json` | 1 KB | Model version, file list, architecture summary, recomposition weights and input description |
| `gnn_v5.1.safetensors` | 7.4 MB | GNN weights, feature and target normalization, and the architecture config (in the safetensors metadata) |
| `gbm_total_v3.skops` | 0.6 MB | Total-time head |
| `gbm_std_v2.skops` | 0.8 MB | Spread head |
| `ood_ranges.json` | <1 KB | Training envelope (part volume, size, distance to cooling) used for out-of-distribution warnings |

None of the weight files can execute code when loaded. The GNN is stored as
[safetensors](https://github.com/huggingface/safetensors) and the heads as
[skops](https://skops.readthedocs.io/). Before it opens a skops file, the loader on GitHub checks
every type the file names against an allowlist of scikit-learn and NumPy types. Training code is
not published.

## Training data

Every reference field was computed with [FEniCSx](https://fenicsproject.org/) (dolfinx) using an
enthalpy-method model of transient heat conduction with latent heat. There is no mold filling,
flow, gas or stress in it.

<p align="center">
  <img src="assets/3dforms.webp" alt="Examples of primitive training geometries" width="640"><br>
  <em>Some of the primitive geometry families from the first training regime.</em>
</p>

| Set | Cases | Content |
|---|---:|---|
| Primitives (original regime) | 9,710 | 42 parametric geometry families, a single alloy, sand mold, different assignments of cooling and adiabatic faces |
| Multi-alloy (v4) | 4,727 | Realistic thick-walled castings in 12 named iron and steel grades plus sampled synthetic alloys, under several boundary-condition regimes |
| Real parts (E) | 500 | Prepared production geometries with varied boundary conditions |

<p align="center">
  <img src="assets/mesh.webp" alt="Tetrahedral training mesh with tagged boundary groups" width="640"><br>
  <em>Training meshes carry three physical groups: Volume, Cooling (mold contact) and
  Adiabatic.</em>
</p>

Training meshes have roughly 7,500 to 10,000 nodes. The target at each node is the time its
temperature drops through the solidus. Material properties come from handbook data. The raw
simulation corpus (about 100 GB) is not distributed.

One data issue shaped this project more than any model choice. The solver stored temperatures
in dolfinx degree-of-freedom order, while coordinates and geometric features were in gmsh node
order, so about 60% of the nodal targets were silently attached to the wrong positions. Every
early model failure traced back to it. The training set was rebuilt with a recovered
permutation between the two orderings. If you train on solver output, check node identity across
every solver and mesh interface before tuning a model.

## Training procedure

v5.1 came out of a chain of fine-tunes, not a single training run.

1. SP1 trained a 12-block two-head GNN from scratch on the primitive set, with the union graph,
   local-modulus features and a loss on both nodal and total time.
2. FT2b fine-tuned it on the multi-alloy set. The FiLM layers started at zero, so training began
   from exactly the SP1 function.
3. A follow-up pass re-encoded one sparse boundary input as a field diffused along the mesh,
   because the network had learned to suppress it, and picked the checkpoint on gate metrics
   instead of validation loss alone.
4. FT3 fine-tuned on real production geometries (set E) with the normalization statistics
   frozen. Its epoch-21 checkpoint is the v5.1 GNN.
5. The total head (v3), the absolute-error spread head (v2) and the recomposition weights were
   selected on held-out real-part validation rows.

Training ran on one RTX 3060 (12 GB) with batch size 2, at 15 to 20 minutes per epoch. To be
promoted, v5.1 had to pass four gates on held-out data: two gold-standard real parts, 24
held-out real parts, an extrapolation split and a check against collapsed field variance.

## Evaluation

Do not compare numbers across regimes. Total times in the primitive regime are tens of seconds;
on real geometry they run to hundreds or thousands.

| Evaluation | n | Nodal MAE | Hot-spot overlap | What it shows |
|---|---:|---:|---:|---|
| Held-out primitive geometry families | family split | 4.26 s | 0.917 | Generalization to unseen synthetic shape families in the original regime |
| Unseen real single-body castings | 6 | 80.3 s median | 0.738 median | Generalization to new real geometry, scored against a FEniCSx reference on the same mesh |

Hot-spot overlap compares the 10% of nodes that solidify last in the prediction with the same
set in the reference. 1.0 means the two sets are identical.

One of the six real parts reached a hot-spot overlap of only 0.035, so the median says nothing
certain about any one part.

> **Important.** The training data and the real-geometry references come from the same FEniCSx
> heat-transfer model, the same material database and the same boundary assumptions. These
> numbers measure how well the surrogate carries over to new geometry. They are not evidence of
> physical accuracy against instrumented castings.

## Intended use and limitations

Good uses:

- comparing early casting concepts;
- finding candidate last-to-solidify regions;
- deciding which designs need a full process simulation;
- early quoting and design discussions, as long as the limitations stay visible.

Out of scope:

- final process sign-off or contractual acceptance;
- guaranteed defect or porosity prediction;
- mold filling, flow, gas, stress or distortion;
- unattended decisions on geometry outside the training envelope;
- safety-critical decisions without an independent engineering review.

Known limitations:

1. The physics is pure conduction with latent heat, and the reference solver shares the model's
   property and boundary assumptions.
2. This release covers one setup: a single casting whose whole outer surface cools into a
   green-sand mold, in `EN-GJS-400-15`, `G20Mn5` or a custom property set.
3. Meshes much coarser or finer than the training density (about 7,500 to 10,000 nodes) are
   outside what the model has seen. At that density, large thin-walled parts can be
   under-resolved.
4. The training envelope in `ood_ranges.json` only covers part size and distance to
   cooling. It is not a calibrated uncertainty estimate.

Treat each prediction as a screening hypothesis: where should an engineer look first, and is
the concept ready for a full simulation? Keep the mesh, the inputs and the model version with
every result, and send high-consequence cases to a calibrated simulation and a physical review.

## Reproducibility

The published files were converted from the internal training checkpoints and checked against
them:

| Release file | Internal artifact | Check |
|---|---|---|
| `gnn_v5.1.safetensors` | `ckpt_ep021.pt` (FT3, epoch 21) | Every tensor bit-identical; same config and normalization |
| `gbm_total_v3.skops` | `hgb_v3.joblib` | Identical predictions on 10,002 probe rows spanning each feature's training split range |
| `gbm_std_v2.skops` | `hgb_std_v2.joblib` | Identical predictions on 10,002 probe rows spanning each feature's training split range |

The published inference code was also run side by side with the internal pipeline on five
production parts (two of them with both alloys), plus a fixed-size mesh with a custom property
set: every output field matched exactly. The heads were trained with scikit-learn 1.5.2, and other versions may refuse to load them or change
their output. Other versions used: Python 3.11, PyTorch 2.3.0, PyTorch Geometric 2.7.0,
NumPy 2.x.

## License and citation

Code and weights are released under the [Apache License 2.0](LICENSE). The tables in
`data/properties/` of the GitHub repository are modelling inputs, not certified material
specifications.

```bibtex
@software{miknevic_casting_gnn_2026,
  author  = {Miknevic, Eugen},
  title   = {Casting Solidification GNN: a graph neural network surrogate for
             rapid solidification screening of metal castings},
  version = {5.1},
  year    = {2026},
  url     = {https://huggingface.co/eugenmik/castsolid-gnn}
}
```

Contact: Eugen Miknevic, eugenmiknevic@gmail.com. If you have failure cases, questions or
measured data from real castings, please post them in the
[Community tab](https://huggingface.co/eugenmik/castsolid-gnn/discussions).