STAR-XPC

STAR-XPC predicts how much of an X-ray beam passes through a polymer composite shield. You give it a polyethylene slab loaded with fillers, a slab thickness and a tube voltage. It returns the transmission ratio I/I₀ in milliseconds. The penEasy (PENELOPE) Monte Carlo simulations it was trained to replace take hours per point. The training data were generated with attenupy, which designs, runs and collects these penEasy shielding studies.

The composite is treated as a set of ingredients. Each ingredient is described by a ChemBERTa embedding of its SMILES and its weight fraction, and a Set Transformer maps the set plus four physics features to I/I₀. Because the chemistry enters through SMILES, you can also register materials the model was not trained on.

Quickstart

pip install torch numpy safetensors huggingface_hub
import sys
from huggingface_hub import snapshot_download

repo = snapshot_download("cmpe-netlab41/STAR-XPC")
sys.path.insert(0, repo)
from star_xpc import STARXPC

model = STARXPC.from_pretrained(repo)

# 40 wt% Bi + 30 wt% W in polyethylene (the remaining 30 wt%), 80 kVp beam
y = model.predict({"Bi": 0.40, "W": 0.30}, thickness_mm=[0.25, 0.5, 1.0, 1.5], kvp=80)
print(y)  # I/I0 at each thickness

predict works out the composite's density and mean excitation energy from the composition, so the composition, thickness and voltage are all you need.

Inputs

Argument Meaning Unit
fillers Filler weight fractions of the whole composite, e.g. {"Bi2O3": 0.4, "BaSO4": 0.4}. Up to 10 fillers fraction (0.4 = 40 wt%)
thickness_mm Slab thickness, scalar or array mm
kvp X-ray tube voltage, scalar or array (broadcast with thickness) kVp
polymer_fraction Polyethylene weight fraction. Defaults to 1 - sum(fillers) fraction
mass_den Composite density. Computed if omitted g/cm³
mean_excit Composite mean excitation energy. Computed if omitted eV

Materials

The model ships with the materials it was trained on:

  • Fillers: B, B4C, BN, Ba, BaSO4, Bi, Bi2O3, Cu, Sb, Sn, W, WO3
  • Matrix: polyethylene (C2H4)

Registering a new material

Any other filler can be registered from its SMILES. It is embedded with ChemBERTa (seyonec/ChemBERTa-zinc-base-v1, pinned revision) using exactly the recipe the training materials went through. For example, lead and gadolinium, which are not in the training data:

# pip install transformers
model.add_material("Pb", smiles="[Pb]", density=11.34)
model.add_material("Gd", smiles="[Gd]", density=7.90)

model.predict({"Pb": 0.70}, thickness_mm=[0.3, 0.6, 1.2], kvp=100)
model.predict({"Gd": 0.40, "W": 0.30}, thickness_mm=1.0, kvp=120)

Registered this way, Pb and Gd reach NMSE % 0.14 and 0.16 against penEasy at 100 kVp. That is in line with the held-out sets of trained materials (see results).

  • Formula: read from the name when the name is a chemical formula ("Pb", "Gd2O3"). Otherwise pass formula="...".
  • Density: in g/cm³, used to compute the composite density. Without a formula or density, pass mass_den= and mean_excit= to predict instead.
  • Style: write SMILES the way the training materials were written. Use pure elements as bracket atoms ([Pb]), oxides as covalent structures (O=[Gd]O[Gd]=O, O=[W](=O)=O), and salts as ions ([Ba++].[O-]S(=O)(=O)[O-]). Keep SMILES under 64 tokens, since longer ones are truncated and you get a warning.
  • Unregistered names: using one raises an error that shows the add_material call to make.
  • Accuracy: predict warns once for each material that is not in the training data. Treat those predictions as estimates and confirm important ones with Monte Carlo. attenupy sets up and runs those simulations.

Physics features

  • Mean excitation energy: the Bragg additivity rule with PENELOPE's elemental I-values for Z = 1–99, so it also works for registered materials. This is what PENELOPE's material program does, and it matches the training data to within 10⁻⁴.
  • Density: the inverse rule of mixtures. It uses the training densities for the bundled materials and the density you pass for registered ones.
  • If you know better: some boron-containing training samples used a slightly different boron density. If you have the exact values from PENELOPE, pass mass_den= and mean_excit=.

Batch prediction from a table

predict_frame takes a pandas DataFrame in the training-data schema: molecule_1..10, ratio_1..10 (filler wt%), polymer, polymer_ratio (fraction), geo_number_mm (in units of 0.01 mm, so 100 = 1 mm), spc (kVp), mass_den and mean_excit.

import pandas as pd
y = model.predict_frame(pd.read_csv("my_samples.csv"))

Training domain and limitations

  • Target: the energy deposited in an air detector behind the slab, divided by the same quantity with the slab replaced by air. It comes from penEasy simulations of a fixed slab-and-detector geometry.
  • Beams: four spectra. 30 kVp is ISO N-30, 80 kVp is IEC RQR6, 100 kVp is IEC RQR8 and 120 kVp is IEC RQR9. Other voltages are interpolation between those spectra and are untested.
  • Thickness: 0.15–2.5 mm.
  • Composition: polyethylene at 20–40 wt%, with one to several fillers making up the remaining 60–80 wt%.
  • Output range: the output layer is unbounded, so predictions near 0 or 1 can land slightly outside that range.
  • Warnings: predict warns when an input falls outside the training domain.

Held-out test results

None of these test sets were used in training. Each has 10 thicknesses.

Test set NMSE % RMSE MAPE (10 % trimmed)
Bi–W, 80 kVp 0.153 0.0119 7.59 %
W–Cu–B, 100 kVp 0.022 0.0074 1.04 %
Bi–Sb, 120 kVp 0.051 0.0111 2.85 %
Bi₂O₃–BaSO₄, 80 kVp 0.209 0.0147 5.15 %
BaSO₄–WO₃, 100 kVp 0.023 0.0076 1.98 %
Bi₂O₃, 120 kVp 0.013 0.0058 1.09 %
Pb 70 wt% (registered), 100 kVp 0.141 0.0143 4.93 %
Gd 70 wt% (registered), 100 kVp 0.163 0.0177 4.28 %

NMSE % = 100 · MSE / mean(y²).

How it works

Each composite becomes a set of 11 slots (up to 10 fillers plus the polymer, zero-padded). Each slot is a 773-d vector:

  • the ingredient's 768-d ChemBERTa embedding, which is the mean of the last 4 hidden layers, masked-mean-pooled over tokens and L2-normalised;
  • its share of the composition;
  • four z-scored global features: tube voltage, composite density, thickness and mean excitation energy.

A Set Transformer (2 ISAB blocks with 64 inducing points, PMA with 8 seeds and a SAB; width 256, 32 heads) turns the set into a single output.

In the training data, filler amounts were stored in wt% and the polymer amount as a fraction, and the two were renormalised together. star_xpc.py reproduces that encoding, so give predict ordinary weight fractions.

Files

File Contents
star_xpc.py Model definition and the STARXPC predictor
model.safetensors Set Transformer weights
embeddings.safetensors ChemBERTa embeddings of the 13 training materials
config.json Architecture, feature normalisation statistics, material table, training domain

Citation

Downloads last month
16
Safetensors
Model size
3M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support