STAR-XPC
STAR-XPC predicts how much of an X-ray beam passes through a polymer composite shield. You give it a polyethylene slab loaded with fillers, a slab thickness and a tube voltage. It returns the transmission ratio I/I₀ in milliseconds. The penEasy (PENELOPE) Monte Carlo simulations it was trained to replace take hours per point. The training data were generated with attenupy, which designs, runs and collects these penEasy shielding studies.
The composite is treated as a set of ingredients. Each ingredient is described by a ChemBERTa embedding of its SMILES and its weight fraction, and a Set Transformer maps the set plus four physics features to I/I₀. Because the chemistry enters through SMILES, you can also register materials the model was not trained on.
Quickstart
pip install torch numpy safetensors huggingface_hub
import sys
from huggingface_hub import snapshot_download
repo = snapshot_download("cmpe-netlab41/STAR-XPC")
sys.path.insert(0, repo)
from star_xpc import STARXPC
model = STARXPC.from_pretrained(repo)
# 40 wt% Bi + 30 wt% W in polyethylene (the remaining 30 wt%), 80 kVp beam
y = model.predict({"Bi": 0.40, "W": 0.30}, thickness_mm=[0.25, 0.5, 1.0, 1.5], kvp=80)
print(y) # I/I0 at each thickness
predict works out the composite's density and mean excitation energy from the
composition, so the composition, thickness and voltage are all you need.
Inputs
| Argument | Meaning | Unit |
|---|---|---|
fillers |
Filler weight fractions of the whole composite, e.g. {"Bi2O3": 0.4, "BaSO4": 0.4}. Up to 10 fillers |
fraction (0.4 = 40 wt%) |
thickness_mm |
Slab thickness, scalar or array | mm |
kvp |
X-ray tube voltage, scalar or array (broadcast with thickness) | kVp |
polymer_fraction |
Polyethylene weight fraction. Defaults to 1 - sum(fillers) |
fraction |
mass_den |
Composite density. Computed if omitted | g/cm³ |
mean_excit |
Composite mean excitation energy. Computed if omitted | eV |
Materials
The model ships with the materials it was trained on:
- Fillers: B, B4C, BN, Ba, BaSO4, Bi, Bi2O3, Cu, Sb, Sn, W, WO3
- Matrix: polyethylene (C2H4)
Registering a new material
Any other filler can be registered from its SMILES. It is embedded with ChemBERTa
(seyonec/ChemBERTa-zinc-base-v1, pinned revision) using exactly the recipe the
training materials went through. For example, lead and gadolinium, which are not
in the training data:
# pip install transformers
model.add_material("Pb", smiles="[Pb]", density=11.34)
model.add_material("Gd", smiles="[Gd]", density=7.90)
model.predict({"Pb": 0.70}, thickness_mm=[0.3, 0.6, 1.2], kvp=100)
model.predict({"Gd": 0.40, "W": 0.30}, thickness_mm=1.0, kvp=120)
Registered this way, Pb and Gd reach NMSE % 0.14 and 0.16 against penEasy at 100 kVp. That is in line with the held-out sets of trained materials (see results).
- Formula: read from the name when the name is a chemical formula (
"Pb","Gd2O3"). Otherwise passformula="...". - Density: in g/cm³, used to compute the composite density. Without a
formula or density, pass
mass_den=andmean_excit=topredictinstead. - Style: write SMILES the way the training materials were written. Use pure
elements as bracket atoms (
[Pb]), oxides as covalent structures (O=[Gd]O[Gd]=O,O=[W](=O)=O), and salts as ions ([Ba++].[O-]S(=O)(=O)[O-]). Keep SMILES under 64 tokens, since longer ones are truncated and you get a warning. - Unregistered names: using one raises an error that shows the
add_materialcall to make. - Accuracy:
predictwarns once for each material that is not in the training data. Treat those predictions as estimates and confirm important ones with Monte Carlo. attenupy sets up and runs those simulations.
Physics features
- Mean excitation energy: the Bragg additivity rule with PENELOPE's elemental
I-values for Z = 1–99, so it also works for registered materials. This is what
PENELOPE's
materialprogram does, and it matches the training data to within 10⁻⁴. - Density: the inverse rule of mixtures. It uses the training densities for the bundled materials and the density you pass for registered ones.
- If you know better: some boron-containing training samples used a slightly
different boron density. If you have the exact values from PENELOPE, pass
mass_den=andmean_excit=.
Batch prediction from a table
predict_frame takes a pandas DataFrame in the training-data schema:
molecule_1..10, ratio_1..10 (filler wt%), polymer, polymer_ratio
(fraction), geo_number_mm (in units of 0.01 mm, so 100 = 1 mm), spc
(kVp), mass_den and mean_excit.
import pandas as pd
y = model.predict_frame(pd.read_csv("my_samples.csv"))
Training domain and limitations
- Target: the energy deposited in an air detector behind the slab, divided by the same quantity with the slab replaced by air. It comes from penEasy simulations of a fixed slab-and-detector geometry.
- Beams: four spectra. 30 kVp is ISO N-30, 80 kVp is IEC RQR6, 100 kVp is IEC RQR8 and 120 kVp is IEC RQR9. Other voltages are interpolation between those spectra and are untested.
- Thickness: 0.15–2.5 mm.
- Composition: polyethylene at 20–40 wt%, with one to several fillers making up the remaining 60–80 wt%.
- Output range: the output layer is unbounded, so predictions near 0 or 1 can land slightly outside that range.
- Warnings:
predictwarns when an input falls outside the training domain.
Held-out test results
None of these test sets were used in training. Each has 10 thicknesses.
| Test set | NMSE % | RMSE | MAPE (10 % trimmed) |
|---|---|---|---|
| Bi–W, 80 kVp | 0.153 | 0.0119 | 7.59 % |
| W–Cu–B, 100 kVp | 0.022 | 0.0074 | 1.04 % |
| Bi–Sb, 120 kVp | 0.051 | 0.0111 | 2.85 % |
| Bi₂O₃–BaSO₄, 80 kVp | 0.209 | 0.0147 | 5.15 % |
| BaSO₄–WO₃, 100 kVp | 0.023 | 0.0076 | 1.98 % |
| Bi₂O₃, 120 kVp | 0.013 | 0.0058 | 1.09 % |
| Pb 70 wt% (registered), 100 kVp | 0.141 | 0.0143 | 4.93 % |
| Gd 70 wt% (registered), 100 kVp | 0.163 | 0.0177 | 4.28 % |
NMSE % = 100 · MSE / mean(y²).
How it works
Each composite becomes a set of 11 slots (up to 10 fillers plus the polymer, zero-padded). Each slot is a 773-d vector:
- the ingredient's 768-d ChemBERTa embedding, which is the mean of the last 4 hidden layers, masked-mean-pooled over tokens and L2-normalised;
- its share of the composition;
- four z-scored global features: tube voltage, composite density, thickness and mean excitation energy.
A Set Transformer (2 ISAB blocks with 64 inducing points, PMA with 8 seeds and a SAB; width 256, 32 heads) turns the set into a single output.
In the training data, filler amounts were stored in wt% and the polymer amount
as a fraction, and the two were renormalised together. star_xpc.py
reproduces that encoding, so give predict ordinary weight fractions.
Files
| File | Contents |
|---|---|
star_xpc.py |
Model definition and the STARXPC predictor |
model.safetensors |
Set Transformer weights |
embeddings.safetensors |
ChemBERTa embeddings of the 13 training materials |
config.json |
Architecture, feature normalisation statistics, material table, training domain |
Citation
- Downloads last month
- 16