OMatGRPO weights and evaluated structures
Weights and generated structures for the paper "Reinforcement Learning on the Discrete Composition Channel of a
Crystal Generator: Validated Gains and Reward Hacking" (arXiv:2610.03880). The code is at
https://github.com/paprakash/OMatGRPO, and the weights load with it. Its scripts/download_assets.py downloads
these files, checks them against SHA256SUMS and unpacks the CIFs.
Contents
prior/prior.safetensors the pretrained OMatG model every run starts from (KL reference)
prior/train.yaml its OMatG configuration
models/<identifier>/final_model.safetensors weights at the end of training (750 rollouts)
models/<identifier>/resolved_config.json the run's settings, as the code's --print_config prints them
structures/<identifier>/cifs.zip the CIFs that LeMat-GenBench scored for the paper (folder cifs/)
structures/<identifier>/structures_summary.csv one row per generated structure (see below)
structures/<identifier>/<identifier-in-file>_comprehensive_multi_mlip_hull_<time>.json the LeMat-GenBench result
SHA256SUMS sha256 of every file
The two MP-20 references of the reward are not included. scripts/build_references.py rebuilds them from
public data.
Weights
All weights are float32 state dicts of the OMatG network (73 tensors, 12,403,912 values), converted from the PyTorch checkpoints of the paper's runs. Every tensor equals its source exactly.
The prior is our own OMatG model, pretrained on MP-20 for de novo generation with stochastic position and lattice
channels and masked discrete flow matching for species. It is not one of the models released by the OMatG
authors. prior/train.yaml has its configuration. The hardware, wall time and seed of the pretraining were not
recorded.
| identifier | paper name | training run |
|---|---|---|
arityguard_creatrelax |
OMatGRPO | arityguard_creatrelax_750 |
canonical_creatrelax |
discovery | canonical_creatrelax_750 |
sparseworst_creatrelax |
penalty routing | sparseworst_creatrelax_750 |
arityguard |
guarded, pre-creativity | canonical_arityguard_750 |
canonical |
discovery, pre-creativity | canonical_species_rl_750 |
sparseworst |
penalty routing, pre-creativity | canonical_sparseworst_750 |
frozen_control |
frozen-composition control | frozen_species_control_750 |
Each identifier is also a run config in the code (configs/runs/<identifier>.yaml). All runs used seed 0, 4 groups
of 16 structures per rollout and 750 rollouts.
Structures
These are the exact sets behind the paper's LeMat-GenBench numbers, copied from the original evaluation folders
(not regenerated). Each set was generated with the paper's protocol: 2,500 structures, consistent sampler (64
time points, species noise 0), seed 42, chunks of 100, relaxation with FIRE and UMA uma-s-1p2 including the cell
(force tolerance 0.05 eV/Å, at most 500 steps), cell and masked-species guards before and after the relaxation.
The best-of-N set is the exception for generation: its pool of 48,000 structures was sampled from the prior with
seed 4242, and the 2,500 selected structures were then relaxed and scored with the same protocol.
cifs.zip holds the folder cifs/ with the relaxed structures that passed both guards and have a finite E_hull.
The CIFs are zipped because a Hugging Face repository holds at most 20,000 files. LeMat-GenBench drops CIFs that
its pymatgen cannot parse, so its structure count can be slightly lower than the number of files.
| identifier | what it is | CIFs | mSUN (LeMat-GenBench) |
|---|---|---|---|
R0 |
the prior | 2497 | 336 |
bestofn48k_top2500 |
best-of-N baseline: the top 2,500 of 48,000 structures sampled from the prior | 2500 | 192 |
frozen_control |
frozen-composition control | 2495 | 245 |
canonical |
discovery, pre-creativity | 2490 | 1063 |
canonical_creatrelax |
discovery | 2463 | 939 |
sparseworst |
penalty routing, pre-creativity | 2494 | 789 |
sparseworst_creatrelax |
penalty routing | 2495 | 841 |
arityguard |
guarded, pre-creativity | 2486 | 982 |
arityguard_creatrelax |
OMatGRPO | 2485 | 1138 |
LeMat-GenBench ran at commit 58e6eae3e4a6c87c22171cf069123ecc4e2fa7e6 with the preset
comprehensive_multi_mlip_hull, and all three potentials of its stability ensemble (ORB, MACE, UMA) scored every
valid structure. mSUN counts metastable, unique and novel structures and excludes structures on the hull.
structures_summary.csv columns: idx (index in the generated set of 2,500), formula, n_elements, eh_rel
(E_hull of the relaxed structure against the LeMat-Bulk UMA hull, eV/atom), ref_rel (hull reference entries
in its chemical system), valid, trusted, sparse (fewer than 12 reference entries), deep_below (E_hull below
-0.1 eV/atom), unique, novel_mp20, novel_alexmp, msun_mp20, msun_alexmp (an internal check against
MP-20 and Alex-MP-20, not LeMat-GenBench), and cif (file name in cifs/, empty for structures that were not
exported).
License
MIT for the weights. The structures were computed with the UMA potential (facebook/UMA, distributed under its
own license) and are derived from a model trained on MP-20, which comes from the Materials Project (data licensed
CC BY 4.0; A. Jain et al., 2013, https://doi.org/10.1063/1.4812323; MP-20: T. Xie et al., 2021,
https://arxiv.org/abs/2110.06197). E_hull values use the LeMat-Bulk-MLIP-Hull convex hull (LeMaterial,
https://huggingface.co/datasets/LeMaterial/LeMat-Bulk-MLIP-Hull), which is not included here.
Citation
Please cite the paper and OMatG.
@misc{prakash2026reinforcementlearningdiscretecomposition,
title={Reinforcement Learning on the Discrete Composition Channel of a Crystal Generator: Validated Gains and Reward Hacking},
author={Pawan Prakash and Philipp Höllmer and Addis Fuhr and Peter Hirschfeld and P. Ganesh and Stefano Martiniani and Richard Hennig},
year={2026},
eprint={2610.03880},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2610.03880},
}
@inproceedings{hoellmer2025,
title={Open Materials Generation with Stochastic Interpolants},
author={Philipp H{\"o}llmer and Thomas Egg and Maya Martirossyan and Eric
Fuemmeler and Zeren Shui and Amit Gupta and Pawan Prakash and Adrian
Roitberg and Mingjie Liu and George Karypis and Mark Transtrum and Richard
Hennig and Ellad B. Tadmor and Stefano Martiniani},
booktitle={Forty-second International Conference on Machine Learning},
year={2025},
url={https://openreview.net/forum?id=gHGrzxFujU},
archivePrefix={arXiv},
eprint={2502.02582},
primaryClass={cs.LG},
}