OMatGRPO weights and evaluated structures

Weights and generated structures for the paper "Reinforcement Learning on the Discrete Composition Channel of a Crystal Generator: Validated Gains and Reward Hacking" (arXiv:2610.03880). The code is at https://github.com/paprakash/OMatGRPO, and the weights load with it. Its scripts/download_assets.py downloads these files, checks them against SHA256SUMS and unpacks the CIFs.

Contents

prior/prior.safetensors                      the pretrained OMatG model every run starts from (KL reference)
prior/train.yaml                             its OMatG configuration
models/<identifier>/final_model.safetensors  weights at the end of training (750 rollouts)
models/<identifier>/resolved_config.json     the run's settings, as the code's --print_config prints them
structures/<identifier>/cifs.zip             the CIFs that LeMat-GenBench scored for the paper (folder cifs/)
structures/<identifier>/structures_summary.csv  one row per generated structure (see below)
structures/<identifier>/<identifier-in-file>_comprehensive_multi_mlip_hull_<time>.json  the LeMat-GenBench result
SHA256SUMS                                   sha256 of every file

The two MP-20 references of the reward are not included. scripts/build_references.py rebuilds them from public data.

Weights

All weights are float32 state dicts of the OMatG network (73 tensors, 12,403,912 values), converted from the PyTorch checkpoints of the paper's runs. Every tensor equals its source exactly.

The prior is our own OMatG model, pretrained on MP-20 for de novo generation with stochastic position and lattice channels and masked discrete flow matching for species. It is not one of the models released by the OMatG authors. prior/train.yaml has its configuration. The hardware, wall time and seed of the pretraining were not recorded.

identifier paper name training run
arityguard_creatrelax OMatGRPO arityguard_creatrelax_750
canonical_creatrelax discovery canonical_creatrelax_750
sparseworst_creatrelax penalty routing sparseworst_creatrelax_750
arityguard guarded, pre-creativity canonical_arityguard_750
canonical discovery, pre-creativity canonical_species_rl_750
sparseworst penalty routing, pre-creativity canonical_sparseworst_750
frozen_control frozen-composition control frozen_species_control_750

Each identifier is also a run config in the code (configs/runs/<identifier>.yaml). All runs used seed 0, 4 groups of 16 structures per rollout and 750 rollouts.

Structures

These are the exact sets behind the paper's LeMat-GenBench numbers, copied from the original evaluation folders (not regenerated). Each set was generated with the paper's protocol: 2,500 structures, consistent sampler (64 time points, species noise 0), seed 42, chunks of 100, relaxation with FIRE and UMA uma-s-1p2 including the cell (force tolerance 0.05 eV/Å, at most 500 steps), cell and masked-species guards before and after the relaxation. The best-of-N set is the exception for generation: its pool of 48,000 structures was sampled from the prior with seed 4242, and the 2,500 selected structures were then relaxed and scored with the same protocol. cifs.zip holds the folder cifs/ with the relaxed structures that passed both guards and have a finite E_hull. The CIFs are zipped because a Hugging Face repository holds at most 20,000 files. LeMat-GenBench drops CIFs that its pymatgen cannot parse, so its structure count can be slightly lower than the number of files.

identifier what it is CIFs mSUN (LeMat-GenBench)
R0 the prior 2497 336
bestofn48k_top2500 best-of-N baseline: the top 2,500 of 48,000 structures sampled from the prior 2500 192
frozen_control frozen-composition control 2495 245
canonical discovery, pre-creativity 2490 1063
canonical_creatrelax discovery 2463 939
sparseworst penalty routing, pre-creativity 2494 789
sparseworst_creatrelax penalty routing 2495 841
arityguard guarded, pre-creativity 2486 982
arityguard_creatrelax OMatGRPO 2485 1138

LeMat-GenBench ran at commit 58e6eae3e4a6c87c22171cf069123ecc4e2fa7e6 with the preset comprehensive_multi_mlip_hull, and all three potentials of its stability ensemble (ORB, MACE, UMA) scored every valid structure. mSUN counts metastable, unique and novel structures and excludes structures on the hull.

structures_summary.csv columns: idx (index in the generated set of 2,500), formula, n_elements, eh_rel (E_hull of the relaxed structure against the LeMat-Bulk UMA hull, eV/atom), ref_rel (hull reference entries in its chemical system), valid, trusted, sparse (fewer than 12 reference entries), deep_below (E_hull below -0.1 eV/atom), unique, novel_mp20, novel_alexmp, msun_mp20, msun_alexmp (an internal check against MP-20 and Alex-MP-20, not LeMat-GenBench), and cif (file name in cifs/, empty for structures that were not exported).

License

MIT for the weights. The structures were computed with the UMA potential (facebook/UMA, distributed under its own license) and are derived from a model trained on MP-20, which comes from the Materials Project (data licensed CC BY 4.0; A. Jain et al., 2013, https://doi.org/10.1063/1.4812323; MP-20: T. Xie et al., 2021, https://arxiv.org/abs/2110.06197). E_hull values use the LeMat-Bulk-MLIP-Hull convex hull (LeMaterial, https://huggingface.co/datasets/LeMaterial/LeMat-Bulk-MLIP-Hull), which is not included here.

Citation

Please cite the paper and OMatG.

@misc{prakash2026reinforcementlearningdiscretecomposition,
    title={Reinforcement Learning on the Discrete Composition Channel of a Crystal Generator: Validated Gains and Reward Hacking},
    author={Pawan Prakash and Philipp Höllmer and Addis Fuhr and Peter Hirschfeld and P. Ganesh and Stefano Martiniani and Richard Hennig},
    year={2026},
    eprint={2610.03880},
    archivePrefix={arXiv},
    primaryClass={cs.LG},
    url={https://arxiv.org/abs/2610.03880},
}

@inproceedings{hoellmer2025,
    title={Open Materials Generation with Stochastic Interpolants},
    author={Philipp H{\"o}llmer and Thomas Egg and Maya Martirossyan and Eric
    Fuemmeler and Zeren Shui and Amit Gupta and Pawan Prakash and Adrian
    Roitberg and Mingjie Liu and George Karypis and Mark Transtrum and Richard
    Hennig and Ellad B. Tadmor and Stefano Martiniani},
    booktitle={Forty-second International Conference on Machine Learning},
    year={2025},
    url={https://openreview.net/forum?id=gHGrzxFujU},
    archivePrefix={arXiv},
    eprint={2502.02582},
    primaryClass={cs.LG},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Papers for paprakash/OMatGRPO