SolidAcid-0 MLIPs
Machine-learned interatomic potentials (MLIPs) for solid acid proton conductors, built on the MACE architecture.
This repository shares two families of models:
- Local-universal (base) models: trained on the full SolidAcid-0 dataset spanning several solid-acid chemistries, intended as general-purpose potentials across the solid acid compositional space.
- Domain-specific (transfer-learned) models: the base models fine-tuned on individual compounds for higher accuracy within a single chemistry.
The training/fine-tuning datasets are hosted separately on the Hugging Face Hub: https://huggingface.co/datasets/jhaens/solidacid-0_model_dataset
Repository layout
.
├── tiny/ # base local-universal model, size "tiny"
│ ├── config_train.yaml
│ └── tiny_seed{1111..5555}.model
├── small/ # base local-universal model, size "small"
│ ├── config_train.yaml
│ └── small_seed{1111..5555}.model
├── medium/ # base local-universal model, size "medium"
│ ├── config_train.yaml
│ └── medium_seed{1111..3333}.model
├── tiny_transferlearned/ # domain-specific fine-tunes of the tiny base model
│ └── csh2po4/
│ ├── config_finetune.yaml
│ └── tiny_transferlearned_csh2po4_seed{1111..5555}.model
└── small_transferlearned/ # domain-specific fine-tunes of the small base model
├── csh2po4/
├── csh2aso4/
├── cshseo4/
└── cshso4/
├── config_finetune.yaml
└── small_transferlearned_<system>_seed{1111..5555}.model
Each model is provided as a seed ensemble. Different seeds are independently initialized/trained models with identical architecture and data; use them together to estimate committee (ensemble) uncertainty, or pick a single seed for production runs.
Model sizes
All base models are MACE potentials with the same core hyperparameters (see config_train.yaml in each directory). The sizes differ in their equivariant feature content:
| Size | hidden_irreps |
max_ell |
Seeds |
|---|---|---|---|
| tiny | 64x0e |
2 | 1111–5555 |
| small | 128x0e |
3 | 1111–5555 |
| medium | 128x0e + 128x1o |
3 | 1111–3333 |
Shared architecture / training settings (base models):
model: MACE,default_dtype: float64r_max: 6.0,correlation: 3,num_interactions: 2interaction_first: RealAgnosticDensityInteractionBlock,interaction: RealAgnosticDensityResidualInteractionBlockpair_repulsion: True,distance_transform: Agnesiloss: universal,forces_weight: 10,energy_weight: 1.0(no stress)- EMA (
ema_decay: 0.999),amsgrad,weight_decay: 1e-8,clip_grad: 10.0 lr: 0.005,batch_size: 8,max_num_epochs: 25
Transfer-learned (domain-specific) models
The domain-specific models are produced by fine-tuning a single base seed (<size>_seed1111.model) on per-compound data (see each config_finetune.yaml):
- Fine-tuned for
max_num_epochs: 30,lr: 0.005(0.01 forcshso4) multiheads_finetuning: false(single-head naive/transfer fine-tuning)- Per-system atomic reference energies (
E0s) and reference keys (REF_TotEnergy,REF_Force) as recorded in the config - Released as seed ensembles (1111–5555)
Usage
These are standard MACE model files and can be used with the mace-torch package via ASE.
Installation
pip install mace-torch
# optional:
pip install cuequivariance
Single model with ASE
from ase.io import read
from mace.calculators import MACECalculator
atoms = read("your_structure.xyz")
calc = MACECalculator(
model_paths="small/small_seed1111.model",
device="cuda", # or "cpu"
default_dtype="float64",
)
atoms.calc = calc
print(atoms.get_potential_energy())
print(atoms.get_forces())
Ensemble (committee) for uncertainty estimates
Pass all seeds of a given model to average predictions and obtain a per-atom force committee spread:
from mace.calculators import MACECalculator
seeds = [1111, 2222, 3333, 4444, 5555]
calc = MACECalculator(
model_paths=[f"small/small_seed{s}.model" for s in seeds],
device="cuda",
default_dtype="float64",
)
Data
Training and fine-tuning datasets: https://huggingface.co/datasets/jhaens/solidacid-0_model_dataset
The config_train.yaml / config_finetune.yaml files reference the internal split paths (train/val/test) used at training time; the corresponding data is provided in the dataset repository above.
Citation
If you use these models, please cite the corresponding preprint and the MACE publication:
@unpublished{solidacid0,
title={ABC},
author={H{\"a}nseroth, Jonas and von Stackelberg, Rose and Dre{\ss}ler, Christian},
archivePrefix={arXiv},
year={2026},
eprint={2609.xxx}
}
@inproceedings{Batatia2022mace,
title={{MACE}: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields},
author={Ilyes Batatia and David Peter Kovacs and Gregor N. C. Simm and Christoph Ortner and Gabor Csanyi},
booktitle={Advances in Neural Information Processing Systems},
editor={Alice H. Oh and Alekh Agarwal and Danielle Belgrave and Kyunghyun Cho},
year={2022},
url={https://openreview.net/forum?id=YPpSngE-ZU}
}
License
The MLIP models are published and distributed under the MIT License.