SimPoly: Simulation of Polymers with Machine Learning Force Fields Derived from First Principles
Authors:
Gregor N. C. Simm,
Jean Hélie,
Hannes Schulz,
Yicheng Chen,
Guillem Simeon,
Anna Kuzina,
Ernesto Martinez-Baez,
Piero Gasparotto,
Gabriele Tocci,
Chi Chen,
Yatao Li,
Lixue Cheng,
Zun Wang,
Bichlien H. Nguyen,
Jake A. Smith,
Lixin Sun
Abstract:
Polymers are a versatile class of materials with widespread industrial applications. Advanced computational tools could revolutionize their design, but their complex, multi-scale nature poses significant modeling challenges. Conventional force fields often lack the accuracy and transferability required to capture the intricate interactions governing polymer behavior. Conversely, quantum-chemical m…
▽ More
Polymers are a versatile class of materials with widespread industrial applications. Advanced computational tools could revolutionize their design, but their complex, multi-scale nature poses significant modeling challenges. Conventional force fields often lack the accuracy and transferability required to capture the intricate interactions governing polymer behavior. Conversely, quantum-chemical methods are computationally prohibitive for the large systems and long timescales required to simulate relevant polymer phenomena. Here, we overcome these limitations with a machine learning force field (MLFF) approach. We demonstrate that macroscopic properties for a broad range of polymers can be predicted ab initio, without fitting to experimental data. Specifically, we develop a fast and scalable MLFF to accurately predict polymer densities, outperforming established classical force fields. Our MLFF also captures second-order phase transitions, enabling the prediction of glass transition temperatures. To accelerate progress in this domain, we introduce a benchmark of experimental bulk properties for 130 polymers and an accompanying quantum-chemical dataset. This work lays the foundation for a fully in silico design pipeline for next-generation polymeric materials.
△ Less
Submitted 15 October, 2025;
originally announced October 2025.
Understanding multi-fidelity training of machine-learned force-fields
Authors:
John L. A. Gardner,
Hannes Schulz,
Jean Helie,
Lixin Sun,
Gregor N. C. Simm
Abstract:
This study systematically investigates two multi-fidelity strategies used to train machine-learned force fields (MLFFs) -- pre-training/fine-tuning and multi-headed training -- and elucidates the mechanisms underpinning their success. For pre-training and fine-tuning, we uncover a log-log linear relationship between pre-trained and fine-tuned accuracies that holds across model architectures, model…
▽ More
This study systematically investigates two multi-fidelity strategies used to train machine-learned force fields (MLFFs) -- pre-training/fine-tuning and multi-headed training -- and elucidates the mechanisms underpinning their success. For pre-training and fine-tuning, we uncover a log-log linear relationship between pre-trained and fine-tuned accuracies that holds across model architectures, model sizes, and quantum-chemical methods. The success of this approach hinges on the quantity and quality of available pre-training data, and, critically, the inclusion of force labels. We demonstrate that pre-trained representations are inherently method-specific, requiring adaptation of the model backbone during fine-tuning. In contrast, multi-headed models learn method-independent backbone representations, where again the heads' accuracies are log-log linearly related. Relative to pre-training and fine-tuning, these shared representations marginally reduce model performance in most cases. However, this trade-off is offset by practical advantages: multi-headed training extends naturally to multiple labelling methods and enables partial replacement of expensive labels with cheaper alternatives, paving the way towards cost-efficient universal MLFFs.
△ Less
Submitted 2 April, 2026; v1 submitted 17 June, 2025;
originally announced June 2025.