Lin, Dahiya, Cersonsky propose generalized approach to incorporating geometry and directionality into coarse-grained machine-learned potentials University of Wisconsin–Madison; enables anisotropic coarse-grained simulations

Lin, Dahiya, Cersonsky propose generalized approach to incorporating geometry and directionality into coarse-grained machine-learned potentials University of Wisconsin–Madison; enables anisotropic coarse-grained simulations A Generalized Approach for Incorporating Geometry and Directionality into Coarse-Grained Machine-Learned Potentials A Generalized Approach for Incorporating Geometry and Directionality into Coarse-Grained Machine-Learned Potentials Arthur Y. Lin,1 Tejas Dahiya,2 and Rose K. Cersonsky1, 3, 4, a) 1)Department of Chemical and Biological Engineering, University of Wisconsin - Madison, Madison, WI, USA 2)Department of Computer Science, University of Wisconsin - Madison, Madison, WI, USA 3)Department of Materials Science and Engineering, University of Wisconsin - Madison, Madison, WI, USA 4)Data Science Institute, University of Wisconsin - Madison, Madison, WI, USA I. INTRODUCTION Machine-learned interatomic potentials (MLIPs) have dramatically expanded the scope of molecular simulation, enabling the prediction of atomic energies and forces with near ab initio accuracy at costs compatible with large- scale molecular dynamics.1–3 As these models continue to mature, the principal obstacle to computational discovery is increasingly shifting from the accuracy of atomistic interactions to the challenge of accessing the larger length and time scales associated with collective behavior, self- assembly, and phase transitions. Coarse-graining (CG), in which groups of atoms are mapped onto effective particles with reduced degrees of freedom, remains one of the most powerful approaches for bridging this gap.4–8 For any MLIP, its capabilities are determined jointly by the quality of the underlying data, the choice of repre- sentation X, and the architecture f(X), where X varies from simple forms (e.g., distance matrices or graphs) to more complex descriptors (e.g., Smooth Overlap of Atomic Positions [SOAP]9 or Behler-Parinello symme- try functions10). A representation is complete if distinct data objects map to distinct data points; in the atomistic MLIP community there has been considerable discussion on what constitutes a complete descriptor, as this has direct implications to the ceiling of predictive accuracy in subsequent models.11–14. Simply put, while expressive architectures can approximate complex functions f(X), they cannot recover information that is not encoded in X, and degeneracy in data representation can explicitly limit model performance. However, unlike atomistic machine-learned potentials, the variables describing a coarse-grained system are not fixed, but instead chosen. Historically, much of coarse-graining has therefore been built around isotropic particles.4 This choice is attractive for both concep- tual and computational reasons. Interactions become functions of intermolecular separation alone, simulation methodologies are well-established, and many successful coarse-grained models have been developed within this framework.15 Such approaches have proven remarkably a)Electronic mail: rose.cersonsky@wisc.edu successful, particularly in biomolecular systems where coarse-graining has enabled simulations inaccessible at atomistic resolution.15 Yet there are also many examples where isotropic descriptions struggle to reproduce experi- mentally observed behavior, particularly in systems whose organization is governed by packing, local structure, or directional interactions.16,17 One possible explanation of these difficulties is that geometry itself carries critical information. Consider two molecular configurations possessing similar intermolecular separations but different relative orientations. Depending on the system, these configurations may exhibit substan- tially different interaction energies. When orientation is removed from the coarse-grained representation, how- ever, both are represented by the same set of variables R. The resultant coarse-grained configuration is there- fore associated not with a single underlying interaction energy, but with a distribution of possible interaction en- ergies. As this distribution broadens, the construction of a coarse-grained potential becomes increasingly difficult, as distinct molecular environments become indistinguishable within the representation. This observation naturally suggests retaining orienta- tional and geometric information within coarse-grained models. However, doing so raises a separate question: how should such information be represented? While atomistic machine-learned potentials have undergone rapid develop- ment over the past decade, the overwhelming majority of descriptor frameworks and architectures were developed for collections of isotropic atoms. Extending these ap- proaches to anisotropic particles requires representations that remain sensitive to particle geometry while preserv- ing the rotational and translational symmetries of the underlying problem. Several recent approaches have begun to address this challenge through anisotropic extensions to machine- learning architectures.18? ,19 Of particular interest are many-body density expansion methods, which have proven highly successful in atomistic machine learning due to their systematic improvability and clear connection to local structure.9,20,21 AniSOAP extended the SOAP formal- ism to non-spherical particles through anisotropic density expansions, enabling the explicit incorporation of particle shape and orientation into a many-body descriptor,22,23, and similar efforts expanded the Chebyshev polynomial arXiv:2609.01911v1 [physics.chem-ph] 1 Sep 2026 ...

September 3, 2026