Self-Supervised Learning and Representation Interpretation for Atomistic Systems

Background Self-supervised learning has become a major paradigm in modern representation learning, including for molecular and materials data. The central idea is that useful representations can be learned from structure, context, or predictive objectives without requiring dense task labels. In atomistic systems, this is particularly appealing because labels can be scarce, expensive, or difficult to define consistently across chemistry and materials domains. A growing body of work explores contrastive learning, BYOL-style self-distillation, masked prediction, and joint-embedding predictive objectives as ways to learn representations that preserve chemical and geometric information. Yet, a key open issue remains: what exactly do these learned representations encode, and which objective is most appropriate for preserving the right information for scientific applications? This question is not only about downstream benchmark performance, but also about whether the representation reflects chemically meaningful structure, geometric information, rare environments, and robust variation across the data manifold. ...

September 24, 2026