Generative Models for Molecules and Materials: Guidance, Acceleration, and Sampling Dynamics

Background Generative modeling has become one of the most active areas of machine learning for scientific discovery, particularly in chemistry and materials science. Recent work has shown that diffusion models, flow-based generators, and related stochastic generative frameworks can produce chemically plausible molecules and crystalline structures, often while respecting geometric and physical constraints. The field is moving quickly, but many open questions remain about what makes these models reliable, controllable, and computationally efficient for atomistic systems. ...

September 24, 2026

Self-Supervised Learning and Representation Interpretation for Atomistic Systems

Background Self-supervised learning has become a major paradigm in modern representation learning, including for molecular and materials data. The central idea is that useful representations can be learned from structure, context, or predictive objectives without requiring dense task labels. In atomistic systems, this is particularly appealing because labels can be scarce, expensive, or difficult to define consistently across chemistry and materials domains. A growing body of work explores contrastive learning, BYOL-style self-distillation, masked prediction, and joint-embedding predictive objectives as ways to learn representations that preserve chemical and geometric information. Yet, a key open issue remains: what exactly do these learned representations encode, and which objective is most appropriate for preserving the right information for scientific applications? This question is not only about downstream benchmark performance, but also about whether the representation reflects chemically meaningful structure, geometric information, rare environments, and robust variation across the data manifold. ...

September 24, 2026

Uncertainty Estimation in Atomistic Machine Learning: Calibration, Robustness, and Detection of Distribution Shift

Background Uncertainty estimation is becoming increasingly important in scientific machine learning, especially when models are used to support decision-making in chemistry, materials, and molecular design. In atomistic systems, where data may be noisy, sparse, or shifted relative to the training distribution, a model that produces a single point estimate can be misleading. In such settings, understanding whether a prediction is reliable is often as important as the prediction itself. A variety of uncertainty methods have been developed in deep learning, including ensembles, Bayesian approximations, calibration techniques, and evidential approaches. However, these methods do not always transfer cleanly to atomistic models, where structure, symmetry, local geometry, and domain shifts can interact in complex ways. The challenge is not only to estimate uncertainty but also to determine what kind of uncertainty is useful: epistemic uncertainty, aleatoric uncertainty, calibration under distribution shift, or the ability to detect out-of-distribution chemistry. ...

September 24, 2026