Self-Supervised Learning and Representation Interpretation for Atomistic Systems

Background Self-supervised learning has become a major paradigm in modern representation learning, including for molecular and materials data. The central idea is that useful representations can be learned from structure, context, or predictive objectives without requiring dense task labels. In atomistic systems, this is particularly appealing because labels can be scarce, expensive, or difficult to define consistently across chemistry and materials domains. A growing body of work explores contrastive learning, BYOL-style self-distillation, masked prediction, and joint-embedding predictive objectives as ways to learn representations that preserve chemical and geometric information. Yet, a key open issue remains: what exactly do these learned representations encode, and which objective is most appropriate for preserving the right information for scientific applications? This question is not only about downstream benchmark performance, but also about whether the representation reflects chemically meaningful structure, geometric information, rare environments, and robust variation across the data manifold. ...

September 24, 2026

Geometric Analysis of Deep Learning

Background Modern deep neural networks, especially those in the overparameterized regime with a very large number of parameters, perform impressively well. Traditional learning theories contradict these empirical results and fail to explain this phenomenon, leading to new approaches that aim to understand why deep learning generalizes. A common belief is that flat minima [1] in the parameter space lead to models with good generalization properties. For instance, such models may learn to extract high-quality features from the data, known as representations. However, it has also been shown that models with equivalent performance can exist at sharp minima [2, 3]. These contradictory findings motivate us to study optimization, learned representations, and their impact on generalization from a geometric perspective. ...

November 17, 2025 · Georgios Arvanitidis

Enhancing Relative Representations using Custom Weighted Mahalanobis Distance

Background Relative representations are a powerful tool in machine learning and data analysis, where data points are represented based on their distances or similarities to a set of reference points called anchors. Traditional methods often rely on similarity measures, such as cosine similarity, which are invariant under rotation and scaling but may not satisfy the properties of a metric space, particularly the triangle inequality. This limitation can hinder the effectiveness of certain algorithms that require a proper distance metric. ...

January 15, 2025