首页 > AI前沿 > Similarity Pairing with Energy Mover's Distance for Self-Supervised Pre-Training at the LHC

Similarity Pairing with Energy Mover's Distance for Self-Supervised Pre-Training at the LHC

arXiv机器学习 2026-09-16 02:44 5 阅读 查看原文

Many self-supervised methods for training foundation models at the Large Hadron Collider (LHC) rely on data augmentations to encourage the model to embed events into a representation space invariant to certain physical or detector symmetries.

A common challenge arises from the large freedom in choosing a proper set of augmentations on which downstream performance depends.

The implementation of augmentations involves either modifying existing events, potentially breaking the event fidelity, or simulating more event variants, which is computationally intensive.

In this work

We present a data-driven method of pairing events by their similarity via the energy mover's distance (EMD), which measures how similar two events are in terms of the work required to transform one into the other.

With this approach, distinct events are sampled and matched by their similarity to serve as views for learning invariance, keeping the physics content of each event intact without handcrafted distortions.

We demonstrate this augmentation-free pairing method by pre-training on QCD jets via self-distillation and show that it can yield semantic jet embeddings with downstream discrimination power comparable to or better than an augmentation-based baseline.