Double Descent, Overparametrization and Scaling Laws in Particle Physics Data
Technical University of Munich
Hosted by PHYSTAT
Abstract
Matthias Vigl examines whether the computational scaling strategies behind modern machine learning can benefit particle physics. Empirical scaling laws relate model performance to computing resources and help balance model size against training-data volume. For a fixed dataset, increasing model capacity beyond the interpolation threshold can improve generalization through double descent, but gains eventually encounter limits imposed by the data. The talk then considers scaling data and model size together, deriving compute-optimal relations for transformer-based jet taggers trained on as many as billions of simulated jets. It studies how training hyperparameters and input representations change these relations. The observed behaviour suggests that larger computing budgets could yield useful gains in particle physics, especially with a move toward more general-purpose foundation models.
Topics
Related Seminars
Understanding machine learning via exactly solvable statistical physics models
Related research
Learning and prediction in artificial deep neural networks: scaling, data manifolds, and universality
More on scaling laws
Can machine learning learn new physics, or do we need to put it in by hand?"\
Related research