David Persson presents Polar Express, a method for polar decomposition and the matrix sign function motivated by the Muon neural-network optimizer. It uses matrix multiplications suited to GPUs and adapts each polynomial update through a minimax problem to reduce worst-case error. The talk covers convergence, finite-precision implementation in bfloat16 and validation-loss improvements when training GPT-2 on FineWeb data. This is an in-person PACM IDeAS seminar at Princeton.
We use essential cookies to run the site. Optional analytics and public-page session replay help us improve World Wide. Learn more.