Parallel Computing seminars
May 2026
Asynchronous Methods on AMD GPU-Based Systems
Katarzyna Swirydowicz· Advanced Micro Devices (AMD)
Fri, May 8 · 13:00 UTC · Providence, United States · In person
Katarzyna Swirydowicz uses an asynchronous solver on an AMD system to examine the practical implementation of computational linear algebra across CPUs and GPUs. The talk introduces the relevant computational ideas, programming models, and software tools, then considers how algorithmic structure interacts with hardware capabilities. This case study illustrates both the opportunities and implementation challenges of asynchronous methods for large-scale scientific computing on GPU-accelerated systems.
Linear AlgebraComputer ScienceSeries: Institute for Computational and Experimental Research in Mathematics (ICERM), Brown UniversityVideo+2 more
Multigrid methods on high performance computers
Matthias Bolten· Bergische Universität Wuppertal
Wed, May 6 · 14:30 UTC · Providence, United States · In person
Matthias Bolten discusses the scalability of multigrid solvers for linear systems arising from discretized partial differential equations. On modern supercomputers, heterogeneous CPUs and GPUs and the widening gap between computation, network, and memory speeds complicate parallelization. Classical multigrid analysis relies on tightly coupled multiplicative components, whereas additive and asynchronous variants relax this coupling. The talk compares approaches to improving high-performance multigrid scalability, including asynchronous execution.
Linear AlgebraComputational MathematicsSeries: Institute for Computational and Experimental Research in Mathematics (ICERM), Brown UniversityVideo+3 more
Straggler-Tolerant Iterative Methods for Linear Systems and Eigenvector Computations with Partial Matrix-Vector Products
Vasileios Kalantzis· IBM Research
Tue, May 5 · 20:00 UTC · Providence, United States · In person
Vasileios Kalantzis develops iterative linear algebra algorithms that tolerate incomplete matrix-vector products in controller-worker cloud systems. Richardson and Chebyshev schemes solve linear systems using randomly available product entries, replacing missing entries with zero. For dominant eigenvectors, modified power iterations substitute zeros, previous entries, or averages of partial iterates for delayed components. The talk presents convergence results in expectation and numerical experiments on sparse matrices for both problem classes.
Linear AlgebraComputational MathematicsSeries: Institute for Computational and Experimental Research in Mathematics (ICERM), Brown UniversityVideo+3 more
Acceleration and Adaptive Selection in Asynchronous Iterative Solvers
Evan Coleman· University of Mary Washington
Tue, May 5 · 18:30 UTC · Providence, United States · In person
Evan Coleman studies how asynchronous solvers can recover convergence quality while tolerating stale data, stragglers, and variable delays. At the coordinator, Anderson acceleration connects asynchronous stationary iterations to Krylov methods with changing preconditioners and flexible GMRES. Controlled-delay experiments on high-performance computing infrastructure show that its effectiveness depends on the iteration's coupling density. At the worker, residual-weighted randomized coordinate descent includes Boltzmann weights that interpolate between uniform and greedy selection while preserving convergence guarantees. Both approaches seek better use of computation when information is inconsistent.
Linear AlgebraComputational MathematicsSeries: Institute for Computational and Experimental Research in Mathematics (ICERM), Brown UniversityVideo+3 more
Asynchronous preconditioners and linear solvers
Erik Boman· Sandia National Laboratories
Tue, May 5 · 14:30 UTC · Providence, United States · In person
Erik Boman discusses preconditioning for asynchronous linear solvers. Inner products create synchronization requirements in Krylov methods, while preconditioners can also improve iterations such as Richardson's method. The talk focuses on asynchronous incomplete factorizations and introduces ATS-ILU, an iterative incomplete LU method with synchronous and asynchronous versions that performs competitively with ParILU.
Linear AlgebraApplied MathematicsSeries: Institute for Computational and Experimental Research in Mathematics (ICERM), Brown UniversityVideo+3 more
November 2021
Efficient GPU training of SNNs using approximate RTRL
James Knight· University of Sussex
Wed, Nov 3 · 17:15 UTC
Last year’s SNUFA workshop report concluded “Moving toward neuron numbers comparable with biology and applying these networks to real-world data-sets will require the development of novel algorithms, software libraries, and dedicated hardware accelerators that perform well with the specifics of spiking neural networks” [1]. Taking inspiration from machine learning libraries — where techniques such as parallel batch training minimise latency and maximise GPU occupancy — as well as our previous research on efficiently simulating SNNs on GPUs for computational neuroscience [2,3], we are extending our GeNN SNN simulator to pursue this vision. To explore GeNN’s potential, we use the eProp learning rule [4] — which approximates RTRL — to train SNN classifiers on the Spiking Heidelberg Digits and the Spiking Sequential MNIST datasets. We find that the performance of these classifiers is comparable to those trained using BPTT [5] and verify that the theoretical advantages of neuron models with adaptation dynamics [5] translate to improved classification performance. We then measured execution times and found that training an SNN classifier using GeNN and eProp becomes faster than SpyTorch and BPTT after less than 685 timesteps and much larger models can be trained on the same GPU when using GeNN. Furthermore, we demonstrate that our implementation of parallel batch training improves training performance by over 4⨉ and enables near-perfect scaling across multiple GPUs. Finally, we show that performing inference using a recurrent SNN using GeNN uses less energy and has lower latency than a comparable LSTM simulated with TensorFlow [6].
End of results.