Robustifying Asynchronous SGD via Soft Throttling
arXiv preprint arXiv:2609.39357, 2026
TL;DR. We made asynchronous SGD Byzantine-robust with a surprisingly simple trick, and it turns out the trick helps even when there are no attackers.
The figure below is the whole algorithm in one picture. Can you see what \(q\) does to the fast clients?

Abstract. Asynchronous SGD is a popular algorithm for distributed learning where each client’s gradient update is applied on arrival. This leads to a speed-up, but also an increased vulnerability to attacks, as fast clients can dominate the total update. We introduce Throttle, a Byzantine-robust generalization of asynchronous SGD where the key idea is to exponentially down-weight updates from faster clients by a factor \(q\). Both asynchronous SGD (\(q=1\)) and synchronous Byzantine-robust SGD (\(q\to\infty\)) correspond to specific settings of Throttle. We provide a theoretical analysis of the convergence rate and validate the robustness to attacks both theoretically and empirically. Remarkably, our experiments show that this down-weighting mechanism can also improve performance over standard asynchronous SGD even in the non-Byzantine setting.
Joint work with Maxime Meyer, Yuki Takezawa, Makoto Yamada, and Anastasia Koloskova.
