CAIDAS-Beitrag erhält Auszeichnung als herausragender Beitrag auf der RLC 2026
18.08.2026Wir freuen uns, bekannt geben zu dürfen, dass die Arbeit „Gradient Iterated Temporal-Difference Learning“ von Théo Vincent, Kevin Gerhardt, Yogesh Tripathi, Habib Maraqten, Adam White, Martha White, Jan Peters und Carlo D'Eramo auf der Reinforcement Learning Conference (RLC) 2026 in Montreal, Kanada, mit dem „Outstanding Paper Award on Empirical Reinforcement Learning Research“ ausgezeichnet wurde.
The Reinforcement Learning Conference is one of the leading peer-reviewed venues dedicated specifically to reinforcement learning research. Its Outstanding Paper Awards recognize a small number of papers each year across several categories, with the Empirical Reinforcement Learning Research award honoring work that makes significant contributions to the empirical practice of the field, including new methodologies, benchmarks, and evaluation techniques carried out with a high standard of experimental rigor.
The awarded paper addresses a long-standing tension in temporal-difference (TD) learning between stability and speed. Most TD methods use semi-gradient updates that learn quickly but can diverge, as illustrated by the classical Baird's counterexample. Gradient TD methods fix this stability issue but have historically been slower and therefore less widely adopted. Building on the recently introduced idea of iterated TD learning, which learns a sequence of action-value functions in parallel to speed up training, the authors propose Gradient Iterated TD learning: a method that computes gradients over the moving targets in this sequence, combining the stability of gradient TD methods with competitive learning speed. Evaluations across a range of benchmarks, including Atari games, show that the approach matches the speed of semi-gradient methods while retaining the theoretical guarantees of gradient TD — a result no prior gradient TD method has demonstrated.
Carlo D'Eramo leitet die Forschungsgruppe „Reinforcement Learning and Computational Decision-Making“ am CAIDAS der Universität Würzburg.
