---
title: "Yongtao Wu defended his PhD thesis at EPFL; adaptive steepest descent with optimal convergence rates"
sdDatePublished: "2026-08-25T18:08:00Z"
source: "https://actu.epfl.ch/news/phd-defense-of-yongtao-wu/"
topics:
  - name: "science and technology"
    identifier: "medtop:13000000"
locations:
  - "Lausanne"
---


Yongtao Wu defended his PhD thesis at EPFL; adaptive steepest descent with optimal convergence rates

PhD Defense of Yongtao Wu - EPFL

PhD Defense of Yongtao Wu

Import & publish the news

Are you sure you want to import this news into ?

This news will be sent to its subscribers.

On August 24th, 2026, Yongtao Wu, a PhD student at LIONS lab, successfully defended his PhD thesis. The thesis, entitled " Optimization in Modern Machine Learning: Steepest Descent Theory and Trustworthy Models " was supervised by Professor Volkan Cevher. Congratulations to Yongtao!

Optimization algorithms have a long history and play a key role in many aspects of modern machine learning. This thesis studies both the theoretical foundations of optimization algorithms in the era of large language models (LLMs) and their applications to trustworthy machine learning systems. Concretely, first, we study the universal convergence property of steepest descent. We design an adaptive steepest descent method that automatically adapts to both the smoothness constant of the objective and the level of stochastic noise, and attains optimal convergence rates in both deterministic and stochastic settings. Second, we investigate the convergence properties of normalized steepest descent methods from a mean-field perspective and show their difference from a class of unnormalized steepest descent methods, including standard gradient descent. Third, we study the role of optimization in trustworthy LLM alignment. We formulate the multi-step alignment process as a minâ max optimization problem corresponding to a two-player Markov game, and propose algorithms based on optimistic online mirror descent to solve it with theoretical guarantees. Fourth, we investigate discrete optimization problems that arise in trustworthy LLMs. In particular, we design character-level adversarial attacks for LLMs based on a query-based method and a simplex projection gradient method, demonstrating the vulnerability of LLMs to adversarial input with character-level perturbation. Finally, we include the benchmark services developed for the area of trustworthy machine learning. Overall, this thesis advances both the theoretical understanding of steepest descent methods and the application of optimization techniques to the development of reliable and trustworthy machine learning systems.

Source: Laboratory for Information and Inference Systems

All Laboratory for Information and Inference Systems news

All School of Engineering | STI news

ABSTRACT Optimization algorithms have a long history and play a key role in many aspects of modern machine learning. This thesis studies both the theoretical foundations of optimization algorithms in the era of large language models (LLMs) and their applications to trustworthy machine learning systems. Concretely, first, we study the universal convergence property of steepest descent. We design an adaptive steepest descent method that automatically adapts to both the smoothness constant of the objective and the level of stochastic noise, and attains optimal convergence rates in both deterministic and stochastic settings. Second, we investigate the convergence properties of normalized steepest descent methods from a mean-field perspective and show their difference from a class of unnormalized steepest descent methods, including standard gradient descent. Third, we study the role of optimization in trustworthy LLM alignment. We formulate the multi-step alignment process as a minâ max optimization problem corresponding to a two-player Markov game, and propose algorithms based on optimistic online mirror descent to solve it with theoretical guarantees. Fourth, we investigate discrete optimization problems that arise in trustworthy LLMs. In particular, we design character-level adversarial attacks for LLMs based on a query-based method and a simplex projection gradient method, demonstrating the vulnerability of LLMs to adversarial input with character-level perturbation. Finally, we include the benchmark services developed for the area of trustworthy machine learning. Overall, this thesis advances both the theoretical understanding of steepest descent methods and the application of optimization techniques to the development of reliable and trustworthy machine learning systems.