---
title: "Zhenyu Zhu defends PhD thesis at LIONS lab, EPFL on Aug 24, 2026; Four main contributions advancing overparameterized nets"
sdDatePublished: "2026-08-25T18:08:00Z"
source: "https://actu.epfl.ch/news/phd-defense-of-zhenyu-zhu-2/"
topics:
  - name: "scientific research"
    identifier: "medtop:20000735"
  - name: "scientific innovation"
    identifier: "medtop:20000736"
  - name: "technology and engineering"
    identifier: "medtop:20000756"
  - name: "information technology and computer science"
    identifier: "medtop:20000763"
locations:
  - "Lausanne"
---


Zhenyu Zhu defends PhD thesis at LIONS lab, EPFL on Aug 24, 2026; Four main contributions advancing overparameterized nets

PhD Defense of Zhenyu Zhu - EPFL

PhD Defense of Zhenyu Zhu

Import & publish the news

Are you sure you want to import this news into ?

This news will be sent to its subscribers.

On August 24th, 2026, Zhenyu Zhu, a PhD student at LIONS lab, successfully defended his PhD thesis. The thesis, entitled "T heoretical Studies of Learning in Overparameterized Neural Networks " was supervised by Professor Volkan Cevher. Congratulations to Zhenyu!

Modern deep learning often relies on highly overparameterized neural networks, which can interpolate training data, learn useful representations, and adapt large pretrained models to downstream tasks. This thesis studies theoretical aspects of learning in such models, focusing on how parameterization, optimization dynamics, and model geometry shape generalization, feature learning, and efficient adaptation.

The thesis contains four main contributions. First, we analyze deep ReLU networks under lazy training and establish near Bayes-optimal test-error guarantees under suitable conditions. We also discuss how these guarantees relate to interpolation results available for closely related optimization settings. Second, we study score matching with deep ReLU networks and provide finite-sample guarantees for score-function estimation, with applications to causal discovery and score-based generative modeling. Third, we analyze feature learning dynamics in teacher--student ReLU networks and show how gradient descent aligns and balances learned features. Finally, we study Low-Rank Adaptation (LoRA) for foundation model fine-tuning and propose parameterization-aware algorithms, iLoRA and LoRA, to improve the flatness, stability, and generalization of low-rank fine-tuning.

Overall, this thesis shows that the behavior of overparameterized neural networks is shaped not only by their expressive power, but also by the geometry of their parameterization and the dynamics of optimization.

Source: Laboratory for Information and Inference Systems

All Laboratory for Information and Inference Systems news

All School of Engineering | STI news

ABSTRACT Modern deep learning often relies on highly overparameterized neural networks, which can interpolate training data, learn useful representations, and adapt large pretrained models to downstream tasks. This thesis studies theoretical aspects of learning in such models, focusing on how parameterization, optimization dynamics, and model geometry shape generalization, feature learning, and efficient adaptation. The thesis contains four main contributions. First, we analyze deep ReLU networks under lazy training and establish near Bayes-optimal test-error guarantees under suitable conditions. We also discuss how these guarantees relate to interpolation results available for closely related optimization settings. Second, we study score matching with deep ReLU networks and provide finite-sample guarantees for score-function estimation, with applications to causal discovery and score-based generative modeling. Third, we analyze feature learning dynamics in teacher--student ReLU networks and show how gradient descent aligns and balances learned features. Finally, we study Low-Rank Adaptation (LoRA) for foundation model fine-tuning and propose parameterization-aware algorithms, iLoRA and LoRA, to improve the flatness, stability, and generalization of low-rank fine-tuning. Overall, this thesis shows that the behavior of overparameterized neural networks is shaped not only by their expressive power, but also by the geometry of their parameterization and the dynamics of optimization.