Researchers' deep dictionary network-based foundation model for ultra-low-dose CT denoising; outperforms across multi-organ benchmarks.

Researchers’ deep dictionary network-based foundation model for ultra-low-dose CT denoising; outperforms across multi-organ benchmarks. A deep dictionary network-based foundation model for ultra-low-dose CT denoising 1 A deep dictionary network-based foundation model for ultra-low-dose CT denoising Baoshun Shi, Shuangyi Yang, Ke Jiang, Bin Zhu, Zhanli Hu, and Huazhu Fu Abstract— Ultra-low-dose computed tomography (ULDCT) reduces radiation exposure but suffers from severe noise that degrades diagnostic image quality. Existing deep learning-based denoising methods are typically trained in an organ-specific fashion, resulting in limited generalization across heterogeneous multi- organ imaging scenarios. Foundation models present a promising all-in-one paradigm for unified multi-organ denoising. However, their architectures suffer from poor interpretability and rely on heuristic training strategies. To address these limitations, we propose an architecture- interpretable foundation model based on the deep dictionary network (DDN) for unified multi-organ ULDCT denoising. Inspired by multilayer sparse representation theory, DDN cascades convolutional sparse coding layers with iterative soft-thresholding, providing inherent architectural interpretability. Furthermore, a dynamic dictionary module and a threshold generation module are embedded within each layer to enhance representation ability. We conduct DDN pre-training on more than one million multi-organ normal-dose CT images by recovering clean images from Gaussian-noised inputs. Sparse regularization is additionally imposed on latent feature representations, guiding the network to learn compact and noise-robust priors. The complete architecture is jointly fine-tuned on multi-organ ULDCT datasets, enabling a single unified model to perform denoising across diverse anatomical regions. Extensive experiments validate that our proposed method achieves state-of-the- art performance and consistently surpasses competing ULDCT methods across all multi-organ benchmarks under the few-shot learning setting. This work was supported by the National Natural Science Foundation of China under Grant No. 62371414, by the Hebei Natural Science Foundation under Grant No. F2025203070, by the Beijing Natural Sci- ence Foundation–Haidian Original Innovation Joint Fund, Key Research Program under Grant L242067, by the General Open Fund Project of State Key Laboratory of Medical Imaging Science and Technology Systems, and by the Scientific Research Cultivation Project (Science and Engineering)-Basic Innovation Research Cultivation Project (Sci- ence and Engineering, Post-2021) under Grant No. 2025LGZD002. (Corresponding author: Baoshun Shi, e-mail: shibaoshun@ysu.edu.cn.) Baoshun Shi, Shuangyi Yang, and Ke Jiang are with the School of Information Science and Engineering, Yanshan University, Qinhuangdao 066004, Hebei, China, and also with the Hebei Key Laboratory of Information Transmission and Signal Processing, Yanshan University, Qinhuangdao 066004, Hebei, China. Bin Zhu is with the Department of Orthopedics, Beijing Friendship Hospital, Capital Medical University, Beijing, China. Zhanli Hu is with the Lauterbur Research Center for Biomedical Imag- ing, Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China. Huazhu Fu is with the Institute of Advanced Intelligence and Com- puting (IAIC), Agency for Science, Technology and Research (A*STAR), Singapore 138632. Index Terms— Ultra-low-dose CT, foundation model, deep dictionary network, model interpretability. I. INTRODUCTION C OMPUTED tomography (CT) is an indispensable imag- ing modality for clinical diagnosis and disease screening. However, the associated X-ray exposure may increase the risk of cancer and other adverse health effects [1]. Ultra-low-dose CT (ULDCT) can reduce radiation exposure by lowering the tube current or incident photon flux [2], but the reduced photon count introduces severe quantum noise into the projection data. After the filtered back-projection (FBP) operator, the reconstructed images often contain strong noise and streak- like artifacts that obscure fine anatomical structures and reduce diagnostic reliability. Effective ULDCT denoising methods should be developed to suppress severe noise while retaining clinically relevant details. In recent years, deep learning-based methods have achieved promising performance in ULDCT image denoising by learn- ing the mapping from noisy images to normal-dose CT images [3]–[5]. Despite their effectiveness, most existing methods are still trained in an organ-specific manner, where a separate model is built for each anatomical region. Such a paradigm has two main limitations: storing multiple organ-specific models increases storage and deployment costs of deep neural net- works (DNNs), while each model learns only from its own anatomical data, making it difficult to fully exploit general CT image priors shared across organs. As a result, the gen- eralization capability of existing denoising methods remains limited in heterogeneous multi-organ ULDCT scenarios. Foundation models improve the generalization of DNNs through large-scale pre-training and downstream adaptation [6]. Representative visual foundation models include self- distillation approaches [7] and masked image modeling meth- ods [8]. From the perspective of network architecture, most existing methods adopt ViT- or Transformer-based backbones [7], [8] and recent methods focus on state space model (SSM), which can effectively model long-range dependencies with linear computational complexity [9]. From the perspective of application, these advances have also extended founda- tion models to medical applications, including computational pathology and scalable medical image encoding. However, most network architectures of existing foundation models remain difficult to interpret and primarily target generic visual representation learning or medical image understanding, rather arXiv:2609.16031v1 [eess.IV] 11 Sep 2026 ...

September 16, 2026