Welcome to Francis Academic Press

Academic Journal of Computing & Information Science, 2026, 9(6); doi: 10.25236/AJCIS.2026.090611.

Analyzing the Decoder Bottleneck in SAM-Style 3D Segmentation Distillation

Author(s)

Haojun Pei

Corresponding Author:
Haojun Pei
Affiliation(s)

School of Information Science and Engineering, Dalian Polytechnic University, Dalian, China

Abstract

SegVol builds on the Segment Anything Model (SAM) for 3D medical image segmentation. It handles many anatomical targets well, but at 180.9M parameters, it is too heavy for routine clinical hardware. We study how encoder weight initialization and decoder configuration affect the outcome of knowledge distillation for this model. We test 7 configurations. The results point to two decisive factors—the decoder needs to be trainable, and encoder weights should be copied from alternating teacher layers rather than consecutive ones. Our best student uses an independent decoder fine-tuned with pseudo-labels, plus skip-layer weight transfer from teacher layers {0,2,4,6,8,10}. It keeps 95.1% of the teacher's Dice (0.557 vs. 0.586), cuts 23.5% of parameters (138.4M vs. 180.9M), and runs 43.1% faster. Centered Kernel Alignment (CKA) shows why a frozen shared decoder fails: the deepest student layer can only match teacher layer 10, whereas the independent decoder's last layer reaches layer 11. Decoder trainability is the critical lever in volumetric segmentation distillation.

Keywords

Knowledge distillation, SAM, 3D medical image segmentation, Model compression, Vision transformer

Cite This Paper

Haojun Pei. Analyzing the Decoder Bottleneck in SAM-Style 3D Segmentation Distillation. Academic Journal of Computing & Information Science (2026), Vol. 9, Issue 6: 94-101. https://doi.org/10.25236/AJCIS.2026.090611.

References

[1] A. Kirillov et al., "Segment Anything," in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 3992–4003.

[2] Y. Du, F. Bai, T. Huang, and B. Zhao, "SegVol: Universal and Interactive Volumetric Medical Image Segmentation," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2024.

[3] G. Hinton, O. Vinyals, and J. Dean, "Distilling the Knowledge in a Neural Network," in NIPS Deep Learning Workshop, 2014, arXiv:1503.02531.

[4] V. Sanh, L. Debut, J. Chaumond, and T. Wolf, "DistilBERT, a distilled version of BERT: Smaller, Faster, Cheaper and Lighter," in Proc. EMC2 Workshop (NeurIPS), 2019.

[5] C. Zhang et al., "Faster Segment Anything: Towards Lightweight SAM for Mobile Applications," arXiv:2306.14289, 2023.

[6] H. Shu et al., "TinySAM: Pushing the Envelope for Efficient Segment Anything Model," arXiv:2312.13789, 2023.

[7] H. Touvron et al., "Training Data-Efficient Image Transformers & Distillation Through Attention," in Proc. Int. Conf. Mach. Learn. (ICML), PMLR 139, 2021, pp. 10347–10357.

[8] A. Romero et al., "FitNets: Hints for Thin Deep Nets," in Proc. Int. Conf. Learn. Represent. (ICLR), 2015.

[9] S. Zagoruyko and N. Komodakis, "Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer," in Proc. Int. Conf. Learn. Represent. (ICLR), 2017.

[10] Y. Tian, D. Krishnan, and P. Isola, "Contrastive Representation Distillation," in Proc. Int. Conf. Learn. Represent. (ICLR), 2020.

[11] O. Ronneberger, P. Fischer, and T. Brox, "U-Net: Convolutional Networks for Biomedical Image Segmentation," in Proc. Int. Conf. Med. Image Comput. Comput.-Assist. Intervent. (MICCAI), LNCS 9351, 2015, pp. 234–241.

[12] F. Isensee et al., "nnU-Net: a Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation," Nature Methods, vol. 18, no. 2, pp. 203–211, 2021.

[13] J. Ma et al., "Segment Anything in Medical Images," Nature Communications, vol. 15, no. 1, Art. no. 654, 2024.

[14] H. Wang et al., "SAM-Med3D: Towards General-Purpose Segmentation Models for Volumetric Medical Images," in Proc. Eur. Conf. Comput. Vis. (ECCV) Workshops, 2024.

[15] S. Kornblith, M. Norouzi, H. Lee, and G. Hinton, "Similarity of Neural Network Representations Revisited," in Proc. Int. Conf. Mach. Learn. (ICML), 2019, pp. 3519–3529.