Welcome to Francis Academic Press

Academic Journal of Computing & Information Science, 2026, 9(8); doi: 10.25236/AJCIS.2026.090801.

DASR-Net: A Deformable Sampling and Attentive State-Space Restoration Network for In-Loop Filtering in Versatile Video Coding

Author(s)

Yue Wen1, Long Xu2, Xuande Zhang1

Corresponding Author:
Xuande Zhang
Affiliation(s)

1School of Electronic Information and Artificial Intelligence, Shaanxi University of Science and Technology, Xi’an, 710021, China

2Faculty of Information Science and Engineering, Ningbo University, Ningbo, 315211, China

Abstract

The rapid development of ultra-high-definition video and streaming media applications has imposed increasingly stringent requirements on reconstruction quality and coding efficiency in H.266/Versatile Video Coding (VVC). Although existing CNN- and Transformer-based learned in-loop filtering methods can improve reconstruction quality, they still face challenges in jointly modeling spatially irregular local artifacts, long-range structural dependencies, and cross-component luma–chroma correlations. To address these issues, this paper proposes a Deformable Sampling and Attentive State-Space Restoration Network (DASR-Net) for learned in-loop filtering in VVC. The proposed network employs a Deformable Sampling Module to adaptively model and suppress spatially irregular local artifacts, including blocking artifacts along block boundaries and ringing artifacts around texture edges. An Attentive State-Space Module is introduced to model long-range structural dependencies and non-local texture correlations, while a Luma-Guided Chroma Fusion module exploits structural cues from the luma component to guide chroma restoration. Experimental results show that, under the All-Intra configuration, DASR-Net achieves average BD-rate savings of 8.96%, 17.64%, and 19.86% for the Y, U, and V components, respectively, relative to the VTM-23.13 anchor, demonstrating the effectiveness of the proposed method.

Keywords

Versatile Video Coding, Learned In-Loop Filtering, Deformable Sampling, Attentive State-Space Modeling, Luma-Guided Chroma Restoration

Cite This Paper

Yue Wen, Long Xu, Xuande Zhang. DASR-Net: A Deformable Sampling and Attentive State-Space Restoration Network for In-Loop Filtering in Versatile Video Coding. Academic Journal of Computing & Information Science (2026), Vol. 9, Issue 8: 1-12. https://doi.org/10.25236/AJCIS.2026.090801.

References

[1] Bross B, Wang Y K, Ye Y, et al. Overview of the Versatile Video Coding (VVC) Standard and its Applications[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 31(10): 3736–3764.

[2] Sullivan G J, Ohm J R, Han W J, et al. Overview of the High Efficiency Video Coding (HEVC) Standard[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2012, 22(12): 1649–1668.

[3] Karczewicz M, Hu N, Taquet J, et al. VVC In-Loop Filters[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 31(10): 3907–3925.

[4] Dong C, Deng Y, Loy C C, et al. Compression Artifacts Reduction by a Deep Convolutional Network[C]. 2015 IEEE International Conference on Computer Vision (ICCV), 2015: 576–584.

[5] Jia C, Wang S, Zhang X, et al. Content-Aware Convolutional Neural Network for In-Loop Filtering in High Efficiency Video Coding[J]. IEEE Transactions on Image Processing, 2019, 28(7): 3343–3356.

[6] Song X, Yao J, Zhou L, et al. A Practical Convolutional Neural Network as Loop Filter for Intra Frame[C]. 2018 25th IEEE International Conference on Image Processing (ICIP), 2018: 1133–1137.

[7] Ding D, Kong L, Chen G, et al. A Switchable Deep Learning Approach for In-Loop Filtering in Video Coding[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2020, 30(7): 1871–1887.

[8] Chen S, Chen Z, Wang Y, et al. In-Loop Filter with Dense Residual Convolutional Neural Network for VVC[C]. 2020 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), 2020: 149–152.

[9] Wang M Z, Wan S, Gong H, et al. Attention-Based Dual-Scale CNN In-Loop Filter for Versatile Video Coding[J]. IEEE Access, 2019, 7: 145214–145226.

[10] Huang Z, Guo X, Shang M, et al. An Efficient QP Variable Convolutional Neural Network Based In-loop Filter for Intra Coding[C]. 2021 Data Compression Conference (DCC), 2021: 33–42.

[11] Huang Z, Sun J, Guo X, et al. One-for-All: An Efficient Variable Convolution Neural Network for In-Loop Filter of VVC[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(4): 2342–2355.

[12] Li Y, Zhang L, Zhang K. Convolutional Neural Network Based In-Loop Filter For VVC Intra Coding[C]. 2021 IEEE International Conference on Image Processing (ICIP), 2021: 2104–2108.

[13] Ma D, Zhang F, Bull D R. MFRNet: A New CNN Architecture for Post-Processing and In-loop Filtering[J]. IEEE Journal of Selected Topics in Signal Processing, 2021, 15(2): 378–387.

[14] Zhu L, Zhang Y, Li N, et al. Neural Network Based Multi-Level In-Loop Filtering for Versatile Video Coding[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(11): 12092–12096.

[15] Man H, Wang H, Lu R, et al. Content-Aware Dynamic In-Loop Filter With Adjustable Complexity for VVC Intra Coding[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2025, 35(6): 6114–6128.

[16] Kathariya B, Li Z, Wang H, et al. Multi-Stage Spatial and Frequency Feature Fusion using Transformer in CNN-Based In-Loop Filter for VVC[C]. 2022 Picture Coding Symposium (PCS), 2022: 373–377.

[17] Kathariya B, Li Z, Geert V der A. Joint Pixel and Frequency Feature Learning and Fusion via Channel-Wise Transformer for High-Efficiency Learned In-Loop Filter in VVC[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(5): 4070–4083.

[18] Zhang H, Liu Y, Jung C, et al. RTNN: A Neural Network-Based In-Loop Filter in VVC Using Resblock and Transformer[J]. IEEE Access, 2024, 12: 104599–104610.

[19] Liang J, Cao J, Sun G, et al. SwinIR: Image Restoration Using Swin Transformer[C]. 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2021: 1833–1844.

[20] Zamir S W, Arora A, Khan S, et al. Restormer: Efficient Transformer for High-Resolution Image Restoration[C]. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022: 5718–5729.

[21] Gu A, Dao T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces[C]. First Conference on Language Modeling (COLM), 2023.

[22] Guo H, Li J, Dai T, et al. MambaIR: A Simple Baseline for Image Restoration with State-Space Model[C]. Computer Vision – ECCV 2024, 2025: 222–241.

[23] Guo H, Guo Y, Zha Y, et al. MambaIRv2: Attentive State Space Restoration[C]. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025: 28124–28133.

[24] Dai J, Qi H, Xiong Y, et al. Deformable Convolutional Networks[C]. 2017 IEEE International Conference on Computer Vision (ICCV), 2017: 764–773.

[25] Xia Z, Pan X, Song S, et al. Vision Transformer with Deformable Attention[C]. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022: 4784–4793.

[26] Agustsson E, Timofte R. NTIRE 2017 Challenge on Single Image Super-Resolution: Dataset and Study[C]. 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2017: 1122–1131.