School of Computer Science and Engineering, Shenyang Jianzhu University, Shenyang 110168, China
| Abstract: | To address the challenges of large scale variations and strong background interference in dense crowd images, this paper proposes a crowd counting model named Multi-scale Selective Density Aggregation Network (MSDA-Net). The model introduces three collaborative modules: the Multi-scale Tuning Module (MTM), the Selective Attention Module (SAM), and the Density Context-aware Module (DCF), which collectively enhance feature representation and density map quality, thereby improving counting accuracy in complex congested scenes. Specifically, the MTM captures multi-scale contextual information through multi-branch atrous convolutions, alleviating the impact of drastic scale variations. The SAM employs spatial and channel attention mechanisms to suppress background noise and highlight foreground crowd regions. The DCF models the contextual dependencies of density maps using convolutional LSTM architecture, enhancing continuity between adjacent regions to generate smoother and more accurate density maps. A joint loss function combining Euclidean distance loss and structural similarity (SSIM) loss is adopted to optimize pixel-level regression and local structural consistency simultaneously. Extensive experiments are conducted on four public benchmarks, including ShanghaiTech Part_A, ShanghaiTech Part_B, UCF_CC_50, and UCF_QNRF. The proposed MSDA-Net achieves superior performance on all datasets, e.g., MAE of 58.1 on Part_A, 6.9 on Part_B, 205.3 on UCF_CC_50, and 84.8 on UCF_QNRF, outperforming many state-of-the-art methods. Ablation studies further validate the individual contribution of each module. Overall, MSDA-Net demonstrates strong generalization ability and practical application value in real-world crowded scenes. |
| Keywords: | Dense Crowd Counting; Multi-scale Modulation; Attention Mechanism; Context Modeling; MSDA-Net |
| DOI: | 10.57237/j.cst.2026.03.001 |
| 1. | 国家自然科学基金 (62073227); 辽宁省科技厅项目 (2023JH2/101300212) |
| [1] | ZANNELLA C, BOVE S. Smart city context and the role of crowd analysis [J]. Sustainable Cities and Society, 2021, 68: 102761. https://www.sciencedirect.com/science/article/pii/S2210670721000226 |
| [2] | ZHANG J, ZHENG Y, QI D. Deep spatio-temporal residual networks for citywide crowd flows prediction [C] // Proceedings of the AAAI Conference on Artificial Intelligence. 2017: 1655-1661. https://doi.org/10.1609/aaai.v31i1.10735 |
| [3] | SINDAGI V A, PATEL V M. A survey of recent advances in CNN-based single image crowd counting and density estimation [J]. Pattern Recognition Letters, 2018, 107: 3-16. https://www.sciencedirect.com/science/article/pii/S0167865517302111 |
| [4] | HELBING D, FARKAS I, VICSEK T. Simulating dynamical features of escape panic [J]. Nature, 2000, 407(6803): 487-490. https://doi.org/10.1038/35035023 |
| [5] | WANG B, LIU H, SAMARAS D, et al. Distribution matching for crowd counting [C] // Advances in Neural Information Processing Systems (NeurIPS). 2020: 159-171. https://doi.org/10.48550/arXiv.2009.13077 |
| [6] | WANG Z, BOVIK A C, SHEIKH H R, et al. Image quality assessment: from error visibility to structural similarity [J]. IEEE Transactions on Image Processing, 2004, 13(4): 600-612. https://doi.org/10.1109/TIP.2003.819861 |
| [7] | DALAL N, TRIGGS B. Histograms of oriented gradients for human detection [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2005: 886-893. https://doi.org/10.1109/CVPR.2005.177 |
| [8] | IDREES H, SALEEMI I, SEIBERT C, et al. Multi-source multi-scale counting in extremely dense crowd images [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2013: 2547-2554. https://ieeexplore.ieee.org/document/6615976 |
| [9] | LEMPITSKY V, ZISSERMAN A. Learning to count objects in images [C] // Advances in Neural Information Processing Systems (NeurIPS). 2010: 1324-1332. https://papers.nips.cc/paper/2010/hash/819f46e52c25763a55cc642422644317-Abstract.html |
| [10] | ZHANG Y, ZHOU D, CHEN S, et al. Single-image crowd counting via multi-column convolutional neural network [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016: 589-597. |
| [11] | Li Y, Zhang X, Chen D. CSRNet: dilated convolutional neural networks for understanding the highly congested scenes [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. Washington: IEEE Computer Society, 2018: 1091-1100. https://doi.org/10.1109/CVPR.2018.00120 |
| [12] | CHENG Z Q, LI J X, DAI Z, et al. Learning spatial awareness to improve crowd counting [C] // Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2019: 6152-6161. https://openaccess.thecvf.com/content_ICCV_2019/html/Cheng_Learning_Spatial_Awareness_to_Improve_Crowd_Counting_ICCV_2019_paper.html |
| [13] | CAO X, WANG Z, ZHAO Y, et al. Scale aggregation network for accurate and efficient crowd counting [C] // Proceedings of the European Conference on Computer Vision (ECCV). 2018: 734-750. https://doi.org/10.1007/978-3-030-01228-1_45 |
| [14] | CHEN L, ZHANG H, XIAO J, et al. SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioning [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2017: 5659-5667. https://openaccess.thecvf.com/content_cvpr_2017/html/Chen_SCA-CNN_Spatial_and_CVPR_2017_paper.html |
| [15] | MA Z, WEI X, HONG X, et al. Bayesian loss for crowd count estimation with point supervision [C] // Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2019: 6142-6151. https://openaccess.thecvf.com/content_ICCV_2019/html/Ma_Bayesian_Loss_for_Crowd_Count_Estimation_With_Point_Supervision_ICCV_2019_paper.html |
We invite active, qualified and high profile scientists and researchers to join as Editorial Board Members.
Join UsScholars with a strong interest in reviewing are invited to join the reviewer panel to ensure the quality of the research to be published.
Join Us