1. 大连交通大学, 自动化与电气工程学院, 辽宁大连 116028
2. 天津轨道交通运营集团有限公司, 天津 300380
3. 大连交通大学, 计算机与通信工程学院, 辽宁大连 116028
| 摘 要: | 人体姿态估计网络模型性能逐渐提高,过深的网络结构带来的是更为庞大的参数量和复杂计算量。针对以上问题,提出了FastPose-Lite轻量型人体姿态估计网络,该网络由GSE-ResNet特征提取网络、上采样DUC模块以及CBAM模块组成。其中构成GSE-ResNet特征提取网络的GBNK基础模块是由Ghost模块和SE模块组成,一方面,为减少参数量和计算量提出利用Ghost模块代替传统卷积模块;另一方面,为保证网络模型性能不变,引入SE注意力机制模块。在空间和通道两方面为增强了网络模型对特征信息的处理能力,将CBAM模块引入到上采样DUC模块之间,减少了上采样过程中带来的损失。在COCO数据集上的实验结果表明,本文提出的FastPose-Lite相对于FastPose网络模型的参数量和计算量分别减少了51.4%和50.8%;与SHN、CPN和SimpleBaseline等常见的热门网络模型相比,FastPose-Lite网络模型不仅参数量和计算量更少,而且预测精度更高。 |
| 关 键 词: | 人体姿态估计; FastPose; Ghost模块; 注意力机制; 轻量化 |
| DOI: | 10.57237/j.cst.2023.02.003 |
1. School of Automation and Electrical Engineering, Dalian Jiaotong University, Dalian 116028, China
2. Tianjin Rail Transit Group Corporation, Tianjin 300380, China
3. School of Computer and Communication Engineering, Dalian Jiaotong University, Dalian 116028, China
| Abstract: | The performance of human pose estimation network model is gradually improved, and the over-deep network structure brings a large number of parameters and complex calculation. To solve these problems, the FastPose-Lite lightweight human pose estimation network is proposed, which is composed of GSE-ResNet feature extraction network, up-sampling DUC module and CBAM module. The basic GBNK module of GSE-ResNet feature extraction network is composed of Ghost module and SE module. On the one hand, Ghost module is proposed to replace the traditional convolutional module in order to reduce the number of parameters and calculation. On the other hand, in order to keep the performance of the network model unchanged, SE attention mechanism module is introduced. In order to enhance the processing ability of the network model to the feature information in both spatial and channel aspects, CBAM module is introduced into the up-sampling DUC module to reduce the loss in the up-sampling process. The experimental results on the COCO dataset show that the proposed FastPose-Lite reduce the number of parameters and calculation by 51.4% and 50.8% respectively compared with the FastPose network model. Compared with the common popular network models such as SHN, CPN, and SimpleBaseline, the FastPose-Lite network model not only has fewer parameters and computations, but also has higher prediction accuracy. |
| Keywords: | Human Pose Estimation; FastPose; Ghost Module; Attention Mechanism; Lightweight |
| [1] | Dalal N, Triggs B. Histograms of oriented gradients for human detection [C] // 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05). Ieee, 2005, 1: 886-893. |
| [2] | Lowe D G. Distinctive image features from scale-invariant keypoints [J]. International journal of computer vision, 2004, 60 (2): 91-110. |
| [3] | LeCun Y, Boser B, Denker J S, et al. Backpropagation applied to handwritten zip code recognition [J]. Neural computation, 1989, 1 (4): 541-551. |
| [4] | Wei S E, Ramakrishna V, Kanade T, et al. Convolutional pose machines [C] // Proceedings of the IEEE conference on Computer Vision and Pattern Recognition. 2016: 4724-4732. |
| [5] | Newell A, Yang K, Deng J. Stacked hourglass networks for human pose estimation [C] // European conference on computer vision. Springer, Cham, 2016: 483-499. |
| [6] | Fang H S, Xie S, Tai Y W, et al. RMPE: Regional multi-person pose estimation [C] // Proceedings of the IEEE international conference on computer vision. 2017: 2334-2343. |
| [7] | Fang H S, Li J, Tang H, et al. AlphaPose: Whole-Body Regional Multi-Person Pose Estimation and Tracking in Real-Time [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022. |
| [8] | Cao Z, Simon T, Wei S E, et al. Realtime multi-person 2d pose estimation using part affinity fields [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 7291-7299. |
| [9] | Howard A G, Zhu M, Chen B, et al. Mobilenets: Efficient convolutional neural networks for mobile vision applications [J]. arXiv preprint arXiv: 1704.04861, 2017. |
| [10] | Sandler M, Howard A, Zhu M, et al. Mobilenetv2: Inverted residuals and linear bottlenecks [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 4510-4520. |
| [11] | Howard A, Sandler M, Chu G, et al. Searching for mobilenetv3 [C] // Proceedings of the IEEE/CVF international conference on computer vision. 2019: 1314-1324. |
| [12] | Hu J, Shen L, Sun G. Squeeze-and-excitation networks [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 7132-7141. |
| [13] | Bochkovskiy A, Wang C Y, Liao H Y M. Yolov: Optimal speed and accuracy of object detection [J]. arXiv preprint arXiv: 2004.10934, 2020. |
| [14] | He K, Zhang X, Ren S, et al. Deep residual learning for image recognition [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778. |
| [15] | Woo S, Park J, Lee J Y, et al. Cbam: Convolutional block attention module [C] // Proceedings of the European conference on computer vision (ECCV). 2018: 3-19. |
| [16] | Han K, Wang Y, Tian Q, et al. Ghostnet: More features from cheap operations [C] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020: 1580-1589. |
| [17] | Wang P, Chen P, Yuan Y, et al. Understanding convolution for semantic segmentation [C] // 2018 IEEE winter conference on applications of computer vision (WACV). Ieee, 2018: 1451-1460. |
| [18] | Lin T Y, Maire M, Belongie S, et al. Microsoft coco: Common objects in context [C] // European conference on computer vision. Springer, Cham, 2014: 740-755. |
| [19] | Chen Y, Wang Z, Peng Y, et al. Cascaded pyramid network for multi-person pose estimation [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 7103-7112. |
| [20] | Xiao B, Wu H, Wei Y. Simple baselines for human pose estimation and tracking [C] // Proceedings of the European conference on computer vision (ECCV). 2018: 466-481. |
| [21] | Sun K, Xiao B, Liu D, et al. Deep high-resolution representation learning for human pose estimation [C] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019: 5693-5703. |