{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T09:58:59Z","timestamp":1775901539893,"version":"3.50.1"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/501100003453","name":"Natural Science Foundation of Guangdong Province","doi-asserted-by":"crossref","award":["2025A1515010454"],"award-info":[{"award-number":["2025A1515010454"]}],"id":[{"id":"10.13039\/501100003453","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62206060"],"award-info":[{"award-number":["62206060"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>Speech-Preserving Facial Expression Manipulation (SPFEM) is an innovative technique aimed at altering facial expressions in images and videos while retaining the original mouth movements. Despite advancements, SPFEM still struggles with accurate lip synchronization due to the complex interplay between facial expressions and mouth shapes. Capitalizing on the advanced capabilities of Audio-Driven Talking Head Generation (AD-THG) models in synthesizing precise lip movements, our research introduces a novel integration of these models with SPFEM. We present a new framework, Talking Head Facial Expression Manipulation (THFEM), which utilizes AD-THG models to generate frames with accurately synchronized lip movements from audio inputs and SPFEM-altered images. However, increasing the number of frames generated by AD-THG models tends to compromise the realism and expression fidelity of the images. To counter this, we develop an adjacent frame learning strategy that finetunes AD-THG models to predict sequences of consecutive frames. This strategy enables the models to incorporate information from neighboring frames, significantly improving image quality during testing. Our extensive experimental evaluations demonstrate that this framework effectively preserves mouth shapes during expression manipulations, highlighting the substantial benefits of integrating AD-THG with SPFEM.<\/jats:p>","DOI":"10.1145\/3795139","type":"journal-article","created":{"date-parts":[[2026,3,17]],"date-time":"2026-03-17T20:19:13Z","timestamp":1773778753000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Exploring Talking Head Models with Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-9898-1546","authenticated-orcid":false,"given":"Zhenxuan","family":"Lu","sequence":"first","affiliation":[{"name":"Guangdong University of Technology, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0731-4585","authenticated-orcid":false,"given":"Zhihua","family":"Xu","sequence":"additional","affiliation":[{"name":"Guangdong University of Technology, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8336-5109","authenticated-orcid":false,"given":"Zhijing","family":"Yang","sequence":"additional","affiliation":[{"name":"Information Engineering, Guangdong University of Technology, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7374-9491","authenticated-orcid":false,"given":"Feng","family":"Gao","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1398-9965","authenticated-orcid":false,"given":"Yongyi","family":"Lu","sequence":"additional","affiliation":[{"name":"Guangdong University of Technology, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7817-8306","authenticated-orcid":false,"given":"Keze","family":"Wang","sequence":"additional","affiliation":[{"name":"Sun Yat-Sen University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5848-5624","authenticated-orcid":false,"given":"Tianshui","family":"Chen","sequence":"additional","affiliation":[{"name":"Information Engineering, Guangdong University of Technology, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,11]]},"reference":[{"key":"e_1_3_1_2_2","volume-title":"Proceedings of the 26th Annual Conference on Computer Graphics and Interactive techniques (SIGGRAPH \u201999)","author":"Blanz Volker","year":"1999","unstructured":"Volker Blanz and Thomas Vetter. 1999. A morphable model for the synthesis of 3D faces. In Proceedings of the 26th Annual Conference on Computer Graphics and Interactive techniques (SIGGRAPH \u201999). DOI: 10.1145\/311535.311556"},{"key":"e_1_3_1_3_2","volume-title":"2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Chen Lele","year":"2019","unstructured":"Lele Chen, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu. 2019. Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). DOI: 10.1109\/cvpr.2019.00802"},{"key":"e_1_3_1_4_2","first-page":"7267","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Tianshui","year":"2024","unstructured":"Tianshui Chen, Jianman Lin, Zhijing Yang, Chunmei Qing, and Liang Lin. 2024. Learning adaptive spatial coherent correlations for speech-preserving facial expression manipulation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 7267\u20137276."},{"issue":"7","key":"e_1_3_1_5_2","doi-asserted-by":"crossref","first-page":"3822","DOI":"10.1007\/s11263-025-02358-x","article-title":"Contrastive decoupled representation learning and regularization for speech-preserving facial expression manipulation","volume":"133","author":"Chen Tianshui","year":"2025","unstructured":"Tianshui Chen, Jianman Lin, Zhijing Yang, Chumei Qing, Yukai Shi, and Liang Lin. 2025. Contrastive decoupled representation learning and regularization for speech-preserving facial expression manipulation. International Journal of Computer Vision 133, 7 (2025), 3822\u20133838.","journal-title":"International Journal of Computer Vision"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","unstructured":"Joon Son Chung Amir Jamaludin and Andrew Zisserman. 2017. You said that? arXiv:1705.02966. Retrieved from https:\/\/arxiv.org\/abs\/1705.02966","DOI":"10.5244\/C.31.109"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","unstructured":"Joon Son Chung and Andrew Zisserman. 2017. Out of time: Automated lip sync in the wild. In Proceedings of Workshops on Computer Vision (ACCV \u201916). C. S. Chen J. Lu K. K. Ma (Eds.) Lecture Notes in Computer Science Vol. 10117 Springer Cham 251\u2013263. DOI: 10.1007\/978-3-319-54427-4_19","DOI":"10.1007\/978-3-319-54427-4_19"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/tpami.2021.3087709"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TBIOM.2021.3049576"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","unstructured":"Paul Ekman and WallaceV Friesen. 1978. Facial action coding system: A technique for the measurement of facial movement. DOI: 10.1037\/t27734-000","DOI":"10.1037\/t27734-000"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"Yuan Gan Zongxin Yang Xihang Yue Lingyun Sun and Yi Yang. 2023. Efficient Emotional Adaptation for Audio-Driven Talking-Head Generation.","DOI":"10.1109\/ICCV51070.2023.02069"},{"key":"e_1_3_1_12_2","unstructured":"Ian J. Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville Yoshua Bengio and Delhi Delhi. 2014. Generative adversarial nets. In NIPS'14: Proceedings of the 28th International Conference on Neural Information Processing Systems Vol. 2 2672\u20132680. DOI: https:\/\/dl.acm.org\/doi\/10.5555\/2969033.2969125"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.123698"},{"key":"e_1_3_1_14_2","first-page":"5784","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Guo Yudong","year":"2021","unstructured":"Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang. 2021. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 5784\u20135794."},{"key":"e_1_3_1_15_2","first-page":"20914","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Gururani Siddharth","year":"2023","unstructured":"Siddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle, and Ming-Yu Liu. 2023. Space: Speech-driven portrait animation with controllable expression. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 20914\u201320923."},{"key":"e_1_3_1_16_2","first-page":"6629","volume-title":"Neural Information Processing Systems","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local nash equilibrium. In Neural Information Processing Systems, 6629\u20136640."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530745"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr46437.2021.01386"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","unstructured":"Justin Johnson Alexandre Alahi and Li Fei-Fei. 2016. Perceptual losses for real-time style transfer and super-resolution. In Proceedings of European Conference on Computer Vision (ECCV \u201916). B. Leibe J. Matas N. Sebe and M. Welling (Eds.) Lecture Notes in Computer Science Vol. 9906 Springer Cham 694\u2013711. DOI: 10.1007\/978-3-319-46475-6_43","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"e_1_3_1_20_2","first-page":"852","volume-title":"Neural Information Processing System","author":"Karras Tero","year":"2021","unstructured":"Tero Karras, Miika Aittala, Samuli Laine, Erik H\u00e4rk\u00f6nen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2021. Alias-free generative adversarial networks. In Neural Information Processing System, 852\u2013863."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00590"},{"key":"e_1_3_1_22_2","first-page":"5549","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lee Cheng-Han","year":"2020","unstructured":"Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo. 2020. Maskgan: Towards diverse and interactive facial image manipulation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 5549\u20135558."},{"key":"e_1_3_1_23_2","first-page":"3387","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liang Borong","year":"2022","unstructured":"Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou, Zhibin Hong, Xiaoguang Han, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang. 2022. Expressive talking head generation with granular audio-visual control. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3387\u20133396."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571746"},{"issue":"9","key":"e_1_3_1_25_2","first-page":"1","article-title":"Multimodal fusion for talking face generation utilizing speech-related facial action units","volume":"20","author":"Liu Zhilei","year":"2024","unstructured":"Zhilei Liu, Xiaoxing Liu, Sen Chen, Jiaxing Liu, Longbiao Wang, and Chongke Bi. 2024. Multimodal fusion for talking face generation utilizing speech-related facial action units. ACM Transactions on Multimedia Computing, Communications, and Applications 20, 9 (2024), 1\u201324.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_26_2","first-page":"11","volume-title":"Ismir","volume":"270","author":"Logan Beth","year":"2000","unstructured":"Beth Logan, et al. 2000. Mel frequency cepstral coefficients for music modeling. In Ismir, Vol. 270, 11."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00049"},{"key":"e_1_3_1_28_2","volume-title":"2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Papantoniou Foivos Paraperas","year":"2022","unstructured":"Foivos Paraperas Papantoniou, Panagiotis P. Filntisis, Petros Maragos, and Anastasios Roussos. 2022. NED. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). DOI: 10.1109\/cvpr52688.2022.01822"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413532"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-019-01210-3"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/taffc.2019.2916031"},{"issue":"1","key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1109\/TASSP.1978.1163055","article-title":"Dynamic programming algorithm optimization for spoken word recognition","volume":"26","author":"Sakoe Hiroaki","year":"1978","unstructured":"Hiroaki Sakoe and Seibi Chiba. 1978. Dynamic programming algorithm optimization for spoken word recognition. IEEE Transactions on Acoustics, Speech, and Signal Processing 26, 1 (1978), 43\u201349.","journal-title":"IEEE Transactions on Acoustics, Speech, and Signal Processing"},{"key":"e_1_3_1_33_2","first-page":"1","volume-title":"International Conference on Learning Representations","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations, 1\u201314"},{"key":"e_1_3_1_34_2","first-page":"104","volume-title":"European Conference on Computer Vision","author":"Kumar Solanki Girish","year":"2022","unstructured":"Girish Kumar Solanki and Anastasios Roussos. 2022. Deep semantic manipulation of facial videos. In European Conference on Computer Vision. Springer, 104\u2013120."},{"key":"e_1_3_1_35_2","volume-title":"Proceedings of the 28th International Joint Conference on Artificial Intelligence","author":"Song Yang","year":"2019","unstructured":"Yang Song, Jingwen Zhu, Dawei Li, Andy Wang, and Hairong Qi. 2019. Talking face generation by conditional recurrent adversarial network. In Proceedings of the 28th International Joint Conference on Artificial Intelligence. DOI: 10.24963\/ijcai.2019\/129"},{"issue":"12","key":"e_1_3_1_36_2","first-page":"1","article-title":"Multi-modal driven pose-controllable talking head generation","volume":"20","author":"Sun Kuiyuan","year":"2024","unstructured":"Kuiyuan Sun, Xiaolong Liu, Xiaolong Li, Yao Zhao, and Wei Wang. 2024. Multi-modal driven pose-controllable talking head generation. ACM Transactions on Multimedia Computing, Communications, and Applications 20, 12 (2024), 1\u201323.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","unstructured":"Zhiyao Sun Yu-Hui Wen Tian Lv Yanan Sun Ziyang Zhang Yaoyuan Wang and Yong-Jin Liu. 2024. Continuously controllable facial expression editing in talking face videos. IEEE Transactions on Affective Computing 15 3 (2024) 1400\u20131413. DOI: 10.1109\/TAFFC.2023.3334511","DOI":"10.1109\/TAFFC.2023.3334511"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.119678"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093474"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3550469.3555382"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-019-01251-8"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58589-1_42"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i3.20154"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr42600.2020.00507"},{"issue":"2","key":"e_1_3_1_45_2","first-page":"1","article-title":"Progressive transformer machine for natural character reenactment","volume":"19","author":"Xu Yongzong","year":"2023","unstructured":"Yongzong Xu, Zhijing Yang, Tianshui Chen, Kai Li, and Chunmei Qing. 2023. Progressive transformer machine for natural character reenactment. ACM Transactions on Multimedia Computing, Communications, and Applications 19, 2s (2023), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_46_2","first-page":"3800","volume-title":"Proceedings of the 32nd ACM International Conference on Multimedia","author":"Xu Zhihua","year":"2024","unstructured":"Zhihua Xu, Tianshui Chen, Zhijing Yang, Chunmei Qing, Yukai Shi, and Liang Lin. 2024. Self-supervised emotion representation disentanglement for speech-preserving facial expression manipulation. In Proceedings of the 32nd ACM International Conference on Multimedia, 3800\u20133808."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i4.16465"},{"key":"e_1_3_1_48_2","first-page":"354","volume-title":"International Conference on Brain Inspired Cognitive Systems","author":"Zheng Si","year":"2023","unstructured":"Si Zheng, Junbin Chen, Zhijing Yang, Tianshui Chen, and Yongyi Lu. 2023. Face reenactment based on motion field representation. In International Conference on Brain Inspired Cognitive Systems. Springer Nature Singapore, Singapore, 354\u2013364."},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019299"},{"key":"e_1_3_1_50_2","first-page":"4176","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhou Hang","year":"2021","unstructured":"Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. 2021. Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 4176\u20134186."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417774"},{"key":"e_1_3_1_52_2","first-page":"2223","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Zhu Jun-Yan","year":"2017","unstructured":"Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. 2017. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE International Conference on Computer Vision, 2223\u20132232."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3795139","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T09:19:13Z","timestamp":1775899153000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3795139"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,11]]},"references-count":51,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3795139"],"URL":"https:\/\/doi.org\/10.1145\/3795139","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,11]]},"assertion":[{"value":"2025-04-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}