{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,20]],"date-time":"2025-12-20T22:11:34Z","timestamp":1766268694203,"version":"3.44.0"},"reference-count":78,"publisher":"Association for Computing Machinery (ACM)","issue":"9","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62276266, 61801198, 62272461, 62172417, 62106268"],"award-info":[{"award-number":["62276266, 61801198, 62272461, 62172417, 62106268"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>Current diffusion model-based image captioning methods generally focus on generating descriptions in a non-autoregressive manner. Nevertheless, it is not trivial to employ such generative models to control the generation of discrete words while pursuing the balance between diversity and accuracy. Inspired by the success of continuous diffusions in image captioning, we introduce the Part-of-Speech (POS) information and classifier-free guidance into the diffusion model, and propose a novel controllable image captioning model, namely POS-Conditional Diffusion Networks (POSCD-Net), which consists of a Diffusion-based POS Generator (DPG) and a Diffusion-based Caption Generator (DCG). The DPG is built to produce diverse syntactic structures for each input image. The diverse POS sequences are further regarded as the control signals of the DCG, which produces the output sentences in a conditional diffusion process. In the DCG, a syntactic control module (SCM) is designed to strengthen the alignment progressively between words and the corresponding POS tags in a cascaded manner. Furthermore, to improve the controllability of POSCD-Net, the classifier-free guidance with learnable parameters is exploited to jointly optimize both the DPG and DCG in a non-autoregressive manner. Extensive experiments on the MSCOCO dataset demonstrate that our proposed method outperforms the state-of-the-art non-autoregressive counterparts and achieves promising performance compared with the autoregressive models.<\/jats:p>","DOI":"10.1145\/3748653","type":"journal-article","created":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T17:12:27Z","timestamp":1753377147000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Syntactic-Conditional Diffusion Networks for Controllable Image Captioning"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2365-6606","authenticated-orcid":false,"given":"Bing","family":"Liu","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology, Xuzhou, China, School of Artificial intelligence, China University of Mining and Technology, Xuzhou, China, and Engineering Research Center of Mine Digitization of the Ministry of Education of the People\u2019s Republic of China, China University of Mining and Technology, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-3905-6128","authenticated-orcid":false,"given":"Wenjie","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology\u2013Xuzhou Campus, Xuzhou, China, School of Artificial intelligence, School of Artificial intelligence, China University of Mining and Technology, Xuzhou, China, and Engineering Research Center of\u00a0Mine\u00a0Digitization of the Ministry of Education of the People\u2019s Republic of China, China University of\u00a0Mining\u00a0and Technology, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5698-8308","authenticated-orcid":false,"given":"Mingming","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Intelligent Manufacturing, Jiangsu Vocational Institute of Architectural Technology, Xuzhou, China, School of Computer Science and Technology, China University of Mining and Technology, Xuzhou, China, School of Artificial intelligence, China University of Mining and Technology, Xuzhou, China, and Engineering Research Center of Mine Digitization of the Ministry of Education of the People\u2019s Republic of China, China University of Mining and Technology, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6728-1773","authenticated-orcid":false,"given":"Hao","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Medical Imaging, Xuzhou Medical University, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6207-0299","authenticated-orcid":false,"given":"Yong","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology, Xuzhou, China, Engineering Research Center of Mine Digitization of the Ministry of Education of the People\u2019s Republic of China, China University of Mining and Technology, Xuzhou, China, and School of Artificial intelligence, China University of Mining and Technology, Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5183-8768","authenticated-orcid":false,"given":"Peng","family":"Liu","sequence":"additional","affiliation":[{"name":"IoT Perception Mine Research Center, China University of Mining and Technology,\u00a0Xuzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,10]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"23716","article-title":"Flamingo: A visual language model for few-shot learning","volume":"35","author":"Alayrac Jean-Baptiste","year":"2022","unstructured":"Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: A visual language model for few-shot learning. In Advances in Neural Information Processing Systems, Vol. 35, 23716\u201323736.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_3_2","first-page":"382","volume-title":"14th European Conference on Computer Vision (ECCV \u201916)","author":"Anderson Peter","year":"2016","unstructured":"Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. Spice: Semantic propositional image caption evaluation. In 14th European Conference on Computer Vision (ECCV \u201916). Springer, 382\u2013398."},{"key":"e_1_3_1_4_2","first-page":"4260","volume-title":"In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919)","author":"Aneja Jyoti","year":"2019","unstructured":"Jyoti Aneja, Harsh Agrawal, Dhruv Batra, and Alexander G. Schwing. 2019. Sequential latent spaces for modeling the intention during diverse image captioning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919), 4260\u20134269."},{"key":"e_1_3_1_5_2","first-page":"65","volume-title":"ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization, 65\u201372."},{"key":"e_1_3_1_6_2","first-page":"32","volume-title":"33th International Conference on Neural Information Processing Systems (NIPS \u201919)","author":"Chen Fuhai","year":"2019","unstructured":"Fuhai Chen, Rongrong Ji, Jiayi Ji, Xiaoshuai Sun, Baochang Zhang, Xuri Ge, Yongjian Wu, Feiyue Huang, and Yan Wang. 2019. Variational structured semantic inference for diverse image captioning. In 33th International Conference on Neural Information Processing Systems (NIPS \u201919), Vol. 32."},{"key":"e_1_3_1_7_2","first-page":"16846","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201921)","author":"Chen Long","year":"2021","unstructured":"Long Chen, Zhihong Jiang, Jun Xiao, and Wei Liu. 2021. Human-like controllable image captioning with verb-specific semantic roles. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201921), 16846\u201316856."},{"key":"e_1_3_1_8_2","first-page":"5659","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Chen Long","year":"2017","unstructured":"Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, Wei Liu, and Tat-Seng Chua. 2017. Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917), 5659\u20135667."},{"key":"e_1_3_1_9_2","first-page":"9472","volume-title":"36th International Conference on Neural Information Processing Systems (NIPS \u201922)","volume":"35","author":"Chen Qi","year":"2022","unstructured":"Qi Chen, Chaorui Deng, and Qi Wu. 2022. Learning distinct and representative modes for image captioning. In 36th International Conference on Neural Information Processing Systems (NIPS \u201922) 35 (2022), 9472\u20139485."},{"key":"e_1_3_1_10_2","first-page":"9962","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201920)","author":"Chen Shizhe","year":"2020","unstructured":"Shizhe Chen, Qin Jin, Peng Wang, and Qi Wu. 2020. Say as you wish: Fine-grained control of image caption generation with abstract scene graphs. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201920), 9962\u20139971."},{"key":"e_1_3_1_11_2","volume-title":"11th International Conference on Learning Representations (ICLR \u201923)","author":"Chen Ting","year":"2022","unstructured":"Ting Chen, Zhang Ruixiang, and Geoffrey Hinton. 2022. Analog bits: Generating discrete data using diffusion models with self-conditioning. In 11th International Conference on Learning Representations (ICLR \u201923)."},{"key":"e_1_3_1_12_2","first-page":"7995","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201918)","author":"Chen Xinpeng","year":"2018","unstructured":"Xinpeng Chen, Lin Ma, Wenhao Jiang, Jian Yao, and Wei Liu. 2018. Regularizing RNNs for caption generation by reconstructing the past with the present. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201918), 7995\u20138003."},{"key":"e_1_3_1_13_2","first-page":"24185","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Zhe","year":"2024","unstructured":"Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 24185\u201324198."},{"key":"e_1_3_1_14_2","first-page":"895","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Chunseong Park Cesc","year":"2017","unstructured":"Cesc Chunseong Park, Byeongchang Kim, and Gunhee Kim. 2017. Attend to you: Personalized image captioning with context sequence memory networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917), 895\u2013903."},{"key":"e_1_3_1_15_2","first-page":"8307","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919)","author":"Cornia Marcella","year":"2019","unstructured":"Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2019. Show, control and tell: A framework for generating controllable and grounded captions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919), 8307\u20138316."},{"key":"e_1_3_1_16_2","article-title":"A neural compositional paradigm for image captioning","volume":"31","author":"Dai Bo","year":"2018","unstructured":"Bo Dai, Sanja Fidler, and Dahua Lin. 2018. A neural compositional paradigm for image captioning. In Advances in Neural Information Processing Systems, Vol. 31.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_17_2","first-page":"2970","volume-title":"IEEE International Conference on Computer Vision (ICCV \u201917)","author":"Dai Bo","year":"2017","unstructured":"Bo Dai, Sanja Fidler, Raquel Urtasun, and Dahua Lin. 2017. Towards diverse and natural image descriptions via a conditional GAN. In IEEE International Conference on Computer Vision (ICCV \u201917), 2970\u20132979."},{"key":"e_1_3_1_18_2","first-page":"712","volume-title":"16th European Conference on Computer Vision (ECCV \u201920)","author":"Deng Chaorui","year":"2020","unstructured":"Chaorui Deng, Ning Ding, Mingkui Tan, and Qi Wu. 2020. Length-controllable image captioning. In 16th European Conference on Computer Vision (ECCV \u201920). Springer, 712\u2013729."},{"key":"e_1_3_1_19_2","first-page":"10695","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919)","author":"Deshpande Aditya","year":"2019","unstructured":"Aditya Deshpande, Jyoti Aneja, Liwei Wang, Alexander G. Schwing, and David Forsyth. 2019. Fast, diverse and accurate image captioning guided by part-of-speech. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919), 10695\u201310704."},{"key":"e_1_3_1_20_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_1_21_2","unstructured":"Jacob Devlin Saurabh Gupta Ross Girshick Margaret Mitchell and C. Lawrence Zitnick. 2015. Exploring nearest neighbor approaches for image captioning. arXiv:1505.04467. Retrieved from https:\/\/arxiv.org\/abs\/1505.04467"},{"key":"e_1_3_1_22_2","first-page":"8780","volume-title":"34th International Conference on Neural Information Processing Systems (NIPS\u201921)","volume":"34","author":"Dhariwal Prafulla","year":"2021","unstructured":"Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat gans on image synthesis. In 34th International Conference on Neural Information Processing Systems (NIPS\u201921), Vol. 34, 8780\u20138794."},{"key":"e_1_3_1_23_2","volume-title":"8th International Conference on Learning Representations (ICLR \u201920)","author":"Dosovitskiy Alexey","year":"2020","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In 8th International Conference on Learning Representations (ICLR \u201920)."},{"key":"e_1_3_1_24_2","first-page":"745","volume-title":"32th International Joint Conference on Artificial Intelligence (IJCAI \u201923)","author":"Fei Zhengcong","year":"2023","unstructured":"Zhengcong Fei and Junshi Huang. 2023. Incorporating unlikely negative cues for distinctive image captioning. In 32th International Joint Conference on Artificial Intelligence (IJCAI \u201923), 745\u2013753."},{"key":"e_1_3_1_25_2","first-page":"3137","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Gan Chuang","year":"2017","unstructured":"Chuang Gan, Zhe Gan, Xiaodong He, Jianfeng Gao, and Li Deng. 2017. Stylenet: Generating attractive visual captions with styles. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917), 3137\u20133146."},{"key":"e_1_3_1_26_2","unstructured":"Yufeng He Zefan Cai Xu Gan and Baobao Chang. 2023. DiffCap: Exploring continuous diffusion on image captioning. arXiv:2305.12144. Retrieved from https:\/\/arxiv.org\/abs\/2305.12144"},{"key":"e_1_3_1_27_2","first-page":"6840","volume-title":"In 34th International Conference on Neural Information Processing Systems (NIPS \u201920)","volume":"33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In 34th International Conference on Neural Information Processing Systems (NIPS \u201920), Vol. 33, 6840\u20136851."},{"key":"e_1_3_1_28_2","volume-title":"NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications","author":"Ho Jonathan","year":"2021","unstructured":"Jonathan Ho and Tim Salimans. 2021. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications."},{"key":"e_1_3_1_29_2","first-page":"499","volume-title":"15th European Conference on Computer Vision (ECCV \u201918)","author":"Jiang Wenhao","year":"2018","unstructured":"Wenhao Jiang, Lin Ma, Yu-Gang Jiang, Wei Liu, and Tong Zhang. 2018. Recurrent fusion network for image captioning. In 15th European Conference on Computer Vision (ECCV \u201918), 499\u2013515."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_1_31_2","first-page":"6271","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919)","author":"Kim Dong-Jin","year":"2019","unstructured":"Dong-Jin Kim, Jinsoo Choi, Tae-Hyun Oh, and In So Kweon. 2019. Dense relational captioning: Triple-stream networks for relationship-based captioning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919), 6271\u20136280."},{"key":"e_1_3_1_32_2","volume-title":"ICLR 2023 Workshop on Multimodal Representation Learning: Perks and Pitfalls","author":"Kornblith Simon","year":"2023","unstructured":"Simon Kornblith, Lala Li, Zirui Wang, and Thao Nguyen. 2023. Classifier-free guidance makes image captioning models more descriptive. In ICLR 2023 Workshop on Multimodal Representation Learning: Perks and Pitfalls."},{"key":"e_1_3_1_33_2","unstructured":"Dianqi Li Qiuyuan Huang Xiaodong He Lei Zhang and Ming-Ting Sun. 2018. Generating diverse and accurate visual captions by comparative adversarial learning. arXiv:1804.00861. Retrieved from https:\/\/arxiv.org\/abs\/1804.00861"},{"key":"e_1_3_1_34_2","first-page":"19730","article-title":"BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In 40th International Conference on Machine Learning, 19730\u201319742.","journal-title":"40th International Conference on Machine Learning"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3343520"},{"key":"e_1_3_1_36_2","first-page":"4328","volume-title":"36th International Conference on Neural Information Processing Systems (NIPS \u201922)","volume":"35","author":"Li Xiang","year":"2022","unstructured":"Xiang Li, John Thickstun, Ishaan Gulrajani, Percy S. Liang, and Tatsunori B. Hashimoto. 2022. Diffusion-lm improves controllable text generation. In 36th International Conference on Neural Information Processing Systems (NIPS \u201922), Vol. 35, 4328\u20134343."},{"key":"e_1_3_1_37_2","first-page":"121","volume-title":"16th European Conference on Computer Vision (ECCV \u201920)","author":"Li Xiujun","year":"2020","unstructured":"Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020. Oscar: Object-semantics aligned pre-training for vision-language tasks. In 16th European Conference on Computer Vision (ECCV \u201920). Springer, 121\u2013137."},{"key":"e_1_3_1_38_2","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out, 74\u201381."},{"key":"e_1_3_1_39_2","first-page":"1922","volume-title":"28th International Conference on Computational Linguistics (ACL \u201920)","author":"Lindh Annika","year":"2020","unstructured":"Annika Lindh, Robert Ross, and John Kelleher. 2020. Language-driven region pointer advancement for controllable image captioning. In 28th International Conference on Computational Linguistics (ACL \u201920), 1922\u20131935."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3107035"},{"key":"e_1_3_1_41_2","first-page":"12954","volume-title":"Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (COLING \u201924)","author":"Liu Guisheng","year":"2024","unstructured":"Guisheng Liu, Yi Li, Zhengcong Fei, Haiyan Fu, Xiangyang Luo, and Yanqing Guo. 2024. Prefix-diffusion: A lightweight diffusion model for diverse image captioning. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (COLING \u201924), 12954\u201312965."},{"key":"e_1_3_1_42_2","first-page":"34892","article-title":"Visual instruction tuning","volume":"36","author":"Liu Haotian","year":"2023","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. In Advances in Neural Information Processing Systems, Vol. 36, 34892\u201334916.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_43_2","volume-title":"6th International Conference on Learning Representations (ICLR \u201918)","author":"Loshchilov Ilya","year":"2018","unstructured":"Ilya Loshchilov and Frank Hutter. 2018. Decoupled weight decay regularization. In 6th International Conference on Learning Representations (ICLR \u201918)."},{"key":"e_1_3_1_44_2","first-page":"23359","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201923)","author":"Luo Jianjie","year":"2023","unstructured":"Jianjie Luo, Yehao Li, Yingwei Pan, Ting Yao, Jianlin Feng, Hongyang Chao, and Tao Mei. 2023. Semantic-conditional diffusion networks for image captioning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201923), 23359\u201323368."},{"key":"e_1_3_1_45_2","volume-title":"7th International Conference on Learning Representations (ICLR \u201919)","author":"Mahajan Shweta","year":"2019","unstructured":"Shweta Mahajan, Iryna Gurevych, and Stefan Roth. 2019. Latent normalizing flows for many-to-many cross-domain mappings. In 7th International Conference on Learning Representations (ICLR \u201919)."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3496029"},{"key":"e_1_3_1_47_2","unstructured":"Junhua Mao Wei Xu Yi Yang Jiang Wang Zhiheng Huang and Alan Yuille. 2014. Deep captioning with multimodal recurrent neural networks (m-rnn). arXiv:1412.6632. Retrieved from https:\/\/arxiv.org\/abs\/1412.6632"},{"key":"e_1_3_1_48_2","volume-title":"AAAI Conference on Artificial Intelligence (AAAI \u201916)","volume":"30","author":"Mathews Alexander","year":"2016","unstructured":"Alexander Mathews, Lexing Xie, and Xuming He. 2016. Senticap: Generating image descriptions with sentiments. In AAAI Conference on Artificial Intelligence (AAAI \u201916), Vol. 30."},{"key":"e_1_3_1_49_2","first-page":"8591","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201918)","author":"Mathews Alexander","year":"2018","unstructured":"Alexander Mathews, Lexing Xie, and Xuming He. 2018. Semstyle: Learning to generate stylised image captions using unaligned text. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201918), 8591\u20138600."},{"key":"e_1_3_1_50_2","unstructured":"Ron Mokady Amir Hertz and Amit H. Bermano. 2021. Clipcap: Clip prefix for image captioning. arXiv:2111.09734. Retrieved from https:\/\/arxiv.org\/abs\/2111.09734"},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","unstructured":"David Nukrai Ron Mokady and Amir Globerson. 2022. Text-only training for image captioning using noise-injected clip. arXiv:2211.00575. Retrieved from https:\/\/arxiv.org\/abs\/2211.00575","DOI":"10.18653\/v1\/2022.findings-emnlp.299"},{"key":"e_1_3_1_52_2","first-page":"311","volume-title":"40th Annual Meeting of the Association for Computational Linguistics (ACL \u201902)","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In 40th Annual Meeting of the Association for Computational Linguistics (ACL \u201902), 311\u2013318."},{"key":"e_1_3_1_53_2","first-page":"2641","volume-title":"IEEE International Conference on Computer Vision","author":"Plummer Bryan A.","year":"2015","unstructured":"Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In IEEE International Conference on Computer Vision, 2641\u20132649."},{"key":"e_1_3_1_54_2","first-page":"647","volume-title":"16th European Conference on Computer Vision (ECCV \u201920)","author":"Pont-Tuset Jordi","year":"2020","unstructured":"Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, and Vittorio Ferrari. 2020. Connecting vision and language with localized narratives. In 16th European Conference on Computer Vision (ECCV \u201920). Springer, 647\u2013664."},{"key":"e_1_3_1_55_2","first-page":"8748","volume-title":"38th International Conference on Machine Learning (ICML \u201921)","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In 38th International Conference on Machine Learning (ICML \u201921). PMLR, 8748\u20138763."},{"key":"e_1_3_1_56_2","unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv:2204.06125. Retrieved from https:\/\/arxiv.org\/abs\/2204.06125"},{"key":"e_1_3_1_57_2","first-page":"4135","volume-title":"IEEE International Conference on Computer Vision (ICCV \u201917)","author":"Shetty Rakshith","year":"2017","unstructured":"Rakshith Shetty, Marcus Rohrbach, Lisa Anne Hendricks, Mario Fritz, and Bernt Schiele. 2017. Speaking the same language: Matching machine to human captions by adversarial training. In IEEE International Conference on Computer Vision (ICCV \u201917), 4135\u20134144."},{"key":"e_1_3_1_58_2","first-page":"12516","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919)","author":"Shuster Kurt","year":"2019","unstructured":"Kurt Shuster, Samuel Humeau, Hexiang Hu, Antoine Bordes, and Jason Weston. 2019. Engaging image captioning via personality. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919), 12516\u201312526."},{"key":"e_1_3_1_59_2","first-page":"3483","volume-title":"28th International Conference on Neural Information Processing Systems (NIPS \u201915)","author":"Sohn Kihyuk","year":"2015","unstructured":"Kihyuk Sohn, Honglak Lee, and Xinchen Yan. 2015. Learning structured output representation using deep conditional generative models. In 28th International Conference on Neural Information Processing Systems (NIPS \u201915), 3483\u20133491."},{"key":"e_1_3_1_60_2","unstructured":"Zhicong Tang Shuyang Gu Jianmin Bao Dong Chen and Fang Wen. 2022. Improved vector quantized diffusion models. arXiv:2205.16007. Retrieved from https:\/\/arxiv.org\/abs\/2205.16007"},{"key":"e_1_3_1_61_2","first-page":"11","article-title":"Visualizing data using t-SNE","volume":"9","author":"Van der Maaten Laurens","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9 (2008), 11.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_62_2","first-page":"4566","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201915)","author":"Vedantam Ramakrishna","year":"2015","unstructured":"Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201915), 4566\u20134575."},{"key":"e_1_3_1_63_2","volume-title":"AAAI Conference on Artificial Intelligence (AAAI \u201918)","volume":"32","author":"Vijayakumar Ashwin","year":"2018","unstructured":"Ashwin Vijayakumar, Michael Cogswell, Ramprasaath Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2018. Diverse beam search for improved description of complex scenes. In AAAI Conference on Artificial Intelligence (AAAI \u201918), Vol. 32."},{"key":"e_1_3_1_64_2","unstructured":"Jianfeng Wang Zhengyuan Yang Xiaowei Hu Linjie Li Kevin Lin Zhe Gan Zicheng Liu Ce Liu and Lijuan Wang. 2022. GIT: A generative image-to-text transformer for vision and language. arXiv:2205.14100. Retrieved from https:\/\/arxiv.org\/abs\/2205.14100"},{"key":"e_1_3_1_65_2","volume-title":"31th International Conference on Neural Information Processing Systems (NIPS \u201917)","volume":"30","author":"Wang Liwei","year":"2017","unstructured":"Liwei Wang, Alexander Schwing, and Svetlana Lazebnik. 2017. Diverse and accurate image description using a variational auto-encoder with an additive Gaussian encoding space. In 31th International Conference on Neural Information Processing Systems (NIPS \u201917), Vol. 30."},{"key":"e_1_3_1_66_2","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919)","author":"Wang Qingzhong","year":"2019","unstructured":"Qingzhong Wang and Antoni B. Chan. 2019. Describing like humans: On diversity in image captioning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919)."},{"key":"e_1_3_1_67_2","first-page":"121475","article-title":"Cogvlm: Visual expert for pretrained language models","volume":"37","author":"Wang Weihan","year":"2024","unstructured":"Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Song XiXuan, et al. 2024. Cogvlm: Visual expert for pretrained language models. In Advances in Neural Information Processing Systems, Vol. 37, 121475\u2013121499.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_68_2","first-page":"6699","volume-title":"2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Wang Yuchi","year":"2024","unstructured":"Yuchi Wang, Shuhuai Ren, Rundong Gao, Linli Yao, Qingyan Guo, Kaikai An, Jianhong Bai, and Xu Sun. 2024. LaDiC: Are diffusion models really inferior to autoregressive counterparts for image-to-text generation? In 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 6699\u20136715."},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3121062"},{"issue":"1","key":"e_1_3_1_70_2","first-page":"1","article-title":"Diverse image captioning via conditional variational autoencoder and dual contrastive learning","volume":"20","author":"Xu Jing","year":"2023","unstructured":"Jing Xu, Bing Liu, Yong Zhou, Mingming Liu, Rui Yao, and Zhiwen Shao. 2023. Diverse image captioning via conditional variational autoencoder and dual contrastive learning. ACM Transactions on Multimedia Computing, Communications, and Applications 20, 1 (2023), 1\u201316.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_71_2","unstructured":"Shitong Xu. 2022. Clip-diffusion-lm: Apply diffusion model on image captioning. arXiv:2210.04559. Retrieved from https:\/\/arxiv.org\/abs\/2210.04559"},{"key":"e_1_3_1_72_2","first-page":"1085","volume-title":"ACM International Conference on Multimedia (MM \u201920)","author":"Yuan Yitian","year":"2020","unstructured":"Yitian Yuan, Lin Ma, Jingwen Wang, and Wenwu Zhu. 2020. Controllable video captioning with an exemplar sentence. In ACM International Conference on Multimedia (MM \u201920), 1085\u20131093."},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3336371"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3243725"},{"key":"e_1_3_1_75_2","first-page":"13096","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201920)","author":"Zheng Qi","year":"2020","unstructured":"Qi Zheng, Chaoyue Wang, and Dacheng Tao. 2020. Syntax-aware action targeting for video captioning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201920), 13096\u201313105."},{"key":"e_1_3_1_76_2","first-page":"8395","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919)","author":"Zheng Yue","year":"2019","unstructured":"Yue Zheng, Yali Li, and Shengjin Wang. 2019. Intention oriented image captions with guiding objects. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201919), 8395\u20138404."},{"key":"e_1_3_1_77_2","article-title":"Polishing network for decoding of higher-quality diverse image captions","volume":"601","author":"Zheng Yue","year":"2022","unstructured":"Yue Zheng, Ya-Li Li, and Shengjin Wang. 2022. Polishing network for decoding of higher-quality diverse image captions. In British Machine Vision Conference, Vol. 601.","journal-title":"British Machine Vision Conference"},{"key":"e_1_3_1_78_2","first-page":"211","volume-title":"16th European Conference on Computer Vision (ECCV \u201920)","author":"Zhong Yiwu","year":"2020","unstructured":"Yiwu Zhong, Liwei Wang, Jianshu Chen, Dong Yu, and Yin Li. 2020. Comprehensive image captioning via scene graph decomposition. In 16th European Conference on Computer Vision (ECCV \u201920). Springer, 211\u2013229."},{"key":"e_1_3_1_79_2","unstructured":"Zixin Zhu Yixuan Wei Jianfeng Wang Zhe Gan Zheng Zhang Le Wang Gang Hua Lijuan Wang Zicheng Liu and Han Hu. 2022. Exploring discrete diffusion models for image captioning. arXiv:2211.11694. Retrieved from https:\/\/arxiv.org\/abs\/2211.11694"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3748653","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,10]],"date-time":"2025-09-10T16:02:40Z","timestamp":1757520160000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3748653"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,10]]},"references-count":78,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3748653"],"URL":"https:\/\/doi.org\/10.1145\/3748653","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2025,9,10]]},"assertion":[{"value":"2024-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-29","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}