{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T17:21:27Z","timestamp":1787505687640,"version":"3.56.0"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"7","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>\n                    AI programming has become a popular topic in recent years. Code suggestion, with code suggestion being a key capability of AI programming. Copilot, an \u201cAI programmer\u201d that provides code suggestions from natural language descriptions, has been launched by GitHub and OpenAI. By far, Copilot has been widely used by millions of developers. However, little work has systematically evaluated the correctness of Copilot\u2019s suggestions. We conducted an empirical study on all 2,033\n                    <jats:italic toggle=\"yes\">LeetCode<\/jats:italic>\n                    problems to assess Copilot\u2019s code generation across four mainstream languages: C, Java, JavaScript, and Python. We have found that: (1) 70.0% of problems received at least one correct suggestion, with language-specific rates of 29.7% (C), 57.7% (Java), 54.1% (JavaScript), and 41.0% (Python); (2) correctness decreases as problem difficulty increases, with acceptance rates of 89.3% (easy), 72.1% (medium), and 43.4% (hard); (3) acceptance rates vary across problem domains from 49.5% to 90.1%, while\n                    <jats:italic toggle=\"yes\">Graph<\/jats:italic>\n                    problems challenge C and Python most, and\n                    <jats:italic toggle=\"yes\">Prefix Sum<\/jats:italic>\n                    and\n                    <jats:italic toggle=\"yes\">Heap<\/jats:italic>\n                    challenge Java and JavaScript most; (4) for the incorrect suggestions, we further summarize 17 types of error reasons accounting for their incorrectness and analyzed possible causes for why these errors occur. We believe our study can provide valuable insights into Copilot\u2019s capabilities and limitations.\n                  <\/jats:p>","DOI":"10.1145\/3715108","type":"journal-article","created":{"date-parts":[[2025,1,27]],"date-time":"2025-01-27T10:53:31Z","timestamp":1737975211000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["Assessing and Analyzing the Correctness of GitHub Copilot\u2019s Code Suggestions"],"prefix":"10.1145","volume":"34","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7556-153X","authenticated-orcid":false,"given":"Ran","family":"Mo","sequence":"first","affiliation":[{"name":"Central China Normal University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6896-1687","authenticated-orcid":false,"given":"Dongyu","family":"Wang","sequence":"additional","affiliation":[{"name":"Central China Normal University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-1966-5517","authenticated-orcid":false,"given":"Wenjing","family":"Zhan","sequence":"additional","affiliation":[{"name":"Central China Normal University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-3385-7403","authenticated-orcid":false,"given":"Yingjie","family":"Jiang","sequence":"additional","affiliation":[{"name":"Central China Normal University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-8055-2425","authenticated-orcid":false,"given":"Yepeng","family":"Wang","sequence":"additional","affiliation":[{"name":"University of South Carolina, Columbia, South Carolina, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2642-5109","authenticated-orcid":false,"given":"Yuqi","family":"Zhao","sequence":"additional","affiliation":[{"name":"Central China Normal University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7258-993X","authenticated-orcid":false,"given":"Zengyang","family":"Li","sequence":"additional","affiliation":[{"name":"Central China Normal University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4239-2009","authenticated-orcid":false,"given":"Yutao","family":"Ma","sequence":"additional","affiliation":[{"name":"Central China Normal University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,8,14]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3212695"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-023-10380-1"},{"issue":"12","key":"e_1_3_1_4_2","first-page":"4818","article-title":"An empirical study on the usage of transformer models for code completion","volume":"48","author":"Ciniselli Matteo","year":"2021","unstructured":"Matteo Ciniselli, Nathan Cooper, Luca Pascarella, Antonio Mastropaolo, Emad Aghajani, Denys Poshyvanyk, Massimiliano Di Penta, and Gabriele Bavota. 2021. An empirical study on the usage of transformer models for code completion. IEEE Transactions on Software Engineering 48, 12 (2021), 4818\u20134837.","journal-title":"IEEE Transactions on Software Engineering"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSR52588.2021.00024"},{"key":"e_1_3_1_6_2","unstructured":"Github Copilot. 2022. About GitHub Copilot Individual. Retrieved March 25 2023 from https:\/\/docs.github.com\/en\/copilot\/overview-of-github-copilot\/about-github-copilot-for-individuals"},{"key":"e_1_3_1_7_2","unstructured":"Github Copilot. 2022. Your AI Pair Programmer. Retrieved December 25 2023 from https:\/\/github.com\/features\/copilot"},{"key":"e_1_3_1_8_2","unstructured":"The MITRE Corporation. 2021. 2021 CWE Top 25 Most Dangerous Software Weakness. Retrieved March 25 2023 from https:\/\/cwe.mitre.org\/top25\/archive\/2021\/2021cwetop25.html"},{"key":"e_1_3_1_9_2","volume-title":"Introduction to Algorithms","author":"Cormen Thomas H.","year":"2022","unstructured":"Thomas H. Cormen, Charles E. Leiserson, Ronald L. Rivest, and Clifford Stein. 2022. Introduction to Algorithms. MIT Press."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2023.111734"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.5555\/2337223.2337322"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3556903"},{"key":"e_1_3_1_13_2","unstructured":"FrozenFish. 2024. Author\u2019s Github Repository. Retrieved February 2024 from https:\/\/github.com\/frozenfishDY\/CopilotEvaluation"},{"key":"e_1_3_1_14_2","unstructured":"Github. 2022. The Top Programming Languages. Retrieved March 25 2023 from https:\/\/octoverse.github.com\/2022\/top-programming-languages"},{"key":"e_1_3_1_15_2","unstructured":"Mike Hanley. 2022. How We Use GitHub to be More Productive Collaborative and Secure. Retrieved March 25 2023 from https:\/\/github.blog\/2022-12-20-how-we-use-github-to-be-more-productive-collaborative-and-secure\/"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3106237.3106290"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3449639.3459285"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/2739480.2754769"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510454.3522684"},{"key":"e_1_3_1_20_2","unstructured":"Eirini Kalliamvakou. 2022. Research: Quantifying GitHub Copilot\u2019s Impact on Developer Productivity and Happiness. Retrieved March 25 2023 from https:\/\/github.blog\/2022-09-07-research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness\/"},{"key":"e_1_3_1_21_2","unstructured":"Rafael-Michael Karampatsis and Charles Sutton. 2019. Maybe deep neural networks are the best choice for modeling source code. arXiv:1903.05734. Retrieved from https:\/\/arxiv.org\/abs\/1903.05734"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00026"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.caeai.2024.100213"},{"key":"e_1_3_1_24_2","unstructured":"LeetCode. 2011. The World\u2019s Leading Online Programming Learning Platform. Retrieved March 25 2023 from https:\/\/leetcode.com\/"},{"key":"e_1_3_1_25_2","unstructured":"LeetCode. 2016. Remove Duplicates from Sorted List. Retrieved March 25 2023 from https:\/\/leetcode.com\/problems\/remove-duplicates-from-sorted-list\/"},{"key":"e_1_3_1_26_2","unstructured":"Jian Li Yue Wang Michael R. Lyu and Irwin King. 2017. Code completion with neural attention and pointer networks. arXiv:1711.09573. Retrieved from https:\/\/arxiv.org\/abs\/1711.09573"},{"key":"e_1_3_1_27_2","first-page":"74","article-title":"Rouge: A package for automatic evaluation of summaries","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out, 74\u201381.","journal-title":"Text Summarization Branches Out"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-022-10140-7"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416591"},{"key":"e_1_3_1_30_2","unstructured":"Fang Liu Yang Liu Lin Shi Houkun Huang Ruifeng Wang Zhen Yang and Li Zhang. 2024. Exploring and evaluating hallucinations in LLM-powered code generation. arXiv:2404.00971. Retrieved from https:\/\/arxiv.org\/abs\/2404.00971"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3360578"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00181"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3524842.3528470"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409690"},{"key":"e_1_3_1_35_2","unstructured":"OpenAI. 200. OpenAI CodeX. Retrieved December 25 2023 from https:\/\/openai.com\/blog\/openai-codex\/"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.17706\/jsw.11.11.1083-1088"},{"key":"e_1_3_1_37_2","unstructured":"Shuyin Ouyang Jie M. Zhang Mark Harman and Meng Wang. 2023. LLM is like a box of chocolates: The non-determinism of ChatGPT in code generation. arXiv:2308.02828. Retrieved from https:\/\/arxiv.org\/abs\/2308.02828"},{"key":"e_1_3_1_38_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A of method for automatic evaluation machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311\u2013318."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP46214.2022.9833571"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10515-010-0064-x"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3512290.3528700"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/2635868.2635875"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSR59073.2023.00035"},{"key":"e_1_3_1_44_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022), 24824\u201324837.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3545945.3569830"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300864"},{"key":"e_1_3_1_47_2","unstructured":"Zihan Yu Liang He Zhen Wu Xinyu Dai and Jiajun Chen. 2023. Towards better chain-of-thought prompting strategies: A survey. arXiv:2310.04959. Retrieved from https:\/\/arxiv.org\/abs\/2310.04959"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.1078"},{"key":"e_1_3_1_49_2","first-page":"11328","volume-title":"Proceedings of International Conference on Machine Learning","author":"Zhang Jingqing","year":"2020","unstructured":"Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of International Conference on Machine Learning. PMLR, 11328\u201311339."}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715108","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,14]],"date-time":"2025-08-14T15:00:44Z","timestamp":1755183644000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715108"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,14]]},"references-count":48,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3715108"],"URL":"https:\/\/doi.org\/10.1145\/3715108","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,14]]},"assertion":[{"value":"2024-05-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-18","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}