Katalog Plus
Bibliothek der Frankfurt UAS
Bald neuer Katalog: sichern Sie sich schon vorab Ihre persönlichen Merklisten im Nutzerkonto: Anleitung.
Dieses Ergebnis aus BASE kann Gästen nicht angezeigt werden.  Login für vollen Zugriff.

Less Is More: DocString Compression in Code Generation

Title: Less Is More: DocString Compression in Code Generation
Authors: Yang, Guang; Zhou, Yu; Cheng, Wei; Zhang, Xiangyu; Chen, Xiang; Zhuo, Terry Yue; Liu, Ke; Zhou, Xin; Lo, David; Chen, Taolue
Contributors: National Natural Science Foundation of China; Short-term Visiting Program of Nanjing University of Aeronautics and Astronautics for Ph.D. Students Abroad; High Performance Computing Platform of Nanjing University of Aeronautics and Astronautics, and the Collaborative Innovation Center of Novel Software Technology and Industrialization; State Key Laboratory of Novel Software Technology, Nanjing University; A*STAR under its 2nd CSIRO and A*STAR: Research-Industry (2+2) Partnership Program
Source: ACM Transactions on Software Engineering and Methodology ; volume 35, issue 2, page 1-31 ; ISSN 1049-331X 1557-7392
Publisher Information: Association for Computing Machinery (ACM)
Publication Year: 2026
Description: The widespread use of Large Language Models (LLMs) in software engineering has intensified the need for improved model and resource efficiency. In particular, for neural code generation, LLMs are used to translate function/method signature and DocString to executable code. DocStrings, which capture user requirements for the code and are typically used as the prompt for LLMs, often contain redundant information. Recent advancements in prompt compression have shown promising results in Natural Language Processing (NLP), but their applicability to code generation remains uncertain. Our empirical study shows that the state-of-the-art prompt compression methods achieve only about 10% reduction, as further reductions would cause significant performance degradation. In our study, we propose a novel compression method, ShortenDoc, dedicated to DocString compression for code generation. Our experiments on six code generation datasets, five open source LLMs (1B to 10B parameters), and one closed-source LLM GPT-4o confirm that ShortenDoc achieves 25–40% compression while preserving the quality of generated code, outperforming other baseline methods at similar compression levels. The benefit of this method is to improve efficiency and reduce the token processing cost while maintaining the quality of the generated code, especially when calling third-party APIs.
Document Type: article in journal/newspaper
Language: English
DOI: 10.1145/3735636
Availability: https://doi.org/10.1145/3735636; https://dl.acm.org/doi/pdf/10.1145/3735636
Accession Number: edsbas.E7C96DBE
Database: BASE