Articles
| Open Access |
https://doi.org/10.55640/ijdsml-06-01-04
Trustworthy and Secure LLM-Assisted Code Generation for Enterprise Software Development
Sai Shiva Reddy Kongari , Department of Information Technology, University of the Cumberlands, Williamsburg, KY, 40769, USA Saurav Kant Kumar , Department of Information Technology, University of the Cumberlands, Williamsburg, KY, 40769, USA Abhishek Kumar , Department of Information Technology, University of the Cumberlands, Williamsburg, KY, 40769, USAAbstract
Large Language Models (LLMs) are increasingly embedded in enterprise software development workflows for source-code generation, debugging, documentation, test creation, refactoring, and design assistance. While these systems improve productivity and reduce repetitive programming effort, their deployment in enterprise environments introduces significant risks related to insecure code suggestions, hallucinated functionality, non-compliant logic, weak explainability, dependency misuse, and over-reliance by developers. This research paper investigates how LLM-assisted code generation can be made secure, reliable, and trustworthy through a structured governance-oriented framework that integrates prompt constraints, retrieval support, policy-based filtering, automated validation, secure coding guardrails, and human-in-the-loop review.
Drawing only on the provided literature, the study positions enterprise code generation within the broader development of LLMs, code-specialized models, hardware generation models, design automation agents, and AI-assisted engineering workflows. Prior work on GPT-4, StarCoder, CodeGen2, Magicoder, VerilogEval, RTLLM, OpenLLM-RTL, ChipNeMo, ChipGPT, AutoChip, and secure hardware generation demonstrates that LLMs can generate technically meaningful artifacts but remain vulnerable to correctness gaps, evaluation instability, insufficient domain grounding, and security-sensitive errors (Achiam, 2023; Li, 2023; Nijkamp et al., 2023; Wei et al., 2023; Liu et al., 2023; Lu et al., 2023; Thakur et al., 2023). This paper extends those insights to enterprise software engineering and proposes a Trustworthy Secure Code Generation Framework organized around six layers: task specification, retrieval-augmented contextualization, constrained generation, security-policy enforcement, validation and testing, and accountable human approval.
The findings suggest that trustworthiness in LLM-assisted enterprise code generation cannot be achieved by model scale alone. Instead, it requires layered controls that combine technical verification, organizational policy, developer accountability, and continuous feedback. The proposed framework reduces the probability of insecure code propagation, improves explainability of generated outputs, supports compliance-oriented review, and transforms LLMs from autonomous code producers into controlled development assistants. The paper concludes that enterprise adoption of LLM-assisted coding should prioritize measurable trust, secure-by-design workflows, traceable decision-making, and hybrid human-AI governance.
Keywords
Large Language Models, Secure Code Generation, Enterprise Software Development, Trustworthy AI, Human-in-the-Loop Review, Retrieval-Augmented Generation, Software Security, Code Validation, AI Governance, Generative AI
References
J. Achiam, “GPT-4 technical report,” 2023, arXiv:2303.08774.
B. Ahmad, S. Thakur, B. Tan, R. Karri, and H. Pearce, “Fixing hardware security bugs with large language models,” 2023, arXiv:2302.01215.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in Proc. NeurIPs, 2015, pp. 1–9.
J. Blocklove, S. Garg, R. Karri, and H. Pearce, “Chip-chat: Challenges and opportunities in conversational hardware design,” 2023, arXiv:2305.13243.
K. Chang, “ChipGPT: How far are we from natural language hardware design,” 2023, arXiv:2305.14019.
K. Chang, “Data is all you need: Finetuning LLMs for chip design via an automated design-data augmentation framework,” 2024, arXiv:2403.11202.
L. Chen, “The dawn of AI-native EDA: Promises and challenges of large circuit models,” 2024, arXiv:2403.07257.
Y. Fu, “GPT4AIGChip: Towards next-generation AI accelerator design automation via large language models,” 2023, arXiv:2309.10730.
E. Goh, M. Xiang, I. Wey, and T. H. Teo, “From English to ASIC: Hardware implementation with large language model,” 2024, arXiv:2403.07039.
Z. He, “ChatEDA: A large language model powered autonomous agent for EDA,” in Proc. MLCAD Workshop, 2023, pp. 1–14.
Q. Jiang, “Mistral 7B,” 2023, arXiv:2310.06825.
R. Kande, “LLM-assisted generation of hardware assertions,” 2023, arXiv:2306.14027.
Z. Liang, “Unleashing the potential of LLMs for quantum computing: A study in quantum architecture design,” 2023, arXiv:2307.08191.
Y. Liu, P. Liu, D. Radev, and G. Neubig, “BRIO: Bringing order to abstractive summarization,” 2022, arXiv:2203.16804.
M. Li, W. Fang, Q. Zhang, and Z. Xie, “SpecLLM: Exploring generation and review of VLSI design specification with large language model,” 2024, arXiv:2401.13266.
R. Li, “StarCoder: May the source be with you!,” 2023, arXiv:2305.06161.
Z. Pei, H.-L. Zhen, M. Yuan, Y. Huang, and B. Yu, “BetterV: Controlled Verilog generation with discriminative guidance,” 2024, arXiv:2402.03375.
M. Liu, “ChipNeMo: Domain-adapted LLMs for chip design,” 2023, arXiv:2311.00176.
M. Liu, N. Pinckney, B. Khailany, and H. Ren, “VerilogEval: Evaluating large language models for Verilog code generation,” 2023, arXiv:2309.07544.
S. Liu, Y. Lu, W. Fang, M. Li, and Z. Xie, “OpenLLM-RTL: Open dataset and benchmark for LLM-aided design RTL generation,” in Proc. IEEE/ACM Int. Conf. Comput. Aided Design (ICCAD), 2024, pp. 1–9.
Y. Lu, S. Liu, Q. Zhang, and Z. Xie, “RTLLM: An open-source benchmark for design RTL generation with large language model,” 2023, arXiv:2308.05345.
M. Nair, R. Sadhukhan, and D. Mukhopadhyay, “Generating secure hardware using ChatgPT resistant to CWEs,” Cryptol. ePrint Arch., IACR, Bellevue, WA, USA, Rep. 2023/212, 2023.
E. Nijkamp, H. Hayashi, C. Xiong, S. Savarese, and Y. Zhou, “CodeGen2: Lessons for training LLMs on programming and natural languages,” 2023, arXiv:2305.02309.
M. Rapp, “MLCAD: A survey of research in machine learning for CAD keynote paper,” IEEE Trans. Comput.-Aided Design Integr. Circuits Syst., vol. 41, no. 10, pp. 3162–3181, Oct. 2022.
C. Shaib, J. Barrow, J. Sun, A. F. Siu, B. C. Wallace, and A. Nenkova, “Standardizing the measurement of text diversity: A tool and a comparative analysis of scores,” 2024, arXiv:2403.00553.
S. Thakur, “Benchmarking large language models for automated Verilog RTL code generation,” in Proc. DATE, 2023, pp. 1–7.
S. Thakur, J. Blocklove, H. Pearce, B. Tan, S. Garg, and R. Karri, “AutoChip: Automating HDL generation using LLM feedback,” 2023, arXiv:2311.04887.
Y. Wei, Z. Wang, J. Liu, Y. Ding, and L. Zhang, “Magicoder: Source code is all you need,” 2023, arXiv:2312.02120.
Z. Yan, Y. Qin, X. S. Hu, and Y. Shi, “On the viability of using LLMs for SW/HW co-design: An example in designing CiM DNN accelerators,” 2023, arXiv:2306.06923.
Y. Zhang, Z. Yu, Y. Fu, C. Wan, and Y. C. Lin, “MG-Verilog: Multi-grained dataset towards enhanced LLM-assisted Verilog generation,” 2024, arXiv:2407.01910.
C. Zhao, “Generative AI-enabled wireless communications for robust low-altitude economy networking,” 2025, arXiv:2502.18118.
X. Gao, X. Zhu, and L. Zhai, “AoI-sensitive data collection in multi-UAV-assisted wireless sensor networks,” IEEE Trans. Wireless Commun., vol. 22, no. 8, pp. 5185–5197, Aug. 2023.
H. Hu, K. Xiong, G. Qu, Q. Ni, P. Fan, and K. B. Letaief, “AoI-minimal trajectory planning and data collection in UAV-assisted wireless powered IoT networks,” IEEE Internet Things J., vol. 8, no. 2, pp. 1211–1223, Jan. 2021.
S. Kaul, R. Yates, and M. Gruteser, “Real-time status: How often should one update?,” in Proc. IEEE INFOCOM, Mar. 2012, pp. 2731–2735.
H. Li, M. Xiao, K. Wang, D. I. Kim, and M. Debbah, “Large language model based multi-objective optimization for integrated sensing and communications in UAV networks,” IEEE Wireless Commun. Lett., vol. 14, no. 4, pp. 979–983, Apr. 2025.
F. Liu, “Evolution of heuristics: Towards efficient automatic algorithm design using large language model,” in Proc. Int. Conf. Mach. Learn., 2024, pp. 32201–32223.
J. Liu, X. Wang, B. Bai, and H. Dai, “Age-optimal trajectory planning for UAV-assisted data collection,” in Proc. IEEE INFOCOM WKSHPS, Apr. 2018, pp. 553–558.
J. Liu, P. Tong, X. Wang, B. Bai, and H. Dai, “UAV-aided data collection for information freshness in wireless sensor networks,” IEEE Trans. Wireless Commun., vol. 20, no. 4, pp. 2368–2382, Apr. 2021.
T. Wu, “A novel AI-based framework for AoI-optimal trajectory planning in UAV-assisted wireless sensor networks,” IEEE Trans. Wireless Commun., vol. 21, no. 4, pp. 2462–2475, Apr. 2022.
R. Zhang, “Generative AI agents with large language model for satellite networks via a mixture of experts transmission,” IEEE J. Sel. Areas Commun., vol. 42, no. 12, pp. 3581–3596, Dec. 2024.
R. Zhang, “Interactive AI with retrieval-augmented generation for next generation networking,” IEEE Netw., vol. 38, no. 6, pp. 414–424, Nov. 2024.
X. Zhang, “Beyond the cloud: Edge inference for generative large language models in wireless networks,” IEEE Trans. Wireless Commun., vol. 24, no. 1, pp. 643–658, Jan. 2025.
B. Zhu, E. Bedeer, H. H. Nguyen, R. Barton, and Z. Gao, “UAV trajectory planning for AoI-minimal data collection in UAV-aided IoT networks by transformer,” IEEE Trans. Wireless Commun., vol. 22, no. 2, pp. 1343–1358, Feb. 2023.
Article Statistics
Downloads
Copyright License
Copyright (c) 2026 Sai Shiva Reddy Kongari, Saurav Kant Kumar, Abhishek Kumar

This work is licensed under a Creative Commons Attribution 4.0 International License.