MARATTO

article

Unlocking LLM Security: Automated Penetration Testing and Source Code Review Based on Vulnerability Analysis and NLP

Abstract

This paper introduces an innovative framework designed to bolster the security of large language models (LLMs) and source code through automated penetration testing and vulnerability analysis. By integrating natural language processing (NLP) and machine learning, the proposed system overcomes the shortcomings of conventional security tools, offering a unified solution for testing LLMs and reviewing source code across multiple programming languages. The architecture employs API-based testing to probe LLMs for vulnerabilities such as prompt injections and data leaks, while leveraging static and dynamic analysis to detect coding flaws like SQL injection and buffer overflows. Experimental evaluations confirm the system's efficacy in identifying security risks, with features including real-time monitoring and an intuitive interface enhancing its accessibility to developers of varying expertise. Despite its promising accuracy and efficiency, challenges such as false positives and reliance on dataset quality are noted, suggesting directions for future refinement. This work advances AIdriven cybersecurity by delivering a scalable, adaptable toolset for securing modern software systems and intelligent models, contributing significantly to the evolving landscape of digital security.

Research topics

  • Web Application Security Vulnerabilities
  • Information and Cyber Security
  • Digital and Cyber Forensics

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/itc-egypt66095.2025.11186635

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.