Taxonomy of Prompt Injection Attacks and Analysis of Defense Mechanisms in Large Language Model-Based Chatbots

Telman Yusifov, Aziz Aghayev, Laman Hasanzada

InterConf · 2026

Large Language Models have become critical infrastructure across healthcare, finance, and software engineering, yet their instruction-following design creates a structural class of security vulnerabilities known as prompt injection. This paper presents a comprehensive taxonomy of prompt injection attacks targeting LLM-based chatbots, organizes the threat landscape across seven categories covering both direct and indirect injection vectors, and evaluates defense mechanisms at three operational layers. The theoretical framework is validated through an empirical benchmark (PI-Bench) testing 210 prompt injection payloads across seven attack categories against five open-weight LLM models in a controlled local environment.

Empirical findings demonstrate attack success rates ranging from 10.5% (LLaMA 3.1 8B) to 83.8% (Mistral 7B), confirming that alignment training quality is the primary determinant of injection resistance rather than model size. Across all models, Jailbreak/Roleplay attacks proved most effective at 34.7% aggregate ASR, while Gradient-Based attacks were least effective at 22.7%. The literature review documents attack success rates reaching 87.2% against state-of-the-art proprietary models and an 86.1% success rate for indirect injection against production applications.

No single defensive measure provides adequate coverage, and a defense-in-depth architecture is necessary for robust protection.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.