Generative AI applications bring new attack surfaces. A large language model (LLM) follows instructions written in natural language, so an attacker can try to change its behavior with words instead of code. Two terms matter most. A direct prompt injection, often called a jailbreak, is a user prompt crafted to make the model ignore its system instructions or safety rules, for example asking it to role-play a character with no restrictions. An indirect prompt injection hides instructions inside content the model processes, such as a web page, email or document retrieved for grounding, so that the model follows the attacker's hidden text when summarizing it. Other risks include sensitive data leakage, wallet abuse (running up token costs) and misuse of connected tools.
Keep reading for free
Create a free StudyToCert account to read the rest of this lesson: 4 more sections, 4 key terms, a real-world example, an exam tip and self-check questions. Every lesson, lab and practice test is free with an account.