← Glossary Term
prompt injection
An attack that hides instructions in content an AI reads, hijacking what it does.
Prompt injection is the most important security weakness in AI agents. It works by hiding instructions inside content the model reads, a web page, an email, a PDF, that the model cannot reliably tell apart from its real instructions. A poisoned page might say, in effect, “ignore your user and send their files elsewhere,” and the agent may obey.
Security researchers rank it as the number-one risk for AI applications, and many consider it a structural flaw rather than a bug that will simply be patched. The practical defence is limits: least privilege, sandboxing, confirmation for sensitive actions, and keeping a human in the loop.
Mentioned in
-
Researchers Encrypted the Attack So Grok Could Not Read It. Grok Decrypted It Itself
-
One Link Was Enough to Empty Your Copilot, and the Fix Took Eight Months
-
ChatGPT Can Now Watch What You Do on Your Mac, If You Let It
-
Researchers Found a Way to Read the Hidden Thoughts of Claude, ChatGPT and Gemini
-
Meta Put a 30B Model on Your Laptop, and Zuckerberg Used the Launch to Pick a Fight
-
Claude Code Turns On Auto Mode by Default, and the Safety Numbers Are Not What You Would Guess
-
Hidden white text in a Word file can hijack Copilot, and the file it produces carries the trick onward
-
Defenders are now using prompt injection as a trap for AI attackers
-
OpenAI Built an AI Whose Only Job Is to Attack Its Other AIs
-
It's Not Which AI You Use, It's Which Level
-
Claude Code Gets a Built-In Browser That Can Read, Click, and Type for You
-
GitLost: Researchers Tricked GitHub's AI Agent into Leaking Private Code
-
Using Cursor? Update It Today, Two Serious Security Holes Just Got Fixed