Prompt injection
Prompt injection is an attack in which text given to an AI model, such as a message, a web page or a document, contains instructions that try to override what the model was told to do. Any product that feeds outside text to a model has to assume some of it will try.
What it means
The risk grows with what the model can do. A model that only writes text can be tricked into saying something embarrassing; one that can send messages, change records or spend money can be tricked into doing it.
Defences work in layers: treat outside text as data rather than instructions, limit what the model may do without approval, check its output before it acts, and test with inputs designed to break the rules.
How we use it
Our method Our own tools treat whatever a visitor types as data to read, never as instructions to follow, and our product builds are tested with inputs designed to mislead the model before anyone else sees them.
Related terms
Published 28 September 2026. All terms