Guardrails
Guardrails are the rules and checks around an AI model that keep its behaviour within safe, useful limits: what it may answer, what it must refuse, what data it may touch and when it hands over to a person. They sit outside the model, so they hold even when the model gets something wrong.
What it means
A model's own training can be steered by clever inputs. Guardrails built around it, such as checks on what goes in, limits on what it can do, and checks on what comes out, don't depend on the model behaving well.
Guardrails need testing like anything else: inputs designed to get round them, attempts to pull out data the system shouldn't share, and requests it should refuse.
How we use it
Our method Before release, we test AI products against inputs designed to mislead them and attempts to extract data they shouldn't share, alongside the evaluation set. It's part of how we know a product works before everyone sees it.
Related terms
Published 28 September 2026. All terms