Partners Running an agency? Partner with us
The Graylemon Journal Build logs, teardowns and method, by email. Subscribe All resources
Follow the studio LinkedIn Instagram
Partner with usBook a call (opens in a new tab)
Glossary

Guardrails

Guardrails are the rules and checks around an AI model that keep its behaviour within safe, useful limits: what it may answer, what it must refuse, what data it may touch and when it hands over to a person. They sit outside the model, so they hold even when the model gets something wrong.

What it means

A model's own training can be steered by clever inputs. Guardrails built around it, such as checks on what goes in, limits on what it can do, and checks on what comes out, don't depend on the model behaving well.

Guardrails need testing like anything else: inputs designed to get round them, attempts to pull out data the system shouldn't share, and requests it should refuse.

How we use it

Our method Before release, we test AI products against inputs designed to mislead them and attempts to extract data they shouldn't share, alongside the evaluation set. It's part of how we know a product works before everyone sees it.

Go further

Releasing in steps, so failures surface before they reach everyone.

Graduated Launch

Published 28 September 2026. All terms

Search Graylemon

↑↓ Move↵ OpenEsc Close