GlossaryFloor 2 · The Harnessthe block and its bolted-on plates: what gets added to itFloor 2 · The Harness
guardrails
No. 041 · v2026-08FR: garde-fousGuardrails are the limits placed around an AI system so that its errors stay harmless: what it cannot trigger on its own, what a person has to approve. Like the railing on a balcony: it does not stop you leaning out, it stops the fall.
What it is not
Guardrails are not a content filter, and they are not an instruction written into the standing instructions either. Filtering what gets said is moderation; guardrails decide what can happen: which action is allowed, over what perimeter, up to what threshold, with what human approval. A sentence politely asking the model not to break anything is not a guardrail, since it influences without constraining. A guardrail is a control belonging to the harness, outside the model, which still holds when the model is wrong or lets itself be manipulated.
In depth
The three places
Guardrails sit in three places, and confusing those places loses the greater part of the benefit. Upstream, you limit what comes in: which sources are readable, which requests are admissible, which data must never join the context. At the moment of acting, you limit what the system can trigger: authorised tools, rights reduced to the strict minimum, ceilings, human confirmation for anything irreversible. Downstream, you check what comes out before it reaches a person or another piece of software: expected shape, sources present, absence of information that was not meant to circulate.
The principle
The principle that governs them fits in one sentence: a guardrail assumes that the model has failed. The point is not to make it wiser, but to make sure that a false, hijacked or absurd output produces no lasting damage. That is why an effective guardrail lives outside the text: a check written into the harness applies whatever happens, whereas an instruction inserted into the context shares the fate of everything else there and can be contradicted by a document read along the way. The right question is therefore never whether the model will obey, but what happens on the day it does not.
The traps
The first trap is excess: limits that are too tight produce a system that refuses everything, which teams get around by going to work elsewhere, out of any supervision. The second is the decorative guardrail, that confirmation request nobody pays attention to any more after the hundredth time: an approval you cannot refuse is not one. The third is believing that they are written once and for all, whereas they are chosen according to the damage feared, tested like the rest of the system, and revisited every time the perimeter of action widens. The last is confusing them with regulatory compliance, which imposes obligations of result without saying how to restrain a particular tool.
Relations where the neighbours live
Check 3 questions · click your answer
Level 1 · Recognise
An assistant asks for your confirmation before sending a message to a customer. What is this arrangement called?
Level 2 · Distinguish
Which measure still protects when the model lets itself be manipulated by a booby-trapped document?
Level 2 · Distinguish
A system refuses to write an insulting text and also refuses a transfer above a certain amount. Is this the same arrangement?
Lexigraph, "Guardrails", v2026-08, https://www.lexigraph.org/en/guardrails/, CC BY 4.0.