Skip to content

GlossaryFloor 2 · The Harnessthe block and its bolted-on plates: what gets added to itFloor 2 · The Harness

moderation

No. 044 · v2026-08FR: modération

Moderation is the sorting of the content that enters and leaves an AI system: what you refuse to process, what you refuse to let through. Like the check at the door of a hall: you look at what crosses the threshold, in both directions.

What it is not

Moderation is not the polite refusal an assistant sometimes gives you. That refusal most often comes from the model itself, shaped during its training to decline certain requests; moderation is an outside device, which inspects the text before it reaches the model or before it reaches you, and which blocks without the model having any say. Nor is it a guardrail: a guardrail limits what the system can do, moderation sorts what it can read and say. The two complete each other, neither replaces the other.

In depth

Two moments

Moderation is exercised at two distinct moments, and confusing them leaves holes. On input, it examines what is about to be submitted: manifestly unlawful requests, content you do not want processed, personal data that has no business in a text sent outside. On output, it examines what is about to be returned: abusive language, dangerous advice, information that was not meant to reach this recipient. Depending on the case, the check rests on explicit rules, on models specialised in classification, or on both, and it always lives in the harness: it is a sorting applied to text, not a property of the model that writes.

Arbitrating between two errors

All moderation arbitrates between two errors, and one is reduced only at the price of the other. Letting through what should have been blocked exposes you to real harm; blocking what should have gone through produces an unusable system, whose incomprehensible refusals push teams to work elsewhere, beyond any oversight. The setting is therefore not technical but editorial: it depends on the audience, on the field and on what you accept risking, and a health service does not keep the same threshold as an advertising copy tool. This arbitration is documented, failing which no one in the organisation knows what is filtered or in the name of what.

The traps

The first trap is to believe moderation neutral: what is judged unacceptable depends on a language, a culture and a period, and a filter tuned to one use handles the others badly, starting with quotation, irony and the ordinary vocabulary of certain trades. The second is circumvention, since a text can be rephrased, translated, cut up or hidden in a document read along the way: a single filter is crossed far more easily than a chain of independent checks. The third is to expect everything of the filter, when solid moderation prevents no disastrous action: it judges content, not effects, and effects come under guardrails. The last is forgetting redress, because a block with no explanation and no possible challenge is an opaque decision, which recent texts no longer tolerate.

Relations where the neighbours live

Check 3 questions · click your answer

Level 1 · Recognise

Where does moderation apply in an AI system?

Level 2 · Distinguish

Your system refuses to write an insult but agrees to delete a customer file on simple request. What is missing?

Level 2 · Distinguish

Your lawyers complain that the assistant refuses to analyse clauses because they mention violent events. What needs adjusting?

No. 044 · v2026-08 · first written in · editorial responsibility Anthony Capirchio

Lexigraph, "Moderation", v2026-08, https://www.lexigraph.org/en/moderation/, CC BY 4.0.

Report

What goes with your message

Entry · Moderation
No. 044 · v2026-08 · /en/moderation

What is this about
0 / 600

It is used to reply to you, and for nothing else. What is recorded