GlossaryFloor 1 · The Modela solid block on its own: the prediction machineFloor 1 · The Model
alignment
No. 060 · v2026-08FR: alignementAlignment is the work that makes a model behave as one wishes: following an instruction, refusing certain requests, adopting a tone. Like the induction that follows a hire, which does not change what a person knows but what they do with it.
What it is not
Alignment is not moderation: moderation filters inputs and outputs from the outside, at run time; alignment modifies the model itself, before it is served. Nor is it guardrails, which are limits placed by the harness around a system. And it is not a guarantee: an aligned model can be circumvented, and its behaviour varies with the wording of the requests.
In depth
The two stages
The work is done in two stages after pre-training. First, learning on dialogues written for the purpose, which teaches the very form of the exchange: a question calls for an answer, an instruction is followed, a dangerous request is declined. Then learning from preferences: the model is shown pairs of answers ranked by people or by another model, and it is adjusted so that it produces more often those that were preferred. It is this second stage that gives the model its apparent personality.
The characteristic flaws
It also produces characteristic flaws, which follow directly from the method. A model trained to produce answers that are appreciated learns to be agreeable, which is not the same thing as being accurate: it tends to approve of the user, to back down from a correct answer when it is challenged, and to prefer a confident answer to an admission of ignorance. These behaviours are not bugs but the logical consequence of an optimisation criterion that measures satisfaction rather than truth.
The underlying difficulty
The underlying difficulty is that there exists no specification of what good behaviour is. Refusals are a permanent compromise: too strict, they make the model useless on legitimate subjects, in medicine or in computer security; too loose, they expose. That slider is an editorial choice made by the model’s publisher, rarely documented, and it differs from one model to another. For an organisation, this means testing the behaviour oneself on one’s own cases, because no datasheet describes it.
Relations where the neighbours live
Check 3 questions · click your answer
Level 1 · Recognise
At what point does alignment take place?
Level 2 · Distinguish
You challenge an answer that was nonetheless accurate, and the model backs down. Where does this behaviour come from?
Level 2 · Distinguish
Two models refuse different requests on the same subject. What should be concluded from this?
Try it 2 practices
Concrete things to try where this term comes up, in ten minutes.
Who works with this 2 roles
The roles for which this term is part of the ordinary work.
Lexigraph, "Alignment", v2026-08, https://www.lexigraph.org/en/alignment/, CC BY 4.0.