Skip to content

GlossaryFloor 3 · The Agentthe block caught in a loop: it starts again until it gets thereFloor 3 · The Agent

observability

No. 026 · v2026-08FR: observabilité

Observability is the ability to understand what a system is doing from what it records, without having to open it up again. Like a car dashboard: the needles repair nothing, they tell you where to look.

What it is not

Observability is not a dashboard. A dashboard answers questions decided in advance, whereas observability has to make it possible to answer a question nobody anticipated, from what has already been recorded. Nor is it the trace itself: the trace is the material, observability is what you manage to draw from it. A system that records everything and lets you find nothing is not observable.

In depth

Three materials

Observability rests on three distinct materials: traces, which recount one execution from end to end; aggregated metrics, which give trends across all executions; and notable events, which signal a break. On a conventional system, one mostly watches availability and speed, and that is largely enough. On a system built around a model, those indicators stay green while quality collapses, since the system answers fast and always, including nonsense. Here observability must therefore bear on content and on behaviour, not only on the machine.

The indicators that speak

On an agent, a few indicators turn out to say more than the others. The rate of goals reached, measured against a criterion checkable outside the model, says what a technical error rate never will. The distribution of turn counts flags loops that are treading water well before anyone complains. Cost per execution and the rate of escalation to a person complete the picture, the first because it drifts silently, the second because it puts a figure on the trust actually placed in the system.

The three traps

The first trap is to measure what is easy rather than what matters: response time is measured effortlessly, the correctness of an open-ended answer calls for a criterion, a judge and discipline. The second is the alert that cannot be set, because a mediocre answer crosses no threshold and wakes nobody, which forces one to run known cases continuously rather than wait for a failure. The third is observability that never comes back down to the people concerned, when only the technical teams see the curves while the business side alone can recognise a wrong answer. Observability is worth something only on condition that it leads to a decision: correcting an instruction, restricting a tool, tightening the autonomy setting.

Relations where the neighbours live

Check 3 questions · click your answer

Level 1 · Recognise

All the technical indicators of an internal assistant are green, and users complain about wrong answers. What is missing?

Level 2 · Distinguish

What is the difference between observability and evals?

Level 2 · Distinguish

Which indicator says the most about the real health of an agent in production?

No. 026 · v2026-08 · first written in · editorial responsibility Anthony Capirchio

Lexigraph, "Observability", v2026-08, https://www.lexigraph.org/en/observability/, CC BY 4.0.

Report

What goes with your message

Entry · Observability
No. 026 · v2026-08 · /en/observability

What is this about
0 / 600

It is used to reply to you, and for nothing else. What is recorded