Skip to content

GlossaryFloor 1 · The Modela solid block on its own: the prediction machineFloor 1 · The Model

multimodal

No. 058 · v2026-08FR: multimodal

A multimodal model handles several kinds of input in the same computation: text, images, sound, without first converting them into a single format. Like a person looking at a chart while listening to the commentary on it, rather than reading its written description.

What it is not

Multimodal is not a chain of specialised models. Recognising the text of an image with a tool, then giving that text to a language model, is not multimodal: it is a chain, and everything the recognition lost, the layout, the colours, a gesture in a photograph, is lost for good. Nor is it a synonym for generative AI: a model can be multimodal on input and produce nothing but text, which is the most common case.

In depth

The mechanism

The mechanism is the same as for text, with one detail. An image is cut into squares, a sound into slices, and each piece is converted into a numerical representation of the same nature as that of a fragment of text. From there, attention treats all these elements alike: nothing in the architecture distinguishes a square of image from a piece of a word. That is what allows a written question to bear on a precise area of an image, with no intermediate pass.

Input and output

Input and output have to be distinguished, which the word covers indiscriminately. Many models read images and write nothing but text: they are called multimodal, rightly, but they draw nothing. Others produce images or sound, and then come under generation in those fields. A system that seems to do everything generally combines several models behind a single interface, which is a matter of harness and not of model.

The practical consequences

The practical consequences are concrete and often underestimated. A scanned document, a screenshot, a hand-annotated diagram become direct inputs, which removes whole chains of conversion. In return, images are expensive: a single one can consume as many tokens as several pages of text, and the bill for document processing shows it. Finally, the quality of reading remains uneven on dense tables and handwriting, which is worth checking on your own documents before committing.

Relations where the neighbours live

Check 3 questions · click your answer

Level 1 · Recognise

You convert a PDF into text with a recognition tool, then give that text to the model. Is this multimodal?

Level 2 · Distinguish

Does a multimodal model necessarily produce images?

Level 2 · Distinguish

What is the main economic side effect of image processing by a model?

No. 058 · v2026-08 · first written in · editorial responsibility Anthony Capirchio

Lexigraph, "Multimodal", v2026-08, https://www.lexigraph.org/en/multimodal/, CC BY 4.0.

Report

What goes with your message

Entry · Multimodal
No. 058 · v2026-08 · /en/multimodal

What is this about
0 / 600

It is used to reply to you, and for nothing else. What is recorded