Multimodal

A multimodal model can process or generate more than one type of data — such as text together with images, audio, or video — within a single system. Multimodality lets a model reason across modalities, for example answering questions about a picture or describing a chart alongside text.

Definition

Multimodal

Traditional language models handle text only. Multimodal models extend the same architecture to accept or produce additional modalities, typically by encoding images, audio, or video into representations the model can attend to jointly with text tokens. This enables tasks like visual question answering, document understanding, and image captioning.

Multimodal capability underpins vision-language models and audio-enabled assistants. On the generation side, some systems also output images or speech. The core benefit is unified reasoning: the model can connect what it reads with what it sees or hears in one context.

Frequently asked questions

What is Multimodal?

A multimodal model can process or generate more than one type of data — such as text together with images, audio, or video — within a single system. Multimodality lets a model reason across modalities, for example answering questions about a picture or describing a chart alongside text.

Sources

Related terms

← Full AI & prompt engineering glossary

Put Multimodal to work with Prompeteer, the Agentic Contextual AI Platform →