Multimodal
A multimodal model can process or generate more than one type of data — such as text together with images, audio, or video — within a single system. Multimodality lets a model reason across modalities, for example answering questions about a picture or describing a chart alongside text.
Definition
MultimodalTraditional language models handle text only. Multimodal models extend the same architecture to accept or produce additional modalities, typically by encoding images, audio, or video into representations the model can attend to jointly with text tokens. This enables tasks like visual question answering, document understanding, and image captioning.
Multimodal capability underpins vision-language models and audio-enabled assistants. On the generation side, some systems also output images or speech. The core benefit is unified reasoning: the model can connect what it reads with what it sees or hears in one context.
Frequently asked questions
What is Multimodal?
A multimodal model can process or generate more than one type of data — such as text together with images, audio, or video — within a single system. Multimodality lets a model reason across modalities, for example answering questions about a picture or describing a chart alongside text.
Sources
Related terms
← Full AI & prompt engineering glossary
Put Multimodal to work with Prompeteer, the Agentic Contextual AI Platform →