Multimodal Model
A model that can process or generate multiple kinds of data, such as text, images, audio, video, or code.
Plain English
It can understand more than words.
Example
A model reads a screenshot, explains the UI bug, and suggests CSS changes.
Why it matters
Multimodal capability makes agents useful in browsers, design tools, documents, and physical-world workflows.