A multimodal model is an AI model that can work with more than one kind of input, such as text, images and audio, and sometimes video too. Show one a photo of a dog, play it a bark and ask 'What is this?', and it can combine all three to answer: 'Dog.'
One model, many kinds of input
Each kind of data is called a modality. Most AI models used to handle just one:
- Text models read and write words.
- Vision models read pictures: they classify photos or spot objects.
- Speech models hear sounds: they turn speech into text, or text into speech.
Each one works well in its own world, but knows nothing about the others. A text model can't look at your screenshot, and a vision model can't read your question about it.
A multimodal model pulls them together. It takes in several modalities at once, connects them, and reasons across them. That matters because real problems are rarely just text: a bug report comes with a screenshot, a lecture is speech plus slides, a medical case is scans plus notes.
How it works: encoders and a shared space
A model can only do maths on numbers, so the first job is turning every kind of input into numbers it can compare.
Step 1: encoders turn data into numbers
Each modality has its own encoder: a part of the model that turns raw data into lists of numbers called embeddings.
- An image encoder typically cuts the picture into small square patches and turns each patch into an embedding. A 224 by 224 pixel image cut into 16 by 16 patches gives 14 × 14 = 196 patches, so the model 'sees' it as 196 embeddings.
- An audio encoder usually turns the sound into a spectrogram, a picture of which frequencies are loud at each moment, and encodes short slices of it.
- A text encoder splits the words into tokens and turns each token into an embedding.
Step 2: map everything into a shared space
On their own, the three encoders would produce numbers that mean nothing to each other. The key step is mapping them into a shared space: one common set of coordinates where similar meanings land close together, whatever form they arrived in.
Here is the whole path for the dog example:
How the model learns the shared space
The model isn't told that a bark means 'dog'. It learns from huge numbers of paired examples: photos with their captions, audio clips with their transcripts, videos with descriptions. During training, the matching pairs are pulled closer together in the shared space and mismatched pairs are pushed apart. CLIP, a well-known model from OpenAI, was trained this way on images and their captions.
After enough examples, the photo of a dog, the sound of a bark and the word 'dog' all point to the same idea. That is what lets the model connect what it sees with what it reads and hears.
What it can do
Once everything shares one space, many tasks open up.
Understanding mixed input:
- Describe images, or read the text inside them.
- Answer questions about charts and notes, such as 'Which month had the most sales?' about a screenshot of a graph.
- Watch a video and explain what happened.
Going the other way. Give it text, and a multimodal system can generate images, speech or even video. The text is encoded into the shared space, and a generator turns that meaning into pixels or sound.
That makes multimodal models useful for:
- Medical analysis: looking at a scan alongside the patient's notes, as an aid to a doctor's judgement.
- Smarter customer support: understanding the screenshot a customer sends as well as their message.
- Creative tools: generating images, voices and video from a description.
- Accessibility: describing images for people who can't see them, or captioning speech for people who can't hear it.
A worked example: asking about an image
In practice you rarely build the encoders yourself. You send a hosted or open model a message that mixes kinds of content, and it handles the rest. Here is the shape of one request to Claude's API, with an image and a question in the same message:
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/jpeg",
"data": "<the photo, base64-encoded>"
}
},
{ "type": "text", "text": "What is this?" }
]
}The image and the text arrive together, so the model can answer the question about the picture, rather than about text alone. Other providers use the same idea with slightly different field names.
Trade-offs
- Cost. Images, audio and video turn into many more tokens than a short question. A single image can use as many tokens as several paragraphs of text, and video far more.
- Errors in a new form. A model that misreads a chart states the wrong number just as confidently as a text model gets a fact wrong, the problem covered in AI hallucinations.
- Not always needed. If your task is only text, a text model is usually simpler and cheaper.
Common mistakes
- Trusting small details in images. Models can miss fine print, count objects wrongly or misread handwriting. Check anything that matters.
- Sending huge files. Very large images are often resized before the model sees them. Crop to the part you care about.
- Assuming every model accepts every input. Many models read images but not audio, or read but can't generate. Check what a model supports before building on it.
- Forgetting privacy. Screenshots and photos often contain more than you meant to share, such as names, emails or faces.
Key takeaways
- Multimodal models work with more than one kind of input: text, images, audio and sometimes video.
- Each kind of data goes through its own encoder, which turns it into numbers.
- Those numbers are mapped into a shared space, so a dog photo, a bark and the word 'dog' point to the same idea.
- That lets them describe images, answer questions about charts and explain videos, and generate images, speech or video from text.
- They cost more to run, and can still misread what they see.