ποΈ Multimodal Models & VLMs
Trace images into language-model inputs, calculate visual token costs, and test whether answers are supported by what the model can see.
On this page
Before you start
Linear Algebra explains vector similarity and matrix shapes. Transformers covers tokens and attention, and Deep Learning Basics covers supervised training and pretrained features.
You will learn to distinguish multimodal retrieval from generation, calculate image/video token budgets, compare visual-language interfaces, and evaluate grounding in a document or image task.
The problem: a plausible answer can ignore the image
Suppose a user asks for an invoice total. A language model can produce a plausible amount, but the correct answer depends on small printed digits and the layout identifying which amount is the total. If preprocessing makes those digits unreadable, fluent reasoning cannot restore the missing evidence.
A multimodal model handles more than one input/output modality. Vision-language models connect visual and language information, but they have different output contracts. CLIP-style embedding models return vectors for comparison; a generative VLM can answer questions or produce captions. A model being multimodal does not imply that it generates every modality or accurately understands every visual task.
First, work through an embedding comparison
Consider two unit image embeddings i_1=[1,0], i_2=[0,1] and two unit text embeddings t_1=[0.8,0.6], t_2=[0.6,0.8]. Unit length makes their dot products cosine similarities.
| Similarity | Text 1 | Text 2 |
|---|---|---|
| Image 1 | 0.8 | 0.6 |
| Image 2 | 0.6 | 0.8 |
With temperature Ο=0.1, Image 1's scaled scores are [8,6]. Its softmax probability for Text 1 is exp(8)/(exp(8)+exp(6)) β 0.8808. This is a relative training/retrieval score over these two candidates, not an 88.08% calibrated probability that the caption is true.
CLIP uses image-to-text and text-to-image softmax contrastive objectives, treating paired examples as positives against batch alternatives. SigLIP uses a pairwise sigmoid loss without a global softmax normalization: match/nonmatch labels contribute separate logistic terms. Negative selection, pair balance, data and scale still matter; sigmoid loss does not eliminate negatives or guarantee better features for every downstream task.
Trace an image into a generative VLM
A common connector design is:
image β vision encoder β visual feature sequence β learned connector β language model β answer
For a patch-based encoder with a 336Γ336-pixel image and 14Γ14-pixel nonoverlapping patches, the grid is 24Γ24 = 576 patch tokens, excluding any special tokens. Suppose each visual feature has width 1,024 and the language model uses width 4,096. A linear connector has weight shape 1024Γ4096 and maps 576Γ1024 features to 576Γ4096 embeddings.
Matching dimensions makes the tensors compatible. Training must still teach the decoder how to use the visual information. A learned connector can be linear, an MLP, or a query-based compressor; visual features can also enter through dedicated cross-attention.
Compare the interfaces, not just their names
| Interface | Information flow | Main cost or limitation |
|---|---|---|
| Projected visual tokens | Visual embeddings enter a joint token sequence | More sequence positions and attention/cache work |
| Dedicated cross-attention | Text queries attend to separate visual features | Extra layers, visual memory and cross-attention work |
| Learned query compressor | A smaller set of queries summarizes visual features | Fixed output budget can discard detail |
LLaVA illustrates connecting a vision encoder to a language decoder. Flamingo uses a resampler and cross-attention. BLIP-2's Q-Former learns a query-based interface between pretrained components. These components can be combined; resampling and cross-attention are not mutually exclusive categories.
In a standard dense joint-attention layer, attention pair count scales with the square of total sequence length. Cross-attention instead has work proportional to text-query count times visual-key count at its inserted layers, plus text self-attention and projection work. Visual features outside the text sequence are not free and still need storage and masking conventions.
Spend the resolution budget where it helps
Downscaling a whole page can erase small text. Alternatives include higher native resolution, region crops, tiling, multi-scale views or OCR/layout tools. Preserve crop coordinates and global context so a local view can still be related to the page.
At 576 patch tokens per crop, four tiles contribute 2,304 tokens. If the pipeline also uses one full-image thumbnail, the total is 2,880, before compression and special tokens. Four tiles are not automatically a model's native layout; count the actual processor output.
Doubling both image dimensions from 336 to 672 with the same 14-pixel patch size changes 576 tokens to 2,304. That is four times as many tokens; the dense visual-to-visual attention pair count alone rises 16-fold. This is not a 16-fold end-to-end latency prediction, because kernels, layer widths, text length, compression and other operations also matter.
There is no universal pixel threshold for reliable OCR. Font size in the original document, blur, crop scale, language and encoder training determine what survives. Test the smallest important text and the difficult layouts at the final preprocessing settings.
Train the interface for the task
One practical recipe first trains a connector on image/text pairs with pretrained backbones frozen, then instruction-tunes selected components on visual tasks. It can reduce adaptation cost, but is not the only valid training schedule. BLIP-2, LLaVA and other systems make different choices about objectives and frozen components.
Caption alignment teaches correspondence; instruction data teaches the expected interaction and answer format. Domain-specific work may benefit from unfreezing vision layers, but needs held-out evaluation for overfitting and forgetting. Preference optimization can help certain response behaviors without guaranteeing visual correctness. Training sample count and compute must be measured for the actual recipe, not inherited from a named architecture.
Ground an answer in a region
Visual grounding links a description to an image region, point or mask. Define the output convention: pixel or normalized coordinates, top-left origin, box ordering, resize/crop transforms and whether endpoints are inclusive. A coordinate-shaped answer may still point to the wrong object.
For a 100Γ100 image, use continuous-coordinate boxes. Ground truth [10,10,50,50] and prediction [20,10,60,50] each have area 1,600 square pixels. Their intersection is 30Γ40 = 1,200; union is 1,600+1,600β1,200 = 2,000. Intersection over union is 0.6. A thresholded score must name its IoU threshold and treatment of multiple objects.
Boxes show where evidence might be. They do not prove that an extracted value or sentence is correct. Link document answers to page/region evidence and separately validate text, units and schema.
Video adds time as well as tokens
Thirty-two frames at 576 tokens each yield 18,432 visual tokens before text or audio. Compressing to 64 tokens/frame reduces that part to 2,048, possibly losing details. One minute at one frame/second gives 60 frames and 34,560 tokens at 576/frame.
Uniform sampling can miss a brief event entirely. A gesture occurring from 0.2 to 0.4 seconds is absent from frames sampled only at 0 and 1 second. Choose temporal sampling for the task, preserve timestamps/order and test event-duration coverage. Motion-aware sampling, temporal encoders and hierarchical summaries are options with their own failure modes. Audio needs an aligned representation; a text transcript alone does not retain every sound cue.
Evaluate whether the answer used visual evidence
Separate reading/perception errors, region/layout errors and answer-generation errors. Compare matched examples whose images change while the question stays fixed; test absent objects, misleading text and unreadable regions. A model that gives the same answer to changed evidence may be relying on priors.
Measure task correctness, grounding, hallucinated objects/attributes/relations, citation support and appropriate uncertainty. Benchmarks such as DocVQA, ChartQA and object-hallucination probes cover different slices; no single score establishes all of them. Prompting the model to describe or reason about an image can help diagnose behavior but can also produce an invented explanation. Attention maps or confident prose are not proofs of grounding.
Check yourself
A system uses four 576-token document tiles and a global thumbnail, then budgets only 2,304 visual tokens. What did it miss?
Solution: The thumbnail adds another 576 tokens, so this hypothetical uncompressed pipeline needs 2,880. Inspect actual preprocessing, special tokens and compression before applying the count to a real model.
Where to go next
Embeddings & Retrieval builds search from vector representations. RAG combines evidence retrieval with generation. Vision-Language-Action Models adds an executable robot-action interface.
References
- Visual Instruction Tuning / LLaVA: connecting pretrained vision and language components.
- SigLIP: pairwise sigmoid versus softmax normalization.
- Flamingo: resampling and cross-attention for image/video-language inputs.
- BLIP-2: a query-based interface between frozen pretrained components.