🤖 Vision-Language-Action Models
Connect images and instructions to robot commands, calculate action precision and timing, and evaluate closed-loop behavior beyond a demo.
On this page
Before you start
Multimodal VLMs explains image/language representations. Deep Learning Basics covers supervised learning, and Calculus & Optimization helps with the flow-matching example.
You will learn to specify a robot action interface, compare discrete and continuous action generation, distinguish planning frequency from motor control, and design an evaluation that measures useful physical behavior.
The problem: “pick up the mug” must become a command
A vision-language model may identify the mug or describe a grasp. A vision-language-action model (VLA) maps visual observations and a language goal, usually with robot state/history, to actions that a controller can execute. Some vision-language models produce embeddings rather than text, so “VLM outputs text” is not a complete definition either.
The policy can be written as π(A given o, l, r), where o is visual observation/history, l is the instruction, r is robot state, and A is one action or an action chunk. A VLA is a policy; it need not contain a learned simulator or search over future plans.
Suppose its output is an end-effector delta: move [0.01, 0, 0] meters in the robot-base frame, hold orientation and close the gripper. That is a one-centimeter translation, not a joint torque. A controller must translate the target into feasible robot motion. The action schema needs coordinate frame, units, timestep, absolute-versus-relative semantics, limits and gripper convention. A valid tensor shape does not establish a valid motor command.
Trace one perception-to-action cycle
- Timestamp camera images and proprioception, the robot's measured joint/gripper state.
- Encode images and the instruction, preserving the camera/robot calibration needed by the task.
- Predict a command or future command sequence in the trained action representation.
- Unnormalize units, check command limits and pass it to the appropriate controller.
- Execute a bounded part, observe again, and handle success, interruption or recovery.
The low-level controller still manages tracking and robot-specific constraints. An image-conditioned policy does not make calibration, feedback, contact sensing or collision handling unnecessary.
Work out discrete action precision
A discrete policy can divide each action dimension into bins and predict their token IDs. Suppose a translation delta ranges from −0.05 m to +0.05 m and uses 256 equal-width bins reconstructed at their centers.
Bin width is 0.10/256 = 0.000390625 m, or 0.390625 mm. Within range, nearest-center reconstruction has error at most half a bin: 0.1953125 mm. Values outside the range clip and can have much larger error. This is representation error only; perception, calibration and motor tracking may contribute more.
Seven action dimensions predicted separately need seven tokens per timestep in a naive encoding. A 50-step chunk would need 350 such tokens. Compression tokenizers can reduce that count; discrete actions do not have to use one token per dimension per step.
Work out continuous action generation
Diffusion and flow-matching policies model a distribution over continuous, typically normalized action chunks. This can represent multiple valid trajectories without averaging them into an invalid middle path. It does not guarantee smoothness, physical feasibility or sub-millimeter accuracy.
For a simple flow-matching construction, let A_0 be a noise chunk, A_1 a demonstrated chunk, and τ a dimensionless interpolation time, not robot time. Train on A_τ=(1−τ)A_0+τA_1, predicting velocity A_1−A_0 conditioned on observation, instruction and robot state.
At inference, integrate the learned velocity field. A simple Euler step is A_next=A_current+Δτ×v_θ(A_current,τ,context). Each integration step evaluates the velocity model. The visual/language context may be cached, but ten integration steps are not a single action-head forward pass.
For a scalar teaching example, start at normalized value 0 and suppose the learned velocity is constantly 0.8. Four steps with Δτ=0.25 produce 0 → 0.2 → 0.4 → 0.6 → 0.8. Real learned fields vary with input and τ; the number of integration steps trades compute against approximation error and task behavior.
Diffusion policies use a denoising/noise-schedule formulation; some samplers are deterministic and others stochastic. Flow matching is not universally faster or more capable. Compare complete inference and closed-loop results.
A chunk is a horizon, not a control rate
Let H be predicted action steps, f the intended command frequency in Hz, and K the number executed before replanning. The chunk horizon is H/f seconds; the open-loop execution interval is K/f seconds.
For H=50, f=50 Hz and K=5, the policy predicts one second of commands but executes 0.1 seconds before taking another observation. This is receding-horizon execution; it does not blindly execute all 50 steps. A 60 ms inference fits within that 100 ms planning interval only if sensing, queuing and transfer fit in the remaining 40 ms. Measure tail latency and handle late/stale chunks.
A separate 200 Hz servo would have a 5 ms update interval. That does not require the large VLA to run at 200 Hz if the controller safely tracks a suitable command reference. Policy frequency, actuator command rate and end-to-end reaction latency are different measurements.
Learn from published designs without ranking demos
| Published design | Useful idea | Limit of the conclusion |
|---|---|---|
| RT-2 | Co-fine-tune vision-language and robot data with actions expressed as tokens | Reported transfer is specific to tasks and evaluation setup |
| OpenVLA | Accessible 7B VLA, fused visual features and adaptation workflow | Shared weights do not remove embodiment/calibration differences |
| π0 | A pretrained VLM combined with a flow-matching action expert | Continuous output alone does not establish reliable dexterity |
| FAST | Discrete cosine transform based action-sequence tokenization | Compression still has representation error and decoding costs |
FAST is a discrete action tokenizer, not a continuous flow head or an FFT-only representation. Its existence is a useful reminder that “discrete versus continuous” is not “incapable versus dexterous.” Model size or a public demonstration cannot substitute for a controlled comparison.
Data quality includes time and embodiment
Demonstrations need images, instructions, action semantics, timestamps and robot state. Shared collections such as Open X-Embodiment help pool data, but a common storage format does not make joint spaces, control rates or coordinate frames identical. Inspect the actual training mixture; a model may use a curated subset rather than every episode in a collection.
Use consistent normalization, embodiment identifiers or adapters where needed, and train/test splits by task, scene, object and collection session. A model may transfer visual recognition while needing new data for a gripper or kinematic chain. Collect failure recovery and corrections, not only successful demonstrations. Simulation and world-model rollouts can extend coverage, but contact and sensor errors must be checked against real evidence.
Evaluate behavior after the robot moves
Offline action error measures agreement with demonstrations; a different valid grasp may score poorly, while a small error can lead to a large closed-loop failure. Evaluate task completion, instruction grounding, contacts/constraint violations, intervention rate, recovery, duration and latency on repeated trials.
Keep held-out objects, scenes, instructions and embodiments distinct. Define success before trials, record all attempts, and use human review where the task outcome is ambiguous. Simulation benchmarks are useful controlled environments; they do not prove transfer to a particular robot. Report uncertainty across independent trials and failures by stage: perception, grounding, planning interface, control and recovery.
Check yourself
A policy emits 40 commands intended for 20 Hz execution. You execute four commands before replanning. What are the horizon and open-loop interval?
Solution: The horizon is 40/20 = 2 seconds; the open-loop interval is 4/20 = 0.2 seconds. Neither number states the model's measured inference speed. That still needs profiling, including worst-case delays and stale-observation handling.
Where to go next
World Models adds prediction for imagined rollouts or planning. Evaluation & Benchmarking helps compare policies without overstating a small set of demonstrations.