How Do Robots See the World?
ROBOT24.com unveils the magic behind machine vision.

When we see an apple on a table, we immediately know what it is and how far away it is.
A robot has to work all of this out from sensor data. First, it identifies the object and understands where it is. Then, it decides what to do with the object; for example, whether to move towards it, avoid it, or pick it up.
That's the problem machine vision helps robots solve. Let's break it down.
The Robot Captures an Image
First, the robot captures an image of its surroundings. A robot may use an RGB camera, depth camera, stereo cameras, LiDAR, or a combination of sensors.
Take a warehouse robot. A normal camera can show that there is a box in front of it, but the robot also needs to know how far away the box is. A depth camera provides that extra information.
A self-driving vehicle has a much harder job. Waymo's autonomous driving system uses cameras, LiDAR, and radar together. Cameras show what is around the vehicle. LiDAR maps the surroundings in 3D, while radar detects nearby objects and tracks their movement.
Robots often struggle with glass, hidden objects, shadows, and reflections. A 2026 study combined vision with camera movement and achieved a 96% success rate when grasping different types of objects.
The Image Is Processed
The captured image may be dark, noisy, blurred, or affected by shadows. The vision software processes the image to identify the objects in it.
Newer robots are also using event cameras. These cameras record changes in brightness instead of capturing complete frames. They are particularly useful for reducing motion blur when objects are moving and allow robots to work better in low-light conditions. Researchers have used them with robotic arms for grasping objects in difficult lighting.
Agricultural robots face strong sunlight and shadows. FieldNet, for example, is a real-time shadow-removal system developed for field robots. Reducing shadows helped improve weed detection.
The Robot Looks for Features
Next, the robot needs to identify the objects in the image.
The robot looks for useful features such as edges, colours, shapes, corners, and textures. For example, a robot sorting objects can use their shape, colour, size, and position to tell them apart.
The Robot Identifies the Object
Once the system finds these features, it can identify the object. A trained vision model compares the patterns in a new image with patterns it learned from many examples in its large database.
Latest vision models can also help robots recognise objects they have not seen before. Research such as DINOBot uses features from vision foundation models to help robots recognise and interact with unfamiliar objects.
3D Vision Tells the Robot How Far Away It Is
The robot has identified the object. Next, it needs to find its location in 3D space.
Depth cameras estimate how far objects are from the camera. Stereo cameras use two viewpoints to calculate depth, while LiDAR measures distances with laser pulses.
This is especially useful for robotic arms. Consider a robot reaching into a bin. The camera identifies a cup. The depth data tells the robot that the cup is 40 cm away, slightly to the left, and tilted at a certain angle. The robot can then find a suitable point to grasp it.
This is the basis of robotic bin picking, where objects are randomly piled together, and the robot must locate and grasp them.
Current systems are also combining vision with touch. Amazon's Vulcan uses stereo vision to estimate space in storage bins. The force sensors tell the robot when it has made contact with an object and how much force it is using.

The Robot Finds a Place to Grab
Knowing where the object is still isn't enough. The robot also needs to find where to place its gripper.
This becomes harder when objects overlap. If one box is partly hidden behind another, the robot has to use the visible information to estimate its position and find a safe grasp point.
Newer vision systems are starting to help with this, too. Google's Gemini Robotics-ER can use 3D and spatial understanding to identify an object and suggest an appropriate grasp and approach path. For example, when shown a mug, it can determine that the handle is a suitable place to grasp it.

The Robot Decides What to Do
A warehouse robot can spot a box, check its position, and move its arm towards it to pick it up. An autonomous vehicle may detect a pedestrian, track their movement, and adjust its path. An inspection robot may read a gauge and report an abnormal reading.
This is where machine vision connects with planning and control. Current vision-language-action systems are pushing this further. For instance, Google's Gemini Robotics can take a visual scene and a language instruction and then help a robot carry out physical actions. It can pick up and move objects and complete several steps in sequence. It can also perform tasks it was not specifically trained for.
The Robot Acts
Finally, the robot acts on its decision.
This has to happen quickly. ABB's conveyor-tracking systems can track moving objects at speeds of up to 1.7 metres per second, allowing robots to pick products without stopping the conveyor. The system can also connect multiple cameras, conveyors, and robots.
[1]https://www.tangramvision.com/blog/sensing-breakdown-waymo-jaguar-i-pace-robotaxi
More articles on this topic
156 articles
ReportRobots vs. Drones: What Is the Real Difference?
ReportHow Animals Inspire Better Robots
ReportFrom Toys to Tools: The Coolest Robots in Schools Today
ReportThis Tiny Robot Hops Like a Frog And Swims Using Elastic Limbs
ReportJapan Tests Robot Bins at Tokyo’s Busiest Railway Station
ReportOctopus-Inspired Muscle Design Simplifies Dexterous Soft Robots
Report5 Jobs Robots Could Take Over by 2035 — And 5 They’ll Struggle to Replace
ReportNvidia Develops Safety Controls as AI Advances Into Robotics
ReportWhy aren't robots in our homes yet?
ReportFlourish Opens Preorders for $3,555 Trainable Home Robot
ReportFive Extreme Sports Where Robots Are Competing
Report6 Unusual Jobs Done By RobotsBUSINESS NEWS WEEKLY LETTER
The Weekly Letter for Robotics Professionals, Summarizing the Most Important Industry Moves, Launches, Deals and Signals.
