Skip to content

Technology

How Humanoid Robots See: The Sensor Stack Behind Robot Perception

Humanoid robots perceive the world through five main sensor categories: RGB cameras, depth sensors, LiDAR, inertial measurement units, and force/torque sensors. Each sensor choice carries distinct tradeoffs in cost, capability, and operational range. Tesla Optimus uses eight cameras with no LiDAR; Boston Dynamics Atlas layers LiDAR with depth cameras; Unitree packs in every sensor type — and the design philosophy each company chooses shapes which markets their robots can serve.

By Cara Voss · March 17, 2026

How Humanoid Robots See: The Sensor Stack Behind Robot Perception

Every time a humanoid robot reaches for an object, avoids an obstacle, or holds its balance on uneven ground, it is drawing on a layered stack of sensors working in concert. Tesla Optimus Gen 3 relies on eight autopilot-derived cameras and nothing else for its external perception. Boston Dynamics Atlas combines LiDAR with depth cameras and inertial sensors. Unitree H1 layers 3D LiDAR, Intel RealSense depth, force-torque joints, and an IMU into one tightly integrated system.

These are not arbitrary choices. Each sensor philosophy carries tradeoffs in cost, weight, robustness, and the kinds of environments a robot can safely operate in. Understanding the sensor stack means understanding why today's humanoids behave the way they do — and where the next bottlenecks lie.

Humanoid robot navigating warehouse with point cloud overlay AI-generated image

A humanoid robot moving through a warehouse with a 3D point cloud overlay showing its real-time depth mapping of the environment.

Key Stats

8

Cameras on Tesla Optimus Gen 3

14x

LiDAR market growth for humanoids by 2030 (IDTechEx)

6-DoF

Force sensors per axis (typical 4 per robot)

5

Main sensor categories in humanoid robots

How Robot Sensing Works

Humanoid robots sense the world through two complementary channels. Exteroception covers everything outside the body: cameras capturing color and texture, depth sensors measuring distance, LiDAR painting detailed 3D maps of the environment. Proprioception covers the robot's own body state: inertial sensors tracking tilt and acceleration, joint encoders reading limb positions, force-torque sensors feeling contact pressure.

Neither channel alone is sufficient. A robot that can see its environment precisely but has no feel for its own balance will fall over. One with perfect proprioception but poor vision will collide with anything that moves. The skill in modern humanoid design is choosing which sensors to combine and how to fuse their outputs into a coherent, real-time model of the robot's situation.

The Five Sensor Categories

• RGB Cameras: The baseline sensor for most humanoids. Standard color cameras capture 2D images at high frame rates. Computer vision models — from classic YOLO object detectors to modern Vision Transformers — process these images to identify objects, people, text, and scene context. Cameras are cheap, light, and power-efficient, which is why vision-only designs remain attractive.

• Stereo Cameras: Two cameras mounted side-by-side, like human eyes. By comparing how the same scene looks slightly different in each lens (parallax), the robot can calculate depth for nearby objects. Stereo vision is effective for short-range 3D reconstruction without the cost of a dedicated depth sensor, but accuracy drops off past a few meters.

• Depth Sensors (RGB-D): Devices like the Intel RealSense line combine a standard color camera with a depth measurement system. Two main technologies apply here: structured light (projecting a known infrared pattern and measuring its distortion) and time-of-flight (measuring how long a light pulse takes to return). Both produce dense point clouds with per-pixel depth values up to roughly 10 meters, enabling real-time 3D mapping of workspaces.

• LiDAR: Pulses laser light in sweeping patterns and measures return time with high precision. A single mid-range LiDAR generates millions of 3D points per second, works in complete darkness, and maintains consistent geometry across lighting conditions. The tradeoffs are cost (mechanical LiDAR units remain expensive), weight, and the fact that LiDAR cannot read color or texture.

• IMU (Inertial Measurement Unit): Combines accelerometers, gyroscopes, and often a magnetometer to track orientation, acceleration, and angular velocity in real time. Every major humanoid robot uses an IMU as the backbone of its balance system. Without accurate IMU data, the robot cannot maintain gait stability or recover from disturbances.

• Force / Torque Sensors: Mounted at joints, wrists, ankles, and feet, these sensors measure the forces and moments acting on each contact point. They allow robots to adjust grip pressure when handling fragile objects, detect unexpected contact during manipulation, and provide the ground reaction data needed for stable bipedal walking. Unitree builds force-torque sensing into all joints of its G1 humanoid.

Under the Hood: The Sensor Fusion Pipeline

Raw sensor data is not useful on its own. Each sensor has noise, latency, and gaps. A camera cannot see through an opaque box. LiDAR cannot tell a white wall from a white door. An IMU drifts over time. The sensor fusion pipeline exists to take all of these imperfect inputs and produce a reliable, unified model of the robot's state and its surroundings.

The Three Fusion Streams

Modern humanoid systems process sensor data in three parallel streams that converge into a shared world model:

• Exteroception stream: Camera images and depth / LiDAR point clouds feed into 3D occupancy mapping. The robot maintains a voxel grid — a 3D grid where each cell is marked as free, occupied, or unknown. This is updated continuously as new sensor frames arrive. SLAM (Simultaneous Localization and Mapping) algorithms process this data to build a map of the environment while simultaneously estimating where in that map the robot is standing.

• Proprioception stream: IMU readings and joint encoder positions feed into a state estimator, typically an Extended Kalman Filter (EKF) or Unscented Kalman Filter (UKF). These algorithms optimally combine noisy measurements over time to produce smooth estimates of body orientation, velocity, and joint positions. This stream runs at high frequency — often 1000 Hz or more — because gait control requires precise, low-latency body state data.

• Haptic stream: Force and torque readings from contact points feed into the manipulation controller. When a robot grasps an object, the grip force is adjusted in real time based on whether the sensor detects slipping or crushing. At the feet, ground reaction forces feed back into the gait planner, enabling the robot to adapt its footfall patterns to uneven terrain.

What Is Zero-Moment Point (ZMP) Control?

ZMP is a concept from bipedal robotics that defines the point on the ground where the robot's net ground reaction force effectively acts. If the ZMP stays inside the robot's support polygon (the area between its feet), the robot is stable and will not tip over. IMU data and foot force sensors feed the ZMP estimator constantly, allowing the control system to shift weight, adjust step timing, and prevent falls before they start.

Vision-Only vs. Multi-Sensor: The Core Design Split

The biggest architectural debate in humanoid sensor design right now is whether to rely purely on cameras or to add active ranging sensors like LiDAR or depth cameras. Both camps have strong arguments.

Factor Vision-Only Vision + LiDAR / Depth
Cost Lower (cameras are commodity hardware) Higher (LiDAR units add hundreds to thousands of dollars)
Weight Lighter, better for endurance Heavier, impacts battery life
Low-light performance Poor without IR illumination Excellent (LiDAR is lighting-independent)
Depth accuracy Estimated via monocular/stereo inference Directly measured, more reliable
AI training data Vast (internet-scale image datasets) Sparse (labeled LiDAR datasets are expensive)
Texture / color Full color and texture LiDAR has no color; requires camera fusion
Scale to production Easier (fewer specialized components) More supply chain complexity
Range Limited by camera optics and AI inference Long-range (LiDAR to 100m+)

Tesla's bet on vision-only is the most aggressive application of this philosophy. The Autopilot camera system accumulated billions of real-world driving miles of training data before being adapted for Optimus. The neural networks running on custom silicon already know how to read depth, identify surfaces, and plan motion purely from pixel data. Scaling this to humanoid bodies is the technical challenge, but the dataset advantage is substantial.

Figure's Helix 02 system takes vision-only even further: a single neural network processes camera pixels and produces full-body motor commands directly, with no intermediate 3D map or explicit planning stage. This end-to-end approach trades interpretability for speed and simplicity.

Humanoid robot hand with force torque sensors AI-generated image

Force-torque sensors embedded in a humanoid robot's wrist and hand allow real-time manipulation feedback, enabling safe grasping and tactile awareness.

Who's Building the Sensor Stack

Each major humanoid maker has landed in a different part of the sensor design space, driven by their software capabilities, target use cases, and production cost goals.

👁️ Tesla Optimus Gen 3

Eight autopilot-derived cameras provide 360-degree visual coverage. No primary LiDAR. IMU for balance. Force sensors at the hands and feet for manipulation and gait feedback. The entire perception stack runs on Tesla's custom FSD (Full Self-Driving) neural processing silicon, repurposed from automotive applications. Target: mass production at automotive scale.

📷 Figure 02

Six cameras covering the robot's full field of view, plus microphones for audio sensing. Purely vision-based perception with no LiDAR or dedicated depth sensors. The Helix 02 AI system runs pixel-to-action control through a single neural network. Designed for warehouse environments where lighting is controlled and floors are flat.

🔴 Boston Dynamics Atlas (Electric)

LiDAR combined with stereo depth cameras and multiple IMUs. Atlas was designed for unstructured, outdoor, and dynamic environments where passive cameras alone are unreliable. The LiDAR provides lighting-independent geometry, while depth cameras fill in close-range detail. Higher cost, but wider operational envelope than vision-only systems.

⚡ Unitree H1 / G1

3D LiDAR, Intel RealSense RGB-D depth cameras, IMU, and force-torque sensors at all joints. Unitree builds the most sensor-dense humanoids at their price point. The G1 is particularly notable for its per-joint force sensing, enabling nuanced manipulation and terrain adaptation. Popular with research labs for this reason.

🏠 1X Neo

Camera-based with soft, compliant sensors throughout the body. Designed specifically for home environments where safe human interaction is paramount. The sensor philosophy prioritizes detecting unexpected contact and avoiding harm over long-range mapping. The Neo is built to move slowly and carefully in domestic spaces rather than operating at factory speed.

🔬 Intel RealSense (Platform)

Not a robot maker, but Intel's RealSense depth camera line has become the default depth sensor for research humanoids and mid-range commercial systems. Showcased at GTC 2026 as the backbone of multiple autonomous humanoid navigation stacks. Structured light and stereo depth variants available, with active IR illumination for low-light operation.

What Sensor Design Means for the Robot Industry

The sensor stack is not just a technical detail — it directly determines which markets a humanoid can serve and at what unit economics. A robot that needs expensive LiDAR cannot be priced for mass consumer deployment. A robot that cannot function in darkness cannot work in a warehouse with spotty lighting. A robot with no force sensing cannot safely hand an object to a human without risk of crushing it.

The BOM Cost Pressure

Bill of materials (BOM) cost is a constant constraint in robotics. LiDAR units that cost $1,000+ each push a robot's base cost well above consumer price points. This is one of the primary reasons Tesla, Figure, and other production-focused companies have moved toward vision-dominant designs. As AI improves at inferring depth from cameras, the performance gap between vision-only and LiDAR-equipped robots is narrowing.

IDTechEx projects 14x growth in the LiDAR market for humanoid robots by 2030. This sounds dramatic, but it reflects the current baseline being low, not an expectation that every humanoid will carry LiDAR. Much of that growth will come from research and industrial-tier robots where cost constraints are less severe.

Edge AI: Bringing Processing Onboard

Early robotics research often offloaded heavy computation to remote servers. The latency this introduces — even on fast local networks — creates problems for real-time balance and manipulation control. The industry is now firmly in an edge AI era: robots need to process sensor data onboard, locally, with millisecond latency.

NVIDIA's Jetson Orin platform has become a standard choice for research humanoids, providing GPU-accelerated inference in a power envelope compatible with battery-operated robots. Tesla uses its own custom silicon derived from the FSD chip. As edge AI chips improve, more sophisticated sensor fusion and perception models become feasible without cloud round-trips.

The Unstructured Environment Problem

Factory floors are the easiest deployment environment for humanoids: controlled lighting, flat surfaces, predictable objects, no children or pets. Homes and city streets are dramatically harder. Lighting is inconsistent. Floors have rugs, stairs, wet patches. Objects are unpredictable. Children and animals move suddenly. Every sensor category struggles more in unstructured environments, and sensor fusion becomes correspondingly more important. This is the core technical challenge that separates today's factory-optimized robots from the home-capable robots companies are promising for the late 2020s.

The 12-Month Outlook

The sensor stack for humanoids is changing faster than almost any other part of the technology. Several clear trends are accelerating through 2026 and into 2027.

• Vision-dominant consolidation: More companies are committing to camera-only or camera-primary designs. The combination of better AI depth estimation, larger training datasets, and cost pressure is making multi-sensor stacks harder to justify for production robots targeting price-sensitive markets.

• Solid-state LiDAR maturation: Traditional spinning LiDAR is expensive and mechanically fragile. Solid-state LiDAR (no moving parts) is falling in price and increasing in resolution. If solid-state units reach the sub-$100 range, the case for including LiDAR in production humanoids becomes much stronger, even for consumer-tier robots.

• Neural sensor fusion: Classical Kalman filter-based fusion is giving way to learned fusion models that can handle complex, nonlinear sensor interactions. Robots trained in simulation learn to weight sensor inputs dynamically based on which signals are most reliable in a given context. This means a robot can automatically down-weight its cameras when lighting is poor and rely more heavily on LiDAR data.

• Tactile skin: Full-body tactile sensing — distributed pressure sensors across a robot's exterior — is moving from research prototypes toward early commercial implementations. 1X Neo's soft sensor approach points in this direction. Whole-body tactile sensing would enable a new category of safe physical interaction that current force-torque joints cannot fully provide.

• GTC 2026 momentum: NVIDIA's GTC 2026 conference saw multiple humanoid makers showcasing Intel RealSense-based navigation stacks and new NVIDIA Isaac Perceptor packages for humanoid perception. The toolchain for building sophisticated sensor fusion pipelines is maturing rapidly, reducing the custom engineering work each company must do from scratch.

Frequently Asked Questions

Why does Tesla Optimus use cameras only and not LiDAR?

Tesla's primary reason is cost and scalability. LiDAR units add significant expense to each robot's bill of materials, which conflicts with Tesla's goal of producing humanoids at automotive scale. Tesla also has a massive advantage with camera-based AI: billions of miles of Autopilot training data have been used to build neural networks that infer depth, identify objects, and plan motion from cameras alone. Repurposing that expertise to Optimus is more practical than rebuilding a sensor fusion stack around new hardware.

What is an IMU and why do all humanoids need one?

An IMU (Inertial Measurement Unit) combines accelerometers and gyroscopes to track a robot's orientation, acceleration, and angular velocity in real time. Without an IMU, the robot's control system cannot know whether the body is tilting, accelerating forward, or rotating — all critical information for maintaining balance during walking. The IMU feeds the state estimator, which runs Extended Kalman Filter algorithms to produce smooth, low-latency body state estimates at rates of 500-1000 Hz. Zero-moment point (ZMP) controllers, which keep humanoids from falling over, depend entirely on this data stream.

What is SLAM and why does it matter for robot navigation?

SLAM stands for Simultaneous Localization and Mapping. It is the process by which a robot builds a map of an unknown environment while tracking its own position within that map at the same time. For humanoids, LiDAR-based SLAM is the most common approach: the robot spins its LiDAR, accumulates millions of 3D points, and uses algorithms like Cartographer or LOAM to stitch those scans into a consistent 3D map. Without SLAM, a robot must either have a pre-built map of its environment or navigate purely reactively based on what sensors see right now, which limits its ability to plan multi-step paths through complex spaces.

How do force-torque sensors make manipulation safer?

Force-torque sensors at the wrist and fingers measure the forces and moments acting on an object in all six degrees of freedom. When a robot grasps an object, the controller monitors grip force continuously. If the sensor detects that the object is starting to slip, it increases grip. If it detects excessive compression (risk of crushing), it reduces grip. At the feet, the same principle applies: ground reaction force data tells the balance controller exactly how much force each foot is applying, enabling real-time gait adjustments on uneven terrain. A robot without this feedback can only estimate contact forces — which leads to broken objects, spilled liquids, and unstable walking on irregular surfaces.

How long before humanoids can navigate homes as well as factories?

Factories are relatively easy because lighting is controlled, floors are flat, and objects are predictable. Homes are harder in almost every dimension: variable lighting, stairs, rugs, clutter, children, pets, and unexpected situations. Most industry observers estimate production-grade home navigation is 3-5 years away from where the leading systems are in 2026. The sensor challenge is manageable — it is the AI training data for the massive variety of home environments that is the harder problem. Some companies like 1X are specifically targeting home safety by using softer sensors and slower speeds, prioritizing avoiding harm over task efficiency.

Perception Is the Platform

The sensor stack a humanoid robot carries is not a feature list — it is a set of commitments about what problems the robot can solve, in what environments, at what cost. Tesla commits to cameras because it has the AI to make them work and the production ambitions that require cheap hardware. Atlas commits to LiDAR because it was designed for the hardest environments, where reliable geometry matters more than unit cost. Unitree commits to dense sensing because research customers need maximum data, not minimum weight.

What all of these systems share is the recognition that no single sensor is enough. Cameras cannot feel. LiDAR cannot see color. IMUs cannot see anything at all. Force sensors cannot map a room. Every perception system worth deploying in the real world fuses multiple data streams, cross-checks them against each other, and builds a richer picture of reality than any single sensor can provide alone.

The next few years will see this fusion get better and cheaper simultaneously. Solid-state LiDAR approaching commodity prices, edge AI chips powerful enough to run full neural fusion pipelines onboard, and training datasets that finally capture the messiness of real homes — these are the ingredients that close the gap between the factory robots of 2026 and the home robots of 2029.

The Bottom Line: The sensor choices inside a humanoid robot determine everything from where it can work to how much it costs to whether it can safely hand you a cup of coffee. Vision-only designs are winning the cost battle now, but multi-sensor systems are winning the capability battle in hard environments. As AI narrows the gap between these two approaches, the humanoid sensor stack will keep shifting — and so will which robots are actually useful outside a controlled factory floor.